25 Frequency Distributions

Makenzie O'Neil

Learning Objectives

By the end of this chapter, you will be able to

  1. Create frequency distribution tables and graphs to describe and display the distribution of a variable.
  2. Interpret frequency distributions in order to interpret and draw reasonable conclusions about various datasets.

As described in the previous chapter, Descriptive statistics refers to a set of techniques for summarizing and displaying data. Let us assume here that the data are quantitative and consist of scores on one or more variables for each of several study participants. Although in most cases the primary research question will be about one or more statistical relationships between variables, it is also important to describe each variable individually. For this reason, over the next few chapters, we will begin by looking at some of the most common techniques for describing single variables.

The Distribution of a Variable

Every variable has a distribution, which is the way the scores are distributed across the levels of that variable. For example, in a sample of 100 university students, the distribution of the variable “number of siblings” might be such that 10 of them have no siblings, 30 have one sibling, 40 have two siblings, and so on. In the same sample, the distribution of the variable “sex” might be such that 44 have a score of “male” and 56 have a score of “female.”

Frequency Tables

One way to display the distribution of a variable is in a frequency table. Table 25.1, for example, is a frequency table showing a hypothetical distribution of scores on the Rosenberg Self-Esteem Scale for a sample of 40 college students. The first column lists the values of the variable—the possible scores on the Rosenberg scale—and the second column lists the frequency of each score. This table shows that there were three students who had self-esteem scores of 24, five who had self-esteem scores of 23, and so on. From a frequency table like this, one can quickly see several important aspects of a distribution, including the range of scores (from 15 to 24), the most and least common scores (22 and 17, respectively), and any extreme scores that stand out from the rest.

Table 25.1. Frequency Table Showing a Hypothetical Distribution of Scores on the Rosenberg Self-Esteem Scale
Self-esteem Frequency
24 3
23 5
22 10
21 8
20 5
19 3
18 3
17 0
16 2
15 1

There are a few other points worth noting about frequency tables. First, the levels listed in the first column usually go from the highest at the top to the lowest at the bottom, and they usually do not extend beyond the highest and lowest scores in the data. For example, although scores on the Rosenberg scale can vary from a high of 30 to a low of 0, Table 25.1 only includes levels from 24 to 15 because that range includes all the scores in this particular data set. Second, when there are many different scores across a wide range of values, it is often better to create a grouped frequency table, in which the first column lists ranges of values and the second column lists the frequency of scores in each range. Table 25.2, for example, is a grouped frequency table showing a hypothetical distribution of simple reaction times for a sample of 20 participants. In a grouped frequency table, the ranges must all be of equal width, and there are usually between five and 15 of them. Finally, frequency tables can also be used for categorical variables, in which case the levels are category labels. The order of the category labels is somewhat arbitrary, but they are often listed from the most frequent at the top to the least frequent at the bottom.

Table 25.2. A Grouped Frequency Table Showing a Hypothetical Distribution of Reaction Times
Reaction time (ms) Frequency
241–260 1
221–240 2
201–220 2
181–200 9
161–180 4
141–160 2

Frequency Graphs

Pie Charts

In a pie chart, each category on a scale is represented by a slice of the pie. The area of the slice is proportional to the percentage of responses in the category. Pie charts are effective for displaying the proportion of responses for a small number of categories. For instance, the pie chart shown in Figure 25.1 represents data from 500 iMac customers, categorized as a previous Mac owner, a previous Windows owner, or a new computer purchaser.

Pie chart of operating systems—Macintosh 71%, None 17%, Windows 12%
Figure 25.1. Pie chart showing the percentage of iMac purchasers who previously owned a Macintosh computer, a Windows computer, or no computer.

Pie charts are not recommended, however, when you have a large number of categories. Additionally, if a pie chart is based on a small number of observations, it can be misleading to label the pie slices with percentages. For example, if just 5 people had been interviewed by Apple Computers, and 3 were former Windows users, it would be misleading to display a pie chart with the Windows slice showing 60%. With so few people interviewed, such a large percentage of Windows users might easily have occurred since chance can cause large errors with small samples. In this case, it is better to alert the user of the pie chart to the actual numbers involved. The slices should therefore be labeled with the actual frequencies observed (e.g., 3) instead of with percentages.

Bar Charts

Bar charts can also be used to represent frequencies of different categories. A bar chart of the iMac purchases is shown in Figure 25.2. Frequencies are shown on the y-axis, and the type of computer previously owned is shown on the x-axis. The bars should not be touching, which indicates separate categories on the x-axis. Typically, the y-axis shows the number of observations in each category rather than the percentage of observations in each category, as is typical in pie charts.

 

Bar chart of buyers’ previous computer: Macintosh ~350, None ~80, Windows ~60 (image description available)
Figure 25.2. Bar chart of iMac purchases as a function of previous computer ownership. [Image Description]

Histograms

histogram is a graphical display of a distribution used for quantitative (interval and ratio) data. It presents the same information as a frequency table, but in a way that is even quicker and easier to grasp. The histogram in Figure 25.3 presents the distribution of self-esteem scores in Table 25.1. The x-axis of the histogram represents the variable, and the y-axis represents frequency. Above each level of the variable on the x-axis is a vertical bar that represents the number of individuals with that score. When the variable is quantitative, as in this example, there is usually no gap between the bars. When the variable is categorical, however, a bar graph is used, and there is usually a small gap between the bars. (The gap at 17 in this histogram reflects the fact that there were no scores of 17 in this data set.)

Histogram of self-esteem scores (15–24) with a peak around 22 and fewer cases at the extremes. (image description available)
Figure 25.3. Histogram Showing the Distribution of Self-Esteem Scores Presented in Table 25.1 [Image Description]

Similar to a frequency table, when there are many different scores across a wide range of values, it is often better to create a grouped histogram, in which the y-axis still represents frequency, but the x-axis represents a range of values for each bar. Each bar represents a range of scores broken into intervals, called class intervals. This can be seen in Figure 25.4, where the data represent the scores of 642 students on a psychology test ranging from 46 to 167; each class interval represents a range of 10 points on the scale, each class interval represents a range of 10 points on the scale.

 

A histogram of scores on a psychology test, with most scores in the center of the distribution and a positive skew. (image description available)
Figure 25.4. Histogram of scores on a psychology test. [Image Description]

Frequency Polygons

Frequency polygons are a graphical device for understanding the shapes of distributions. They serve the same purpose as histograms, but are especially helpful for comparing sets of data. Frequency polygons are also a good choice for displaying cumulative frequency distributions.

To create a frequency polygon:

  • Start just as for histograms, by choosing a class interval.
  • Draw an x-axis representing the values of the scores in your data. Mark the middle of each class interval with a tick mark, and label it with the middle value represented by the class.
  • Draw the y-axis to indicate the frequency of each class. Place a point in the middle of each class interval at the height corresponding to its frequency.
  • Finally, connect the points. You should include one class interval below the lowest value in your data and one above the highest value. The graph will then touch the x-axis on both sides.

The frequency distribution of 642 psychology test scores, shown in Figure 25.4, was used to create the frequency polygon shown in Figure 25.5.

The first label on the x-axis is 35. This represents an interval extending from 29.5 to 39.5. Since the lowest test score is 46, this interval has a frequency of 0. The point labeled 45 represents the interval from 39.5 to 49.5. There are three scores in this interval. There are 147 scores in the interval that surrounds 85. You can easily discern the shape of the distribution from Figure 25.5. Most of the scores are between 65 and 115.

 

Frequency polygon of test scores with a peak near 85 and a long right-skewed tail. (image description available)
Figure 25.5. Frequency polygon for the psychology test scores. [Image Description]

Frequency polygons are useful for comparing distributions. This is achieved by overlaying the frequency polygons drawn for different datasets. Figure 25.6 provides an example. The data come from a task in which the goal is to move a computer cursor to a target on the screen as fast as possible. In 20 of the trials, the target was a small rectangle; in the other 20, the target was a large rectangle. Time to reach the target was recorded on each trial. The two distributions (one for each target) are plotted together in Figure 25.6. The figure shows that, although there is some overlap in times, it generally took longer to move the cursor to the small target than to the large one.

Frequency polygons of response times: large target peaks faster (~550 ms); small target shifts slower (~650–850 ms). (image description available)
Figure 25.6. Overlaid frequency polygons for the cursor task. [Image Description]

Distribution Shapes

When the distribution of a quantitative variable is displayed in a histogram, it has a shape. The shape of the distribution of self-esteem scores in Figure 25.4 is typical. There is a peak somewhere near the middle of the distribution and “tails” that taper in either direction from the peak. The distribution of Figure 25.4 is unimodal, meaning it has one distinct peak, but distributions can also be bimodal, meaning they have two distinct peaks. Figure 25.7, for example, shows a hypothetical bimodal distribution of scores on the Beck Depression Inventory. Distributions can also have more than two distinct peaks, but these are relatively rare in psychological research.

Histogram of Beck Depression Inventory scores: peaks at 10–19 and 50–59; few in 30–39. (Image description available)
Figure 25.7. Histogram Showing a Hypothetical Bimodal Distribution of Scores on the Beck Depression Inventory. [Image description]

Another characteristic of the shape of a distribution is whether it is symmetrical or skewed. The distribution in the center of Figure 25.8 is symmetrical. Its left and right halves are mirror images of each other. The distribution on the left is negatively skewed, with its peak shifted toward the upper end of its range and a relatively long negative tail. The distribution on the right is positively skewed, with its peak toward the lower end of its range and a relatively long positive tail.

Three mini histograms labeled: Negatively Skewed, Symmetrical, Positively Skewed.(image description available)
Figure 25.8. Histograms Showing Negatively Skewed, Symmetrical, and Positively Skewed Distributions [Image Description]

An outlier is an extreme score that is much higher or lower than the rest of the scores in the distribution. Sometimes outliers represent truly extreme scores on the variable of interest. For example, on the Beck Depression Inventory, a single clinically depressed person might be an outlier in a sample of otherwise happy and high-functioning peers. However, outliers can also represent errors or misunderstandings on the part of the researcher or participant, equipment malfunctions, or similar problems. We will say more about how to interpret outliers and what to do about them in the next chapter.

Equity Activity: Income inequality in urban schools

A university has collected data on student participation in extracurricular activities (e.g., clubs, sports teams, community service, etc.). The data shows the number of students from different ethnic groups who participate in extracurricular activities. The university is concerned about potential disparities in involvement across different ethnic groups.

Below are two pie charts showing the number of students enrolled at the university and the number of students who participate in extracurricular activities, broken down by ethnic group:

Two pie charts compare extracurricular participants vs. total university population by ethnicity (image description available)
Figure 25.9. Extracurricular Participation by Ethnic Group [Image Description]

Online Descriptive Statistics

Although many researchers use commercially available software such as SPSS and Excel to analyze their data, there are several free online analysis tools that can also be extremely useful. Many allow you to enter or upload your data and then make one click to conduct several descriptive statistical analyses. Among them are the following.

Rice Virtual Lab in Statistics

VassarStats

Bright Stat

For a more complete list, see the Statpages website.

License & Attribution

“Frequency Distributions” by Makenzie O’Neil is adapted from “Describing Data Using Distributions & Graphs” by Linda R. Cote Ph.D.; Rupa G. Gordon Ph.D., Chrislyn E. Randell Ph.D., Judy Schmitt, and Helena Marvin, which is licensed CC BY-NC-SA 4.0.

“Frequency Distributions” is licensed under CC BY-NC-SA 4.0.


Image Descriptions

Figure 25.2. Vertical bar chart with Number of Buyers on the y-axis (0–400) and Previous Computer on the x-axis (categories: None, Windows, Macintosh). Bars show approximate counts: None ≈ 80, Windows ≈ 60, Macintosh ≈ 350. Macintosh dominates, with far more buyers than the other categories. [Return to Figure 25.2]

Figure 25.3. A histogram shows Self-Esteem on the x-axis (scores 15 to 24) and Frequency on the y-axis (0–12). Bars rise from low counts at 15–18, increase around 19–21, peak near 22 (≈11 cases), then decline at 23–24 (mid-single-digit counts). The distribution is unimodal and clustered in the low-20s. [Return to Figure 25.3]

Figure 25.4. Vertical histogram with Test Score on the x-axis (bins from about 40 to 170) and Frequency on the y-axis (up to 150). Bars are low at the left (≈40–60), rise through 60–80, peak near 80–90 at roughly 140–150 cases, then steadily decline across 90–120 and taper to small counts from 130 to 160+. The distribution is unimodal and positively skewed (right-skewed). [Return to Figure 25.4]

Figure 25.5. Line graph (frequency polygon) with Test Score on the x-axis (35–175) and Frequency on the y-axis (0–160). Open-circle points connect to form a sharp rise from near zero at 35–55, increasing to about 50 at 65, 105 at 75, and peaking around 150 at 85. The line then declines: ~130 at 95, ~80 at 105, ~58 at 115, ~36 at 125, ~10 at 135, and approaches zero by 155–175. The distribution is unimodal and positively skewed (right-skewed). [Return to Figure 25.5]

Figure 25.6. Line chart with Time (msec) on the x-axis (350–1150) and Frequency on the y-axis (0–10). Two polygons are plotted: a teal dashed line with open circles labeled Large target and a magenta solid line with open squares labeled Small target. The large-target distribution rises from 350 ms to a peak frequency of 10 at ~550 ms, then drops to zero by 750 ms. The small-target distribution begins near zero at 450 ms, climbs to ~6 at 650 ms, stays around 5 through 750–850 ms, and then declines, with a small bump near 1050 ms. Overall, times are faster (shorter) for the large target and slower for the small target. [Return to Figure 25.6]

Figure 25.7. Histogram-style bar chart with Beck Depression Inventory Score bins on the x-axis (0–9, 10–19, 20–29, 30–39, 40–49, 50–59, 60–69) and Frequency on the y-axis (0–16). Approximate counts by bin: 0–9 ≈ 3, 10–19 ≈ 14 (largest), 20–29 ≈ 6, 30–39 ≈ 1, 40–49 ≈ 3, 50–59 ≈ 12 (second largest), 60–69 ≈ 3. The distribution is bimodal, concentrated in the 10–19 and 50–59 ranges. [Return to Figure 25.7]

Figure 25.8. Side-by-side silhouettes show distribution shapes. Left (green) has a long left tail and a peak toward the right and is negatively skewed. Center (blue) is balanced with a central peak and is symmetrical. Right (orange) has a long right tail and a peak toward the left and is positively skewed. [Return to Figure 25.8]

Figure 25.9. Side-by-side pie charts with the headings “Participation in Extracurricular” (left) and “Total University Population” (right). A legend is located below maps colors: White (dark blue), Black/African American (orange), Hispanic/Latino (green), Asian/Pacific Islander (red), Native American (purple), Other (teal). The pie graph entitled Participation in Extracurricular counts by slice are White 170, Black/African American 100, Hispanic/Latino 50, Asian/Pacific Islander 30, Native American 30, Other 20. The pie graph entitled Total University Population counts by slice are White 500, Black/African American 250, Hispanic/Latino 300, Asian/Pacific Islander 250, Native American 50, Other 100. [Return to Figure 25.9]

definition

License

Icon for the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License

Critical Research Methods in Psychology Copyright © 2025 by Stephanie D'Costa; Mireille Ukeye; Makenzie O'Neil; and Rebecca Anguiano is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License, except where otherwise noted.