Statistics I dealt with data you could list. Once there is too much of it, the data gets grouped into classes, and every calculation becomes an estimate. That trade — detail for manageability — is the idea behind this whole topic.
Grouped frequency tables
Forty students sat a test marked out of 50:
Mark
Frequency
1–10
4
11–20
7
21–30
12
31–40
9
41–50
8
Three things to be able to state about any class:
Class limits — the values as written, so 21 and 30 for the middle class.
Class boundaries — where one class really ends and the next begins. Marks are whole numbers recorded to the nearest unit, so the 21–30 class runs from 20.5 to 30.5.
Class width — upper boundary minus lower boundary: 30.5−20.5=10.
The midpoint is the average of the two limits:
midpoint=221+30=25.5
The modal class is simply the class with the highest frequency — here 21–30, with 12 students. You cannot name a single modal value from grouped data, because the individual marks are no longer visible.
Mean of grouped data
Since the actual values are lost, assume every value in a class sits at its midpoint. That makes the answer an estimate.
estimated mean=∑f∑fxwhere x is the midpoint
Mark
f
Midpoint x
fx
1–10
4
5.5
22
11–20
7
15.5
108.5
21–30
12
25.5
306
31–40
9
35.5
319.5
41–50
8
45.5
364
Total
40
1120
mean=401120=28
Histograms and frequency polygons
A histogram displays grouped continuous data. It looks like a bar chart with one crucial difference: the bars touch, because the classes are continuous and there is nothing between them.
Draw the bars between the class boundaries, not the limits.
With equal class widths the height is the frequency.
Label both axes.
A frequency polygon joins the midpoints of the tops of the bars with straight lines. It is useful for comparing two distributions on one diagram, since two polygons can be drawn over each other where two sets of bars would collide.
Cumulative frequency
Cumulative frequency is a running total — how many values fall at or below each point.
Mark
Frequency
Cumulative frequency
≤10.5
4
4
≤20.5
7
11
≤30.5
12
23
≤40.5
9
32
≤50.5
8
40
So 23 students scored 30 or fewer, and the final cumulative frequency must equal the total, 40.
Plotting cumulative frequency against the upper class boundary and joining the points with a smooth curve gives an ogive.
Median and quartiles
An ogive gives you three key readings. With n values:
Median (Q2) — read across from 2n
Lower quartile (Q1) — read across from 4n
Upper quartile (Q3) — read across from 43n
interquartile range=Q3−Q1
The interquartile range measures the spread of the middle half of the data, so unlike the range it is unaffected by one extreme value.
For the ordered data 3,5,7,8,10,12,15,18:
There are 8 values, so the median is the mean of the 4th and 5th:
Q2=28+10=9
Split the data at the median. Lower half 3,5,7,8, so
Q1=25+7=6
Upper half 10,12,15,18, so
Q3=212+15=13.5
IQR=13.5−6=7.5
Probability
For equally likely outcomes,
P(event)=total number of outcomesnumber of favourable outcomes
Every probability lies between 0 and 1. An impossible event has probability 0; a certain one has probability 1.
A fair die is rolled. Find the probability of an even number.
The even numbers are 2, 4 and 6 — three of six outcomes:
P(even)=63=21
A bag holds 4 red, 3 blue and 5 green marbles. One is drawn at random.
There are 12 marbles altogether.
P(blue)=123=41
P(not green)=124+3=127
Check: P(green)=125, and 1−125=127 ✓
Two events at once
List the outcomes in a table or grid. Two dice have 6×6=36 equally likely totals. The pairs giving a total of 7 are (1,6),(2,5),(3,4),(4,3),(5,2),(6,1) — six of them: