Showing posts with label tally marks. Show all posts
Showing posts with label tally marks. Show all posts

Tuesday, February 2, 2016

Chapter 1.1 - Analysing large data

In the previous section we saw a simple example of a Frequency distribution table and also a Bar graph. Now we will see another example. This time the Class teacher of class VIII, Division C, wants to know the details about the heights (in cm) of the 32 students in that class.

The first step is to collect the raw data. As the students belong to a particular class, the heights can be collected in serial order, according to the roll number of the students. But even if it is collected in serial order, the data is 'raw' because the different heights will be distributed randomly in the list. The following list shows the raw data:


The Frequency distribution table and the Bar graph are shown in fig.1.9 below:      
Frequency distribution table and bar graphs help to condense and analyse data.
Fig.1.9 Frequency distribution table and Bar graph
The figs. self explanatory. We can see that 155 cm is the smallest height and 165 is the largest height. Six students have a height of 157 cm. Another six students have a height of 158 cm.

• In the above example, the data from a single division of class VIII was considered. So the raw data list is small. 
• The values in the raw data list are close to each other. 155 is the smallest value and 165 is the largest value. 
• There are values in between these two. There are 6 'in between values'. 
• So there are a total of 8 values. Thus there are 8 rows in the Frequency distribution table, and there are 8 bars in the Bar graph. 
• Also note that the values like 157, 158, 159 etc., repeated several times. So many 'tally marks' were accommodated into those rows.

Because of these properties, we obtained a small Table and a small Bar graph. In some cases we get a list of raw data with a large number of values.
• The values may not be close to each other. For example, one value may be 52, another may be 300, and yet another 520. 
• There may be a vast difference between the smallest and the largest value. For example, the smallest may be 45 and the largest may be 580.
• As the list is large, the number of 'in between values' will also be large.
• If there are not many largely repeating values among these 'in between values', we will have to provide a large number of rows in the table. There will also be a large number of bars in the Bar graph.

In such situations, we divide the list into 'groups'. For example, suppose there are a total of 360 values in the raw data. 
• If we decide to form groups, each with 20 values, the total 360 values will become 18 groups.
• If we decide to form groups, each with 40 values, the total 360 values will become 9 groups.

Thus we see that the large number of values are made into small number of  manageable values. Now the question arises: How to form these groups?

The simplest way may seem to make the first 20 into the first group, the next 20 into the second group and so on. But such a grouping will not be of any help in 'analysing' the data. The raw data will still be 'raw'. We need a more 'scientific' method.

We can think of grouping based on 'similarity in properties'. That is., all the members of a group will be having similar properties. This can be explained with the help of an example. Suppose that  there are 360 items in a raw data list. The items are the Electricity bill amount collected from 360 houses in a town in a particular month. [All the bills should be of a particular month. This is because, the electricity consumption in summer months will be greater than in winter months due to the usage of cooling appliances. If we collect the bill from some houses in a summer month, and some others in a winter month, the actual consumption of the whole town cannot be analysed] We want to divide the 360 items into groups. We want to do the division based on 'similarity in properties'.

This can be done as follows: The amounts which are close to each other should come in a group. For example, if a value in a group is 142, the other values that come in the group should be those like 120, 135, 160 etc., Other extreme values like 20, or 450 should not come in the group. 20 and 450 should fall into other appropriate groups. Let us see how this can be done:

We know the 'Number line'. A line in which the counting numbers are arranged in sequential order. The positive portion of it starts from the minimum value of zero, and can go up to any maximum value. This is shown in the fig.1.10 below:
Fig.1.10 Number line from zero to 8
Because of the limitation in space, we can show only upto 8 or 9. But this can be solved by changing the 'scale'. Thus, here is another number line which shows upto 80:
Fig.1.11 Number line from zero to 80
In fig.1.10, one unit represents '1'. But in fig.1.11, one unit represents '10'. In this way we can assume that one unit represents '50' or '100' to show up to 500 or 1000 in a small space.

So now we know about the Number line. Each of the values in the raw data list will fall onto a unique place in the Number line. Now, we divide the number line into equal intervals. This division can be done in many ways. One way is equal intervals of 20. Then, the intervals will be: 0 - 20, 20 – 40, 40 – 60, 60 – 80 etc., This is shown in the fig below:
Fig.1.12 Equal intervals of '20' on the Number line
If one of the values in the raw data list is 32, it will fall in the first red interval. If another value is 53, it will fall in the second green interval. It is possible that, a value which is lying in the last portion of the raw data list will find it's place in the first green interval, if it has a low value (from 0 to 20). 

• Lower values will find their final places in any one of the intervals towards the left of the number line. 
• Higher values will find their final places in any one of the the intervals towards the right of the number line.

So this is a scientific method of grouping. In the table shown in fig.1.9, which we saw earlier above, a row is assigned to every single value in the raw data list. (If a value repeats more than once, it need to be shown in one row only. Even then, for large data lists, the number of rows will become very large) 

But in the above method, the 'intervals' take the place of 'values'. So the number of rows will become small. Once the first column of the table is filled up with the appropriate intervals, 'tally marking' can begin. We will see an example:

Given below is a raw data list of electricity bills of a particular month collected from 38 houses in a locality.



Let us prepare the Frequency distribution table. The smallest entry in the raw data list is 74. The largest entry is 389. So we want the portion from 74 to 389 in the number line. This is shown in the fig.1.13 below:
Fig.1.13 Portion from 74 to 389
• The required portion is marked in red color. 
• The 'length' of this red portion is 389 -74 = 315 
• We want equal intervals in this red portion. Let us try to obtain 10 equal intervals. Then each interval will be 315/10 =31.5
• It is not convenient to mark equal intervals at 31.5. We must try to use multiples of 2, 5, 10, 50 or 100. The closest multiple to 31.5 is 50.
• So we will divide the red portion into equal intervals of 50
• The division of the red portion in this manner should coincide with the divisions on the number line also.

Based on the above, we will get the arrangement as shown below:
Fig.1.14 Equal divisions inside the Required portion
The required portion now contains equal divisions of 50. This is from 100 to 350. There is a portion before 100 and another portion after 350 which do not come in the 'equal divisions'. All the divisions must be 'equal'. So we will extend the red portion to either side upto 50 and 400. Thus the final form will be as shown below:

So, instead of 10, we have obtained 7 equal intervals. We can now prepare the Frequency distribution table. We will see this in the next section.

PREVIOUS       CONTENTS       NEXT                                          


Copyright©2016 High school Maths lessons. blogspot.in - All Rights Reserved

Chapter 1 - Frequency distribution tables and Bar graphs

On many occasions, we may want to collect information and do an analysis of it. For example a Canteen manager in an University campus may want to know which flavours of ice cream are more popular among students. To find it, he sets about to 'collect the information'. The easiest way to do this is to ask the students themselves. So he meets each student and notes down their preferred flavour. The information that he collects on his note book may look like this:

Student 1: Strawberry, Student 2: Vanilla, Student 3: Chocolate, Student 4: Vanilla, Student 5: Butter scotch ..... Student 25: Chocolate.

He collected information from 25 students. Here, the details about students is not necessary. So he needs to note down the flavour only with a serial number. So the list will look like this:

1. Strawberry, 2. Vanilla, 3. Chocolate, 4. Vanilla, 5. Butterscotch, ....... 25. Chocolate.

The information collected in this way is called Data. By just looking at this list, we cannot arrive at a conclusion about the most preferred flavour. This is particularly so, if the list is large. As it is a random list, it is called Raw Data. The raw data will be taken from the 'field' (which is the University campus in our case), to the office. There it will be 'analysed' and made into suitable 'mathematical models'. By this process, many important conclusions can be made. We will now discuss the methods to make some simple and basic mathematical models from the raw data.

The first step is to make an 'ordered list' from the raw data. For this, a table is drawn up as shown in fig.1.1. 
Fig.1.1 Table for making ordered list
The first column shows the 'Flavour'. The second column shows the  Tally marks. The third column shows the 'Number of students'. The first column is easy, and is already filled here. Let us now fill up the second column. For this the complete raw data collected from the 25 students is given below:

Based on this list, we now put the tally marks in the second column. The steps are as follows:

• Take the first entry in the raw data list. It is 'Strawberry'.
• Find 'Strawberry' in the first column.
• Put a '|' mark on the second column 'in line' with strawberry. This is shown in the fig.1.2.
Fig.1.2 Tally mark for the first entry in the Raw data list

Repeat the process:
• Take the second entry. It is 'Chocolate'.
• Find 'Chocolate' in the first column.
• Put a '|' mark on the second column 'in line' with Chocolate. This is shown in the fig.1.3 
Fig.1.3 Tally mark for the second entry in the  data list

Repeat the process:
• Take the third entry. It is 'Strawberry'.
• Find 'Strawberry' in the first column.
• Put a '|' mark on the second column 'in line' with Strawberry. This is shown in the fig.1.4.

Fig.1.4 Tally mark for the third entry in the  data list
In this way all the entries in the raw data should be marked in the 'Tally marks' column. For doing this, we must learn to do a 'special type of marking' when the count for an item reaches 5. It can be explained based on an example: After marking the 14 th item, the table will look like in fig.1.5.
Fig.1.5 Tally marks after the 14 th item
Now take the 15 th item. It is Strawberry. Strawberry has been entered four times so far. The next one will be the fifth. For a fifth entry, we do not put  a '|' mark. Instead, we put a diagonal '\' mark. This diagonal must cross all the four '|' marks. This is shown in the fig.1.6 below:
Fig.1.6 Diagonal Tally mark 
Now we can proceed to enter the rest of the items. The completed table is shown in fig.1.7 below:
Fig.1.7 Completed Table

We can see that the last column shows the total number of students who prefer each flavour.

The counting of tally marks is made much easier by making the diagonal mark after four. Because each group with a diagonal will indicate a 'five'. So we do not have to count with in such groups.

Now let us see some features of the above table.
• The total number of times that a particular item (in our example, items are the 'flavours') occurs in the raw data is called the frequency of that item. For example, the frequency of chocolate is 6. 
• The frequency is the 'total number' of a particular item. This 'total number' is distributed randomly in the raw data. So it is not so easy to get the frequency of a particular item from the raw data.  
• But the table readily gives us the frequency of each item. So the table is called Frequency distribution table.

Now let us see another method to 'present' the information in the frequency distribution table. It is called the Bar graphIt is a pictorial representation of the 'results of the analysis'. The fig.1.8 shows the bar graph of our problem. 

Fig.1.8 Bar Graph
• Each item, is represented by a bar. 
• The 'height of the bar' of an item is equal to it's frequency.
• The widths of all the bars are equal
• The width of all the gaps between the bars are equal.

We can see that the Bar graph gives a better presentation of the results of the analysis. Strawberry is the most preferred flavour, with Vanilla following close behind. In the next section, we will see another example.

    CONTENTS       NEXT                                          


Copyright©2016 High school Maths lessons. blogspot.in - All Rights Reserved