Geography Form Three Notes – Application of Statistics

Geography Form Three Notes – Application of Statistics

These Form Three Geography notes cover physical geography, soil, surveying, map and photograph interpretation, and statistics in a structured mobile-friendly format.

Topic: Application of Statistics

Concept of Statistics

Statistics Explain the concept of statistics Statistics is the study of collection, analysis, interpretation, presentation, and organization of data, Data refers to crude or uninterrupted information.

In applying statisties to, for example, a scientific, industrial, or societal problem, itis necessary to begin with a population or process to be studied. Populations can be diverse topics such as "all persons living in a country" or “every household in a village”. It deals with all aspeets of data including the planning of data collection in terms of the design of surveys and experiments

Types of statistics

Two main statistical methodologies are used in data analysis, namely, descriptive statistics and inferential statistics a. Descriptive statistics summarizes data from a large sample using indexes such as the mean or standard deviation. Descriptive statisties are distinguished from inferential statisties (or inductive statistics), in that descriptive statistics aim to surnmarize a sample, rather than use the data to lear about the population thatthe sample of data is thought to represent

b. Inferential statistics draws conclusions from data that are subject to random variation (e.g, observational errors, sampling variation). Descriptive statistics are most often concerned with two sets of properties of a distribution (sample or population): (i) Central tendeney (or location) — this seeks to characterize the distribution’s central or typical value, (it) Dispersion (or variability) — this characterizes the extent to which members of the distribution depart from its

centre and each other. ‘Types of Statistical Data Differentiate types of statistical data When working with statistics, it's important to recognize the different types of data. Data are the actual pieces of information that you collect through your study. For example, if you ask five of your fiends how many pets they own, they might wive you the following data: 0, 2, 1, 4, 18 (The ih friend might count each of her aquarium fish as @ separate pet), Not all data are

numbers; let’s say you also record the gender of each of your friends, getting the following data: male, male, female, male, female Most data fall into one of two groups: numerical or categorical

  • Numerical data, These data have meaning as a measurement, sueh as a person's height

weight, 1Q, or blood pressure; or they're # count, such as the number of stock shares a person ‘owns, how many teeth a dog has, or how many pages you can read of your favourite book efore you fall asleep. Statisticians also call numerical data quantitative data. Numerical data can be Further broken into two types disetete and continuous. Discrere data represent items that ean be ‘counted; they take on possible values that ean be listed out, The list of possible values may be

fixed (also called finite), or it may go from 0, 1, 2, on to infinity (making it countably infinite). Continuous data represemt measurements; their possible values eannot be counted and ‘can only be described using intervals on the real number line, For example, the exact amount of as purchased at the filling station for cars with 20-gallon tanks would be continuous data from 0 gallons to 20 gallons, represented by the interval (0, 20], inclusive. You might pump 8.40

gallons, of 8.41, or 8.414863 gallons, or any possible number from 0 to 20. In this way, continuous data can be thought of as being uncountably infinite. For ease of recordkeeping statisticians usually pick some point in the number to round off

  • Categorieal data: Categorical data represent characteristics such as a person's gender

marital status, hometown, or the types of movies they like, Categorical data can take on hhumerical values (such as “I” indicating male and “2” indicating female), but those numbers don't have mathematical meaning. and you couldn’t add them together Other names for categorical data are qualitative data, ot Yes/No data

  • Ordinal data’ These data mixes numerical and categorical data. The data fall into

categories, but the mumbers placed on the categories have meaning. For example, rating a restaurant on a scale from 0 (lowest) to 4 (highest) stars gives ordinal data. Ordinal data are often treated as categorical, where the groups are ordered when graphs and charts are made, However, unlike categorical daia, the numbers do have mathematical meaning. For example, if you survey 100 people and ask them to rate a restaurant on a scale from 0 to 4, taking the average of the 100

responses will have meaning This would not be the ease with categorical data Statistical data can be expressed in different levels or scales of measurement. These are: a Nominal seale: This type of seale has qualitative property such that one my decide to express the data as “excellent, "good’, “fai’ or ‘poor’ and maybe use grades. ¢.g. A.B. C, D and so on. Nominal scale may also include numerical values. For example one may decide to let 1,2,

3 and 4 stand for ‘excellent’, ‘good’, ‘fair’ or ‘poor’ or vice versa. >. Ordinal scale: This scale involves ranking, so it is also qualitative in nature. The data involves rank orders or positions among events or objects. These statisties attempt to provide ‘quality or position. For example, if Chacha scored 5% in Geography Test while Tibaijuka scored 95%, then we can sey that the former ranked number 19 while the latter ranked mumber 1 out of

20 students, Sometimes, values such as "4 of the class scored below S0% in Geography may be included in the ranking,

Seale Properties Examples

Religion: I-eatholie; 2-protestant; — 3~Jewish Nominal Indicates a difference, without any implied ordering 4~Muslim, Sother Indicates a difference, and the direction of the Attitude on a subject'I~stronaly disagree, 2-disoetee: Ordinal differenci(e more of Less than) 3donit care /dont know; 4-agree; Sstronely agree Interval Indicates a diffrence, with diectionaity and amount of Temperature in CelausOccupational Prestige (12-96)

diterence in equal itervals Indicates a difference, the cirection of the diference, the amount of the difference in equal intervals, an absolute Ratio 20 Tempetoture in KelvinlncomeYeas of schooling © Taterval scale: This type of seale employs iruly quantitative values and allows the use of mathematical operations such as adding, subtracting, multiplying and dividing. At no time is zero present in this scale, For example, the range of temperature in which rice grows well is 25°C and

48°C; most livestock keepers get between 10 and 15 litres of milk per cow per day. Ratio seale This is & type of scale that is used 10 make comparisons between values or quantities. For example, Ms Iku harvested $0 sacks of maize which is twice Mr Aritamba ‘obtained fiom the same aereage because the former applied fertilizer and good farming practices while the latter did not Variables A variable is anything or characteristic that data may have, or an attribute which changes in

value under given conditions. Variables include population size, age, sex, altitude, temperature and time. There two broad types of variables, namely, independent and dependent variables.

a Am independent variable is @ variable factor which influences the changes of other variables or outcomes, The independent variable is also known as manipulated variable, This is the factor manipulated (controlled) by the researcher, and it produces one or more results known asdependeni variables. There may be more than several dependent variables, because manipulating the independent variables can influence many different things. For example, am

experiment to test the effects of a certain fertilizer, upon plant growth, could measure height, umber of fiuits, and the average weight of the fruits produced. All these are varied analyzed

factors, arising from the manipulation of one independent variable, the amount of fertilizer

b. A dependent variable is an outcome or result that has been influenced by other variables. A dependent variable does not influence or change other variables, The dependent variable responds to independent variable. It is called dependent because it “depends” on the independent vuriabe. In any research, you cannot have a dependent variable without an independent variable Any alteration in the independent variable will change the dependent variable For example, you

phosphorus fertilizer on maize growth. To conduct this experiment, you grow maize in similar (independent variable). Then you measure the height of maize plants (dependent valuable) alter the height you will obtain wil obviously depend on the amount (concentration) ofthe fertilizer applied. And, inthis case, you will obviously get different heights depending on the quantity of fertilizer applied

Graphical Data

Present data graphically After data have been collected, the next step is to present the data in different ways and forms. Some of the forms in which the data may be presented include charts, graph, list, diagrams, Line (linear) graphs Line graphs have unique properties that distinguish them from other graphs. The properties of line graphs ar a fllows: a. The graphs are drawn by plotting a dependent variable agninst an independent variable

and poins are joined by a ine. General procedure for drawing line graphs &. Get the required data for plotting the graph b. Identify the independent and dependent variable, Statistically the independent variables variable available decide on the horizontal spacing of the graph according to graph space available © Draw and divide the vertical and horizontal axes depending on the respective scales £ Plot and join the points to get the graph.

2 Write the ttle of the graph you have drawn. h. Indicate the seale of the graph, i. Show the key for the graph if need be Line graphs can be sub-divided int: a Simple line graphs b Group (comparatives) line graphs © Compound line graphs

  • Divergent line graphs

Simple tine graph Presenting the statistical data by a simple line graph is the most common and popular method. ‘The simple fine graphs are easy to construct and interpret, They have many uses whieh include showing temperature, farm outputs, population, and mineral production, among others Construction procedure: ‘The graph can be drawn after getting the required data Consider the following table which shows the average monthly temperature recorded in a certain weather station

Average monthly temperature for station X Month Jan Feb Mar Apt May Jun Jul Aug Sept Oct Ney Dee

TempeC) 2324 HR HHH HHS

The following procedures may be used 1 Mentify the variables. The dependent variable is temperature and the independent variable is months. 2 Determine a vertical scale. Assume that the graph space available is 6 em vertically

Vertical scale = maximum value of the divided by the graph space available g.30°C/6 em =

5°C per centimetre. Therefore, in the vertical axis (x-axis), em will represent 5°C

  • Determine the horizental scale (y-axis) depending on the available space, Let, for

instance, em represent one month. 4, Draw both axes and label them: y-axis for temperature and x-axis for months,

  • Plot the points and join them by a smooth line to make a curve.
  • Insert the title and seale

The following is a simple line graph showing monthly temperature for station X 3s ae _ We i Bos sof M AM 4 4 8 5 0 WN 8 Average monthly temperature for Station X

Souree: Hypothetical data

Seale © Vertical 1 em3°C *— Horizontal— 1 em-1 month

Advantages of simple line graphs

1 Theyare easy to draw, read and interpret

  • They show specific values of data, so if you are given one variable the other can easily be

determined.

  • They show pattems in daia clearly, meaning that they visibly show how one variable is

affected by the other as it increases of decreases. 4, They enable the viewer to make predictions about the results of data. So they allow for determination of intermediate or continuing values,

  • Itis easy to read the exact values against plotted points on straight line graphs.
  • A broken scale can be used when the value starts at a large number.

Disadvantages of simple line graphs

1 They ean only be used to show the data of one item over time 2 One can change the data ofa line graph by not using consistent scales on the axis.

  • They can give a wrong impression on the continuity of data even when there are periods

when data is not available

  • ‘They do not give a clear visual impression of the actual quantities.

Group (comparative) line graph A group line graph is also known by the following terms:

  • Comparative line graph
  • Composite line graph
  • Multiple tine graph
  • Polygraph

A group Tine graph involves drawing more than one line on the same statistical graph. It shows, the relationship between sets of similar statistics for two or more items, Usefulness of a group line graph

  • Comparing different values or tends in two or more data variables.
  • Examining the possibility ofa relationship existing between the distributions of a number

of variables over time

  • Comparing the distribution of the same variable at different places

Construction: ‘The method of drawing a group line graph is the same as fora simple line graph. Therefore, to draw each single line in a group line graph, follow similar steps used for construction of the simple fine graph The following things should be considered before drawing the graph: 1, The lines drawn should not be uniform in colour, thickness, general appearance, ete (See the graph below in which each line has a different colour),

2 The number of lines that a graph can accommodate should not exceed 5, meaning that not more than 5 items should be compared in a single graph The following table shows banana production (in tonnes) by three villages in Ingwe Division, Tarime district. These dats have been used to plot the group (comparative) line graph as shown below Banana production by three villages

Village/Year Geisangora Liryo Bungurere Nyansineha

2000) 10 15 25 25 2001 20 10 Is 20 2002 3 0 7 Is

Source: Hypothetical data

s § —eisaneora ‘fas ine 3 «+ ——— ——hyansincha é 2000 200 2002 Year Maize production by three villages between 2000 and 2002

Advantages of group line graph

1, The quantity of each component is shown clearly by different line shadinas.

  • ‘Time and space are saved since all the line graphs are drawn at ago as a group.

Disadvantages of group Tine graph

  • The lines can be overcrowded and hence become difficult to read and interpret if many

data are involved.

  • Itdoes not give a clear visual impression of actual quamtties.

‘Compound line graph A compound line graph is used to analyse the total and the individual inputs of the specific commodities or economic sectors. The graph involves drawing two or more lines, each line corresponding to one item in a different year or region. The items are differentiated fom each other or one another by shading differently Construction: The table below is used for construction of the graph. The table contains hypothetical figures for

mineral exports between 2010 and 2012

Year/Mineeal Diamond Gold Fanzanite

2n10 10,000 16,000 20,000 2011 20,000 25,000 32,000 2012 25,000 35.000 0,000 Procedure:

  • Simplify the data to make the presentation work easy by dividing each value by 1000.

Year (Mineral Diamond Gola Tanzanite

2m10 0 6 0 2011 20 25 2 212 a xs 40

  • Add the values for each year to get the cumulative export: 2010 = 10#15+20 = 46, 2011
= 20425+32 = 77; 20112 = 25+35+40 ~100; These values will be used to determine the

uppermost height of the graph. They will also help estimate the seale to be used. In case of the above data, the highest value is 100. So if we want to use the scale of Tem to tonne (1000 tonnes in reality) the uppermost height of our graph will be 100 em (see the graph drawn

  • Plot the values for mineral exports against years on 9 graph, Usually the line graph for

data with the highest values is drawn first Thus, frst draw the Line graph for tanzanite since i has the highest values, followed by that of gold and finally diamond

  • Draw the second Tine graph above the first one to show the next component. To get the

values for plotting the second line graph, add the values of the first item (in this ease, tanzanite)

to that of the second item (gold) for each year, thus: 2010 = 20+6 =36, 2011 = 32425 =57, 2012
= 40-35-75
  • Draw the line graph forthe lat item (diamond) above that ofthe second item. To get the

values for ploting this graph, add the values for the second item to those of the last item, thus

2010 = 36+10 =46; 2011 = 57420 ~67; 2012 = 75+25 =100
  • Shade the component parts between th line graphs using different shadings as shown,
  • Label the axes, show the key and indicate the seale used to construct the graph.

100 + _

. Be Key

g oe £5 Disewond a °Ra _ Ls z MD varenite 2010 zou 2012 Year

Advantages of compound line graph

1, Total values are shown clearly and easily, 2 Iteives good visual impression.

  • Combining all graphs in one saves time and space

Disadvantages of compound line graph

1, Graph construction is difficult and time-consuming

  • involves a lot of ealeulations which are difficult and time-consuming,
  • It is difficult to read and interpret the value for any one commodity for any particular

year Divergent line graph AA divergent line graph is a line graph which shows how variables deviate from the mean. The mean is represented by zero axis drawn horizontally across the graph paper.

Year Yield (onnes)

212 1000 2013, 1500 ror 500 2m 3000 Construction

  • Sum up the values ofall items or commodities. 1000 + 1500 ~ 500 + 3000 = 6000
  • Calculate the arithmetic mean (average) of the values. 6000/4 = 1500 Thus the arithmetic
mean () = 1500
  • Calculate the deviation from the mean of each value as shown in the table below

Deviation from the mean value Year x x 2012 1000 00 2o13 100 ° ous suo 1000 2ols 000 +1500 ttle and seal ofthe graph i. i _— z 2012 2013 ove 2o1s Year 1 It clearly shows how items fluctuate from the mean, 2 it compares the valucs of the items and hence facilitates a sound conclusion,

  • Itshows both the positive (profit) and negative (loss) phenomena

Disadvantages of divergent line graph

  • Itinvolves many calculations and hence time-consuming

2, It might be difficult to interpret if one lacks statistical skills.

  • tis applicable for only one item per graph.

Bar graphs A bar graph is also called bar chart or columnar graph, This method is used to present data which are not continuous. ‘This means that in a bar graph there is no relationship between or among dota Bar graphs emphasize individual amounts and their relative variations. When drawing such sraphs, bar width in a graph is kept constant while bar lengths change in size as per the amount of the independent variable in question

Though the bars ean also be drawn horizontally, they are usually drawn vertically. The bars should be separated from one another by a space

Types of bar graphs.

a Simple bar graphs >. Group or comparative bar graphs © Compound bar graphs Divergent bar graphs Simple bar graph A simple bar graph is drawn to show a single item per bar It mainly represents simple data. Consider the data in the table below which shows the value of sisal exported by Tanzania between 1900 and 1993 Year Sisal expott (Tsk 000) 1990 106126 1991 107430 1992 142601 1993 61180 1994 20s Construction: 1 Choose the appropriate scale. However, note that the table below is not drawn to scale ~

it was drawn using the computer. All hand-drawn graphs must indicate the scale used. For, ‘example, in our graph below, we might have chosen 1 em to represent 10,000 tones, in which case we could obtain the values 5, 10, 15, 20 and 25 that we could have used to plot the graph 2 Draw the axes and insert the bars, Note that all the bars must have the same width and spacing

  • Shade the bars uniformly by using shade, lines, crosses, dots, ete

4, Insert vertical and horizontal seales and the ttle, 250000 138000 —______________gas 100000 E soon F 190 1901 19 1992 1998 Year Tanzania sisal export Seale: 1 cm to 50,000 tonnes

Advantages of a simple bar graph

1 Itis simple to construct, read and interpret 2 Ithas a good visual impression

  • __Itcan be used to compare how the amount of an item varies from time to time.

Disadvantages of a simple bar graph

1, Iti limited to onty one item or commodity and hence not suitable for massive data. 2, Not suitable for continuous data such as temperature. Group (comparative) bar graph A comparative bar graph consists of several bars drawn side by side on the same chart for the purpose of comparison. The technique involves grouping of bars in a chart The graph can be used to show how production of certain commodities varies each year.

Construction: The procedure for construction of the comparative bar graph is similar to that of drawing the simple bar graph except that the simple bar graph contains a single bar while the comparative bar raph comprises of multiple bars.

Consider the data in the table below, showing agricultural production in mettic tonnes YeariCommosity 1986 1987 1988 Sorghum 1200, 5000 804 Tea 9000 7000 0c Tobacco 3000, 000 400 ‘The graph for the data is as shown below 5 Tc

FA Tobacco

E 4 i gz e 1986 a997 set Group (comparative) bar graph showing crop yields in ‘000 kg (1986-1988) Advaniages ofa group bar graph

  • The total values are expressed well for illustration of points

2 Mis easy to construc, read and interpret

  • The importance of each component is shown clearly

Disadvantages of a group bar graph

1, is difficult to compare the totals of each item/component.

  • Trends such as fall and rise cannot be shown easily,

‘Compound (divided) bar graph This 1s a method of data presentation that involves construction of bars which are divided into segments 10 show both the individual and cumulative values of items, The lengih of each segment represents the contribution of an individual item in the total length while that of the whole bar represents the (otal (cumulative) value of the different items in each group.

Construction

  • Get the data needed for presentation. For example, consider the table below. which shows

the number of tourists who visited the named Tanzania National Parks from 1998 to 2002, Year /oark 1998 99 2000 2002 2003 Manyara 120,000 160.000 172,000 70.000 203,000 Serene 175,000 160,000 148,000 185,010 201,000 Tarangire 29,000 30.000 $4,100 79.000 102,000 Mikuni 100,000 110.000 113,000 150.000 183.400

  • Simplify the data (to make the presentation work easy) by dividing each value by 10,000.

Then add the values to get the total for each year. The simplified data are as shown in the table below

  • Determine the seale of the bar length based on the highest total value. In this case, the

highest total value is 68 (20+ 20 + 10+ 18) Recall the construction of the compound line graph! If we choose 1 em to represent tourist (10,000 tourists in reality), then the length of the tallest bar will be 68 em, Note that the maximum height of'a graph for each year equals the cumulative total values for cach year (1.2. 43, 46, 48, $9, 68),

  • Decide on the bar spacing, for example, 1 em apart
  • Draw the axes and label them
  • Start by drawing bars that represent the highest values,
  • The first sets of bars to be drawn are those that represent the highest values. On top of

these, the second highest segments are drawn, The last segments to be drawn are those with the lowest values in general

  • Tomake it easy to follow the rise and fall of individual values, soft line could be drawn

across bars to separate individual segments,

  • Colour or shade the segments to improve the appearance and simplify interpretation
  • Inset the scales, key and ttle

x atin Zio sTwontire = as ss sornget 3 ener g t ° * 2 1098 1999 2000 2001 2002 Year Compound (divided) bar graphs showing tourist visits in 0°000 (1998-2002)

Advantages of compound (divided) bar graph

1 Its easy to read and interpret as the totals are clearly shown

  • It gives a clear visual impression of the total values,
  • It clearly shows the rise and fall in the grand total values

Disadvantages of compound (divided) bar graph

  • The values of individual segments above the first set are difficult to establish because

they don’t start at zero, To get the correct values of the top segments, you have to add the figures, which is difficult for someone not well equipped with statistical skills 2 The graph is very difficult to construct and interpret

  • is not easy to represent a large slumber of components as this would involve very long

bars with many segments Divergent bar graph AA divergent bar graph is a graph which shows the fluctuation of individual items from the mean. Construction: 1 Calculate the arithmetic mean (average) of the items.

2, Subtract the mean from each item.

  • Draw the graph using the resulting values.
  • Insert the scale and ttle of the graph

The data below show the enrolment of Form One students at Mara Secondary School from 1980-1985. Study the table and present the data by a divergent bar graph. Year ‘Number of students 1980 100 981 150 982 1s oss 200 Loa ms Loss 300 Procedure:

  • Find the arithmetic mean

Ix Where: T mean e+ sum of individual items n= the total number of items x _ 1150 Therefore, LE ~ 1150 _ tgp a 6

  • Subtract the mean from each item:

ver Number of students x 1980 100 2 19st 150 2 182 a 17 1983 200 8 1984 ns 2 1985 200 108

  • Choose a suitable seale and construct the graph using the obtained values (X —).

Ee 1980 1981 1982 1983 1984 1986 Year A divergent bar graph showing student enrolment (1980-1985) Advaniages of divergent bar graph

  • Fluctuation in values, which helps to detect the problem in general terms, is shown,

2 Itis important for comparison of positives and negatives:

  • Profit (success) or loss (Failure) can easily be deduced

4, ‘They are simple to construct, read and interpret,

Disadvantages of divergent bar graph

  • Graph construction is time-consuming since it involves many steps.
  • The calculations involved may be difficult to someone who is poor at mathematies.
  • tis limited to analysis of only one variable.

Divided circles (pie charts) A divided circle is also known as pie chart, citcle chart or pie graph. The chart involves dividing the circle into “pie slices” to represent and show relative sizes of data. The size of each slice or segment is always proportional to the value it represents. Divided circles ean appear in two forms: a. Simple divided circles.

b. Proportional divided cireles. AA simple divided circle involves a single set of data whereas the proportional divided circle involves more than one set of data sueh thatthe circles will be proportional to the total quantity that each cirele represents Simple divided eircle Construction:

  • Obtain the data to work on. Study this hypothetical record showing enrolment of Form

(One students in selected Secondary Schools in Tarime District: A table showing student enrolment in selected schools in Tarime District Name of school Number of students Bungurere 80 Nyanungu " Magato 8 Tarime 6 Nyamongo % Total 456

  • Calculate the total number of students as shown in the table
  • Calculate the angle in a circle that would represent the number of students enrolled in

each school, For example, 85 out of 456 students enrolled in Nyansincha Secondary School will be represented in the circle by a segment with an angle of 85/456 *630 ~ 67 degrees This will sive the following results: Name of school Number of students Degrees Nyansineha 8s or Bungurere 0 or Nyanungu 78 or Mazoto 78 o Tame 6 sir Nyamonzo » 3° Total 456 360°

  • Draw acirele of reasonable size
  • Using a protractor, draw a radius from the 6 o'clock mark to the centre of the cirele
  • Starting with the largest seament representing a specific component, measure and draw

its angle from the centre of the citcle.

  • Do the same for other components in ascending order
  • Divide a circle into segments according to the sizes of the angles,
  • Shade the segments and write the ttle and key of the drawn graph.

Key satiyansiocha mbungurere fatiyonunew mMagoto marime atyamongo Student enrolment in selected Secondary Schools in Tarime District

Advantages of divided circles

1, tis easy to compare components as they are represented by angles. 2 Analysis and interpretation of data is easy.

  • It is easy to assess the proportion of individual components against the total
  • Construction of this graphical representation is relatively simple
  • itis easy to determine the value of each component since it is indicated on each segment

6, Visual impression of the individual components is clear and facilitates the understanding of the information in the data

Disadvantages of divided cireles

  • Its time-consuming because it involves 2 lot of calculations

2 The represented actual values remain hidden as the values shown on the faces of the segments may be in percentages of the chart is difficult.

  • When the valuzs of data set vary slightly, itis diffielt to visualize the proportional

The Importance of Statistics to the User Explain the mporiance of ialistcs tothe user

  • It enables the geographers to handle large sets of data and summarize them in a way that

can be easily understood. phenomena, ¢-¢ 10 compare the amount of rainfall and agriculture production or population 3, Sates tamstates data iano maiemalial ways ‘which wake the applicitia of ‘quantitative techniques possible.

chats, et.

  • Statistics give precise rather than generalized information. This offers a lot of satisfaction
  • Sutistcs i very useful for planning at local and national levels, For example, statistes

con census can be used to plan for social services. How Massive Data can be Summarised The massive data collected from the field have to be summarized so a8 to make it easy to read, Frequency distribution A fiequency distribution shows a summarized grouping of data divided into mutually exclusive classes and the number of occurrences in a class. It is a way of showing unorganized data e.2, to show results of am election, income of people for a certain region, sales of a product within @

certain period, student loan amounts, ete. Some of the graphs that can be used with frequency distributions are histograms, line charts, bar charts and pie charts, Frequency distributions are used for both qualitative and quantitative data Frequency distribution helps to determine how many times a certain score occurs in a sample, In statistics, a frequency distribution is a table that displays the frequency of various outcomes in @

sample, Each entry in the table contains the frequency or count of the occurrences of values within a particular group or interval, In this way, the table summarizes the distribution of values inthe sample ‘Consider the following table which shows family size of 20 families which were interviewed in a certain village:3, 2, 2,4, 3, 7,8, 1,3, 6,2,2,4,5, 6,4, 3,4 5, and2, ‘The data can be summarized in a frequency table thus

a Arrange the seores in a descending order from 8 to 1. Iti advised to arrange the scores in ascending order. b. Distribute each score in the sample to determine the number of times each score occurs (frequency) in the data sample.

Score Frequenc; a es [2s 1 I The frequency indicates how many times a score or event appears or occurs in sample However, in each case, itis certainly difficult to deal with individual scores separately. In such cases, a grouped frequency 1s used The steps for making a grouped frequency are as follows

  • Decide about the number of classes. Too many classes or too few classes might not reveal

the basic shape of the data set; also it will be difficult to interpret such a frequency distribution.

The maximum number of classes may be determined by formula: Number of classes = C= 1+
3 3login) or C = \n(appraximately) where n isthe total number of observations in the data
2 Calculate the range of the data (Range = Max — Min) by finding minimum and maximum

data value. Range will be used to determine the class interval or class width

  • Decide about the class interval denote by h and obtained by h = Range/Number of classes

4 Decide the individual class Timits and select a suitable starting point of the first class which is arbitrary, it may be less than or equal to the minimum value. Usually itis started before the minimum value in such a way that the midpoint (the average of lower and upper class limits of the first class) is properly placed,

  • Take an observation and mark a vertical bar ( !) fora class it belongs. A running tally is

epi till the last observation. However, it is not always necessary to show tallies inthe Frequency Distribution Table because the frequency column serves the same purpose

  • Find the frequencies, relative frequency, cumulative frequency ete, as required,

Frequency distribution table ‘Class interval Frequeney ‘Cumulative frequency 0-9 4 4 10-19 ° iB 29-29 8 a 30-39 3 4 40-49 4 28 50-59 7 33 50-69 s 0 0-79 4 4“ 80-89 2 46

Characteristics of the class interval

1 A.score appears only once. That means no score should belong to more than one class. 2 The size of the class interval should be the same. No score should fall in more than one class, Arrange the class intervals in order of ranks as shown in the frequency distribution table above.

  • The class intervals should zlways be continuous.

4, The range of class interval should be between 3 and 20. Thus, the intervals should not be below 3 and not above 20. From the summarized data in the table above, one can identity two concept: Apparent upper limit b. Apparent lower limit These limits (or boundaries) are seen in each class interval. The apparent lower limit opens the class interval while the apparent upper limit closes the class interval.The table above shows 80,

70, 50, 40, 30, 20 and 10 as apparent lower limits and 89, 79, 69, 59, 49, 39, 29, 19 and 9 as the apparent upper limits Apart from the two concepts above, the table has real Limits which are not visible. These are 0.5 below or above the apparent limits, From the above summarized data, other measures of statistics can be deduced. Such measures include the measures of central tendency. measures of dispersion (variability), measures of

relationship (correlation) and measures of relative position.

Simple Statistical Measures and Interpretation

Methods of Presenting Simple and Mixed Data

Describe methods of presenting simple and mixed data Measures of central tendency (averages) A measure of central tendeney is a single value that attempts to describe a set of daia by identifying the central position within that set of data. As such, measures of central tendency are sometimes called measures of central location. They are also classed as summary statistics The mean (often called the average) is most likely the measure of central tendency that you are most

familiar with, but there are others, such as the median and the mode. The mean, median and mode are all valid measures of central tendency, but under different conditions, some measures of central tendency become more appropriate to use than others, In the following sections, we will look at the mean, mode and median, and leam how to calculate them,

The Mean, Mode and Median

Caleulate the mean, mode and median Arithmetic mean The mean (or average) is the most popular and well known measure of central tendency. It can be used with both discrete and continuous data, although its use is most offen with continuous data. The mean is equal to the sum of all the values in the data set divided by the number of values in the data set. So, if;we have n values ina data set and they have values x1, x2, .., x0, the

sample mean, usually denoted by (pronounced x bar), i

= Oat Xn t 4 Xq)

gj. Ree n This formula is usually written in a slightly different manner using the Greek capital letter, Y., pronounced "sigma", which means "sum of." Ix r= Where: T =arithmetic mean x, = individua score n= number of occurences or events

Example 1

In English exam, students obtained the following percentage scores: 45, 42, 35, 86, 40, 56, 87, 40, 35, 74, 68 and $0 ‘The arithmetic mean of the score is: 45+42+35-+86+40+56+87+40+35+74+68+50 _ 688 _ 5, ‘The average score was 54.8%

Advantages of the mean

  • itis rigidly defined by a mathematical formula

2 His easy to understand and calculate ES It is based on all observations

  • It is determined im alll cases
  • is suitable for further mathematical treatment or manipulation

6 Compared to other averages, arithmetic mean is affected least by Actuation of sampling

Disadvantages

1, itis greatly affected by extreme values ofthe data 2, cannot be obtained ifa single observation (item) is missing

  • tis not appropriate in some distributions

Median ‘The median is the middle score for a set of data that has been arranged in order of magnitude Suppose we want to find the median from the data below 65, $5, 89, 56, 35, 14, 56, $5, 87, 45, 92 We first need to rearrange that data in order of magnitude (smallest first) 14, 35, 45, 55, 55, 56, 56, 65, 87, 89, 92 ‘Our median mark is the middle mark – in this case, 56 (highlighted in bold). Its the middle mark because there are 5 scores before it and 5 scores after it. This works fine when you have am odd

umber of scores, but what happens when you have an even number of scores? What if you had only 10 scores? Well, you simply have to take the middle two scores and average the result. So, if we look at the example below 65, 55, 89, 56, 35, 14, $6, $5, 87, 45 We again rearrange the data in order of magnitude (smallest first: 14, 38, 45, 55, 55, 56, 56, 65, 87, 89 Only now we have to take the Sth and 6th score in our data set and average them to get a median

of $5.5,

Advantages of the median

  • Its easy to calculate and understand

2, team also be calculated in qualitative data

  • his appropriate for skewed distribution

4 It is not affected by all extreme observations Hence, it is a better average than the arithmetic mean when extreme observations are present.

  • The values of a median can be obtained graphically.

Disadvantages

  • Itts not suitable for further mathematical treatment.

2 Itisnot rigidly defined

  • tis based on all values or observations

4 Compared to mean, median is more affected by Nuctuation of sampling

  • Im ease of ungrouped data, rearrangement of values in order of magnitude becomes

Mode The mode is the most frequent score in a data set. It represents the highest bar in « bar chart or histogram. You can, therefore, sometimes consider the mode as being the most popular option.

An example of a mode is presented below:

  • mode

ee ee Normally, the mode is used for categorical data where we wish to know the most common category, as illustrated below

Two Modes

¥ X Another problem with the mode is that it will not provide us with a very good measure of central tendency when the most common mark is far away from the rest of the data in the data set, as depicted in the diagram below Mode In the above diagram the mode has a value of 10. We can elearly see, however, that the mode is not representative of the data, which is mostly concentrated around the 2 to 3 value range. To use

the mode to describe the central tendency of this data set would be misleading

Advantages of the mode

1, itis simple to compute 2 It is easy to understand and calculate. In some cases it can be located merely by inspection. The value of the mode can be obtained graphically from the histogram,

  • It gives a rough idea of the differences of the data set.

4 tis the only average that can be used when the data is not numerical

Disadvantages

1, Itis not rigidly defined; hence it is unstable for large samples, 2 tis independent of sample size except under special circumstances,

  • tis mot based on all the values of the data.
  • Mode is not suitable for further mathematical restment
  • As compared to mean, mode is affected to grest extent by the fluctuation of sampling
  • ‘There may be more than one mode (as isthe ease inthe previous graph).
  • ‘There may be no mode at all iF none of the data are the same. $. It may not accurately

represent the data The Significance of Mean, Mode and Median Measures of central tendency are very useful in statistics. Their importance is because of the following reasons: 1 To find representative value: Measures of central tendeney or averages give us one value forthe distribution and this value represents the entice distribution. In this way averages 2 To condense data: Collected and clasifie figures are vast. To condense these figures

condensation have to find the representative values of these distributions. These representative values are found withthe help of measures ofthe central tendency

  • Helpful in further statistical analysis: Many techniques of statistical analysis like

Measures of Dispersion, Measures of Skewness, Measures of Correlation, and Index Numbers are bosed on measures of central tendency. That is why messures of central tendency are also called measures of the first order.

Interpretation of Data using Simple Statistical Measure Inerpret data sing simple statistical measures In the section about averages (mean, mode and median), we learned how to calculate the mean for a given set of data. The data we looked at were ungrouped and the total number of elements in the dataset was not tht arg. The metho isnot always a realistic approach especialy if you are dealing with grouped data Assumed mean (A), lke the name suggests, isa guess or an assumprion ofthe mean, I doesn't

need 10 be comect or even close to the actual ean and choice ofthe assumed ican is at your discretion except for were the question explicitly asks you to use a certain assumed mean value Assumed mean is used to calculate the actual mean as well as the variance and standard deviation, Meas of earl tendency cau be emolaed Tot grouped dt for example

  • Mean =a + Df

y Where vac mtn "num ofthe prodict of equency and deviation N= ton! fequoney

2 ate= t+ ( 4 i

aD Wher: =the lowe iit of the modal lass t= the exces of the modal frequency ovr the frequency of the net lower class f= the excess ofthe modal frequency over the frequency of the net higher cass the modal cls ntervl x, Where L= the lower boundary of he median class Noth otal mimic of freaveney

A= teil mur of tus labs below the odin ina

c= the total umber fiers within the modian clas 1Skechsinerval Calculation of measures of central tendency for grouped data Study the frequency distribution table below Midpom() f _ __V= [ o-4 [ os-as [2 J 2 [ 5-9 [ 45-95 [ 7 6 10-14 9.5— 14.5 12 lo 15-19 14.5- 19.5 WZ 8 me 20-24 19.5—24.5 22 4 ie ee es Calculation fom the table:

Mem )=a+ D4

Whee Nests

df = 30

No

Mean = 12+ 2
Mean =13

2 . (4)

  • Mode = L+ lizal

where r95

i= 10-6=4

fa 10-be2

Assumed mean (A) = 12
Note: we find the class interval by using the class limits as follows: i= upper class limit — lower

Get Well-Formatted PDF Notes

For an easier offline copy with the original document formatting, click below to request the complete PDF notes.

GET WELL-FORMATTED PDF NOTES

banner
Scroll to Top