Data Analysis Total Session Time: 120 minutes
Session 23: Data Analysis
Total Session Time: 120 minutes
Prerequisites
None
Learning Tasks
By the end of this session students are expected to be able to:
Describe data in terms of frequency distribution, percentages and proportion
Use figures to present data.
Explain the difference between mean, mode and median
Calculate the frequencies, percentages, proportion, ratios, rates means, medians, modes
for major variables.
Identify variables that are necessary for analysis of the collected data
Resources Needed
Flip charts, marker pens, and masking tape
Black/white board and chalk/whiteboard markers
Computer and LCD Projector
SESSION OVERVIEW
Activity/
Step Time Content
Method
1 05 minutes Presentation Introduction, Learning Tasks
Presentation Description of Data in Terms of Frequency
2 15 minutes
Brainstorming Distribution, Percentages and Proportion
10 minutes
3 Presentation Using Figures to Present Data.
20minutes Presentation
4 Group Difference Between Mean, Mode and Median
discussion
20 minutes
Calculation of the Frequencies, Percentages,
5 Presentation Proportion, Ratios, Rates Means, Medians, Modes for
Major Variables.
40 minutes
6 Presentation Identification of Variables that are Necessary for
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
228
Analysis of the Collected Data
7 05 minutes Presentation Key Points
8 05 minutes Presentation Evaluation
SESSION CONTENTS
STEP1: Presentation of Session Title and Learning Tasks (5 minutes)
READ or ASK students to read the learning tasks and clarify
ASK students if they have any questions before continuing
STEP 2: Description of Data in Terms of Frequency Distribution,
Percentages and Proportion (10 minutes)
Activity: Brainstorming (5 minutes)
Ask students to brainstorm on the following question
What is frequency distribution, percentages and proportion?
ALLOWS few students to respond
WRITE their responses on the flipchart or board
CLARIFY and SUMMARISE by using the content below
etc., that describe the data.
o Frequency counts
From the data master sheets, simple tables can be made with frequency counts for
each variable.
A frequency count is an enumeration of how often a certain measurement or a
certain answer to a specific question occurs.
For example
Smokers 51
Non-smokers 93
Total 144
If numbers are large enough it is better to calculate the frequency distribution in
percentages (relative frequencies):
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
229
given. In other words, percentages standardize the data.
into categories. This process may include the following steps:
o Inspect all the figures: What is their range? (The range is the difference between the
largest and the smallest measurement.)
o Divide the range into three to five categories. You can either aim at having a
clinic distance) or you can define the categories in such a way that they are each equal
o Construct a table indicating how data are grouped and count the number of
observations in each group.
STEP 3: Using Figures to Present Data (10 minutes)
o Figures make the descriptive data more readable when you have many tables
o Numerical data
Histograms
Line graphs
Scatter diagrams
maps
o Categorical data
Bar charts
Pie charts
o Is simplest and most effective means of illustrating qualitative data
o Bars can either be horizontal or vertical
o Eg.57 Adolescents from Kaloleni streets in Arusha were asked the following question:
How often have you used cannabis for the past one year? This was closed question
with the following possible answers
o Frequently (more than 5 times), Occasionally ( 3 to 5 times), rarely (1 to 2
times)and never
Categories Number Percentage
Frequently 7 12.2
Occasionally 9 15.8
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
230
Rarely 10 17.5
Never 31 54.4
Total 57 100
Pie charts
o Provides quick view of data presented in different form.
o Used in qualitative number with few categories to avoid congestion
Histograms
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
231
o Numerical data are often presented in histograms
o Which are similar to bar charts important difference is that in histogram ‗the bars‘ are
connected (as long as there is no gap between the data where as in bar charts are not
connected as the different categories are distinct entitles)
Line graphs
o Particularly useful for numerical data if you want to show Trend over time
o It is easy to show two or more distribution in one graph as long as difference between
lines are easy to distinguish e.g. age distribution between males and females
STEP 4: The Difference Between Mean, Mode and Median (30 minutes)
Activity: Small Group Discussion (20 minutes)
DIVIDE students into small manageable groups
ASK students to discuss on the following question
ALLOW students to discuss for 10 minutes
ALLOW few groups to present and the rest to add points not mentioned
CLARIFY and SUMMARIZE by using the contents below
o The mean of a data set is also known as the average value. It is calculated by dividing
the sum of all values in a data set by the number of values.
o So, in a data set of 1, 2, 2, 3, 4, 5, we would calculate the mean by adding the values
(1+2+2+3+4+5) and dividing by the total number of values (6). Our mean then is
17/5, which equals 3.4
o The mode is the most common observation of a data set, or the value in the data set
that occurs most frequently.
o The example of the mode in 1, 2, 2, 3, 4, 5, is 2
o The mode is an appropriate measure to use with categorical data
o The median of a data set is the value that is at the middle of a data set arranged from
smallest to largest.
o In the data set 1, 2, 3, 4, 5, the median is 3.
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
232
o In a data set with an even number of observations, the median is calculated by
dividing the sum of the two middle values by two. So in: 1, 2, 2, 3, 4, 5, the median is
(2+3)/2, which equals 2.5.
o The median is appropriate to use with ordinal variables, and with interval variables
with a skewed distribution
STEP 5: Calculation of the Frequencies, Percentages, Proportion, Ratios,
Rates, Means, Medians, Modes for Major Variables (35 minutes)
Frequency distribution
o Frequency distribution is description of data presented in tabular form.
o Gives frequency in each value appears in data
o Count number of response in category
o E.g. Frequency of categorical nominal data
Distribution of course students according to sex
Sex course students Number of course students
Male 34
Female 27
Total 61
Frequency distribution of numerical data
o Frequency distribution of numerical data is similar to that of categorical data except
here data have to be grouped in categories
school (45,43,45,47,54,53,62,38,54,34,45,53,56,42,51, 62,61)
o Procedures
Select group for grouping these data (selected Groups (31-40; 41-50; 51-60; 61-
70)
Count number of measurement (wt. of nurses) in each group:
31-40 II (34, 38)
41-50 IIIIII (42, 43, 45, 45, 45, 47)
51-60 IIIIII (51, 53, 53, 54, 54, 56)
61-70 III (61, 62, 62)
Add up and check totals.
o Groups must not overlap, to avoid confusion
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
233
o There must be continuity from one group to the other (no gaps)
o Groups must range from the lowest to the highest measurements
o Groups should normally be of equal width
o is number of units in the sample with a certain characteristics divided by total of units
in the sample multiplied by 100
o May also be called Relative frequencies
o Standardizes the data and make it easier to compare with similar data obtained in
another sample of different size or origin
Weight(Kg) Number of Pharmacy Relative frequency
students (percent)
31-40 12 26.7
41-50 17 37.8
51-60 11 24.4
61-70 5 11.1
Total 45 100
o Don‘t include missing numbers and not applicable in calculation.
o Don‘t know' is special category that should not counted as missing data.
o Should not be used if total is less than 30, as one unit makes a big difference in terms
of percentage
o Sometimes relative frequency are expressed in proportion instead of percentages
o Definition: proportion is numerical expression that compares one party of study units
to the whole.
o Can be expressed in fraction or decimal
Numerator is part of denominator
o It is numerical expression that indicates relationship in quantity or amount or size
between two or more parts.
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
234
Numerator is not part of denominator
Example: Ratio of module studied in first semester to module to be studied in
second semester
o Is quantity or amount or degree of event or disease measured over specified period of
time.
STEP 6: Identification of Variables that are Necessary for Analysis of the
Collected Data (40 minutes)
(e.g.
sex).
o Variables are something that varies or logical groupings of attributes.
variable.
o Variables and attributes are the derived from the concepts and they are part
of the
operational definition for measurement.
Figure…: Variables and Attributes
Variables Attributes
Age Young
Gender Female
Occupation Lecturer
Race Chinese
Social Class Low
Economic class High
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
235
cause
or at least to influence the problem. Refers to cause, it is what you (or nature)
manipulate.
exposure-condition association.
o To be a confounder
A risk factor must be associated with the risk factor under study.
It must also be a risk factor for the condition/problem being investigated.
A potential confounder is any factor that is believed to have a real effect on
the risk of the problem under study e.g. smoking, age, socioeconomic status and
education level
Types of Variables and scales
to carefully define the problem and each of the factors identified when analyzing
the problem
variable and independent variables (factors influencing the outcome or problem)
Numerical Variables
o When the values of the variables are expressed in numbers
Example of a variable in the form of numbers is person’s age. The
variable ‘age’ can take on different values since a person can be 20
years old, 35 years old and so on.
Weigh (expressed in kilograms or in pounds)
Homes-clinic distance (expressed in kilometers or in minutes
walking distance);
Monthly income (expressed in dollars, Shillings)
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
236
o Some variables may be expressed in categories. For example, the variable
sex has two
distinct categories, male and female.
Figure: Example of categorical variable
Variable Categories
Colour Red, blue, green, yellow
Main type of staple food Maize, rice, millet, cassava
Types of drugs Antibiotics, anti-inflammatory
o Continuous: With this type of data, one can develop more and more accurate
measurements depending on the instrument used. For example, height in
centimeters
(2.546 cm or 2.543216 cm) and Temperature in degrees Celsius (37.2 0 c or
37.190).
o Discrete: These are variables in which numbers can only take full values, e.g.
number of visits to a clinic (0, 1, 2, 3, and 4). Number of sexual partners (0, 1,
2, 3, 4,
and 5).
o Ordinal variables: These are grouped variables that are ordered or ranked in
increasing or decreasing order. For example:
High income (above $3000 per month), middle income ($1000-$3000 per
month), and low income (less than 1000 USD per months).
Disability: no disability, partial disability, serious or total disability.
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
237
Seriousness of a disease: severe, moderate, mild.
Agreement with a statement: fully agree, partially agree, and fully
disagree.
o Nominal variables: The groups in these variables do not have an order or
ranking. For example
Sex: Male, female.
Main food crops: Maize, millet, rice.
Religion: Christian, Muslim, Hindu, Buddhism.
o Most of what we call factors are variables which have negative values.
o Contributing factors in negative phrase are such as lack of knowledge.
o Example:
Factor Variables
Long waiting time Waiting time
Absence of drugs Availability of drugs
Lack of supervision Frequency of supervision
o Quantitative variables:
Variables which have definitive quantitative values (age in years- 24 year
or 4 years or weight or height) and can be manipulated according to the
rules of mathematics.
Ordinal variables: Variables which do not have numerical values but can
be graded (e.g. level of education as primary school, advanced diploma
and bachelor’s degree or quality of services as poor, average, good).
Nominal qualitative variables or attributes (do not belong to either of the
above) examples are marital status-married, single, divorced or gender as
male or female.
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
238
For a selected research problem, you may find that you are interested on
many variables, but to make a research manageable, pick few of them and
leave others.
for causal explanations, it is important to make a distinction between dependent
and independent variables.
of the
problem and the objectives of the study.
variable is the dependent and which the independent ones are.
to speak of associations between variables, unless a causal relationship can be
proven.
of the
problem.
o A confounding variable may either strengthen or weaken the apparent
relationship
between the problem and a possible cause.
Figure : Cause and Effect/outcome
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
239
must be
considered, either at planning stage or while doing data analysis.
For example:
o A relationship is shown between compliance with antimalaria treatment and
severe malaria in under-five children. However, mother’s education may be
related to compliance with the treatment and severe malaria.
o Mother’s education is a potential confounding variable. In order to give a true
picture
of the relationship between compliance with the treatment and severe malaria
in under five children, the influence of mother‘s education should be controlled.
o This could either be done in the research design, e.g., by selecting only
mothers with a
specific level of education, or it can be taken into account in the analysis of
the
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
240
findings. Then the relation between compliance with antimalaria and severe
malaria would be analyzed separately for mothers with different levels of
education.
Figure ; Compliance with antimalaria treatment and Severe malaria
o Related to a number of independent variables, so they influence the problem
indirectly
o In almost every study, background variables appear, such as Age, sex,
educational level, socio-economic status, marital status and religion.
Only background variables important to the study should be measured.
Background variables are notorious ‘confounders.
STEP 7: Key Points (5 minutes)
, etc., that describe the data
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
241
STEP 8: Evaluation (5 minutes)
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
242
References
Hardon A, Boonmongkon P, and Streefland P. et al (2001). Applied Health research,
Anthropology of health and health care, (3rd Ed) Amsterdam, The Netherlands: Het
Spinhuis Publishers
Beaglehole R, Bonita R and Kjellstrom (1993) Basic epidemiology: Geneva, Switzerland:
World Health Organization,
Kothari C.R (1985). Research Methodology – Methods and techniques, (2nd ed); New Delhi,
India; Wiley Eastern Limited
Stewart A (2001). Basic Statistics and epidemiology, A practical guide,; London, United
Kingdom: Radcliffe Medical Press,
Varkevisser, C. M, Pathmanathan, I and Brownlee, A (1991). Designing and Conducting
Health Systems Research Projects, Vol. 2 Part I: Ottawa, Canada: IDRC
Polit, D. F and Beck, C. T (2004). Nursing Research – Principles and Methods, (7th Ed):
Philadelphi, USA: Lippincott Williams & Wilkins,
___________________________________________________________________________
PST 06210 Operational Research NTA Level 6 Semester 2 Facilitator Guide
243
Unataka kutumiwa notes hizi kupitia WhatsApp?Kwa notes zilizopangiliwa vizuri kwa kusoma offline au PDF, bonyeza kitufe hapa chini. Ujumbe wenye Level, Semester, Module na Topic utaandaliwa moja kwa moja.TUMIWA NOTES WHATSAPP