The Stages of Data Collection
Chapter Eighty-Three
Syllabus topic 3.5, "Stages of data collection- conceptualizing problem, laying down hypothesis, defining the variables, choosing the tools of data collection, phase of data collection, data analysis."
Pages 370 to 374 of 451
In one line
A piece of research proceeds in six steps: work out the question, state what you expect, define what you will measure, choose the instrument, collect the material, and analyse it.
In the wording a student can write in an exam: the stages of data collection are conceptualising the problem, in which a vague area of interest is turned into a precise and answerable question; laying down the hypothesis, a tentative proposition to be tested; defining the variables, so that what is to be measured is stated unambiguously and operationally; choosing the tools of data collection appropriate to the question, the population and the resources; the phase of data collection itself; and data analysis, in which the material is processed, interpreted and related back to the hypothesis.
Stage 1: Conceptualising the problem
What it involves. Moving from a topic to a question. "Legal aid" is a topic and cannot be researched. "What proportion of eligible undertrial prisoners in this state's district prisons received legal representation at their first production before a magistrate" is a question, and everything that follows is possible only because it is one.
The steps within it: selecting the broad area; reviewing what is already known, so that the study is not a rediscovery; narrowing to a specific question; deciding the scope, which is the population, the place and the period; and stating the objectives explicitly.
The tests of a good research problem: it is clear and unambiguous; it is researchable, meaning that evidence capable of answering it can be obtained; it is significant, adding something worth knowing; it is feasible with the time, money and access available; and it is ethical.
The commonest failure in student research is at this stage, and it cannot be repaired later: a study whose question is vague produces data that answer nothing, however well collected.
Stage 2: Laying down the hypothesis
What a hypothesis is. A tentative proposition about the relationship between two or more variables, stated in a form that evidence can support or contradict. It is a proposed answer to the research question, held provisionally so that it can be tested.
The characteristics of a good hypothesis: it is clear and conceptually precise; it is empirically testable, so that some observation could count against it; it states a relationship between variables; it is specific rather than sweeping; it is consistent with known facts; and it is simple.
Its forms. The null hypothesis states that there is no relationship, and is the form actually tested statistically, the researcher seeking to reject it. The alternative hypothesis states that there is one. A directional hypothesis states which way the relationship runs; a non-directional one does not.
The Stages of Data Collection
Where hypotheses come from: theory, earlier research, observation, exploratory work, and analogy.
Not every study has one. Exploratory and descriptive studies, from [Types of Methodology], often proceed without hypotheses, and a study that manufactures one for form's sake has misunderstood the stage. Say so if the design does not require it.
Stage 3: Defining the variables
What a variable is. Any characteristic that can take different values across the units studied: age, income, education, caste, the number of adjournments, whether representation was provided.
The kinds. The independent variable is the presumed cause, the one manipulated or treated as prior; the dependent variable is the presumed effect, the one measured; an intervening variable stands between them in the causal chain; and a control or extraneous variable is one held constant or accounted for so that it cannot confuse the relationship.
Operational definition is the heart of this stage and the thing most often skipped. It means stating a variable in terms of the operations by which it will be measured, so that two researchers would classify the same case identically.
The illustration is worth giving. "Delay" cannot be measured. Operationally defined, it becomes: the number of days between the date of institution and the date of final disposal, counted from the court's own register, excluding matters transferred. Now it can be counted, checked and compared, and another researcher can repeat it.
Without operational definitions the data are not comparable, because each investigator applies their own understanding, and the study measures the investigators as much as the world.
Stage 4: Choosing the tools of data collection
What decides the choice, and this is the examinable part: the nature of the question; the population, above all its literacy, dispersion and accessibility; the kind of data required, whether behaviour, opinion, record or measurement; the resources of time, money and trained staff; the accuracy required; and the ethical constraints.
The tools are those of the preceding chapters: observation, interview, questionnaire, schedule, case study, and the documentary sources of [Research Methods: Documentary, Empirical and Survey], applied to a sample drawn as in [Sampling].
Two rules. Pilot the instrument before using it, since defects invisible on paper appear at once in the field. And prefer more than one tool where the resources allow, because triangulation is the strongest protection a design has.
Stage 5: The phase of data collection
Preparation: obtaining permissions and access, which in institutional settings takes longer than anything else; recruiting and training field staff so that instruments are applied identically; and preparing the materials.
Fieldwork: administering the instruments, maintaining the sample by pursuing selected respondents rather than substituting available ones, and recording as the design requires.
The Stages of Data Collection
Supervision and quality control: spot checks, re-interviewing a fraction of respondents, and scrutiny of returns as they arrive rather than at the end, when nothing can be corrected.
Records: keeping a field diary, and recording refusals and non-contacts, because the response rate is itself a finding and a report without it cannot be assessed.
The problems to expect, and naming them is realistic rather than pessimistic: refusal; absence and repeated visits; suspicion of the investigator's purpose; the presence of others during interviews, which [The Interview] identifies as acute in Indian households; seasonal absence in agricultural areas; and pressure from those with an interest in the findings.
Stage 6: Data analysis
Editing. Checking the returns for completeness, legibility, consistency and accuracy, in the field where possible so that errors can still be corrected.
Coding. Assigning symbols or numbers to answers so that they can be counted, which requires a code book and, for open questions, a scheme built after reading a sample of the replies.
Classification and tabulation. Grouping the data into classes and presenting them in tables, which is where the shape of the findings first becomes visible.
Statistical analysis. Measures of central tendency, mean, median and mode; measures of dispersion; percentages and rates; cross-tabulation to examine relationships between two variables; correlation, which measures how far two variables move together; and tests of significance, which assess how likely an observed relationship is to have arisen by chance in a sample.
Interpretation, which is the intellectual work: relating the findings to the hypothesis, explaining what they mean, comparing them with earlier studies, and stating what they do not establish.
Two cautions that belong at the end of Module III.
Correlation is not causation. Two variables moving together may be causally connected either way, or both produced by a third, or associated by chance. This is the caution already given in [Types of Methodology] and it is the commonest error in reading social data.
Statistical significance is not importance. A relationship may be statistically reliable and too small to matter, and a large sample makes trivial differences significant.
Reporting, finally, and it belongs to this stage: the report must state the method, the sample, the response rate and the limitations, because a finding whose method is not disclosed cannot be assessed and is not, in the sense of [Social Research: Nature and Purpose], research at all.
A worked example
The area of interest: legal aid.
1. Conceptualising the problem. Narrowed to: among persons produced before magistrates in three districts of this state in a stated year and entitled to free legal services, what proportion were represented at their first production, and what distinguished those who were from those who were not?
The Stages of Data Collection
2. Hypothesis. Representation at first production is associated with the offence charged and with whether the person was produced in a district headquarters court rather than an outlying one. The null hypothesis is that there is no association.
3. Defining the variables. Dependent: representation at first production, operationally defined as the presence of an advocate recorded in the order sheet of the first production. Independent: the offence category as recorded in the first information report; and the court's location, coded as headquarters or outlying. Control: the year, and the district.
4. Choosing the tools. Documentary work on order sheets and legal services authority records; a schedule administered to a sample of the persons produced, since literacy cannot be assumed; and non-participant observation of first productions on sampled days. Piloted in one court first.
5. The phase of data collection. Permissions from the courts and the prison authority; investigators trained together and matched to respondents where appropriate; a stratified multi-stage sample of courts and then of matters; refusals and non-contacts recorded; and a fraction of the schedules re-checked by a supervisor.
6. Data analysis. Editing and coding; tabulation of representation by offence category and by court location; cross-tabulation; a test of significance on the association; and interpretation. And the honest limits stated: the study measures what the order sheet records, which is not necessarily what occurred; three districts do not represent the state; and it establishes association and not cause.
Quick revision
- MU's six stages: conceptualising the problem; laying down the hypothesis; defining the variables; choosing the tools; the phase of data collection; data analysis.
- A good problem is clear, researchable, significant, feasible and ethical. A vague question cannot be repaired later.
- A hypothesis is a testable tentative proposition about a relationship between variables. Null states no relationship and is what is actually tested; alternative states one; directional and non-directional. Exploratory and descriptive studies may properly have none.
- Variables: independent (cause), dependent (effect), intervening, control. Operational definition states a variable by the operations that measure it, without which data are not comparable.
- Tool choice is decided by the question, the population's literacy and dispersion, the kind of data, resources, accuracy required and ethics. Pilot, and triangulate.
- Collection: permissions, training, maintaining the sample rather than substituting, supervision and spot checks, and recording refusals, since the response rate is a finding.
- Analysis: editing, coding, classification and tabulation, statistics (central tendency, dispersion, cross-tabulation, correlation, significance), interpretation, reporting.
- Correlation is not causation, and statistical significance is not importance.
The Stages of Data Collection
Test yourself
1. Name MU's six stages of data collection. Conceptualising the problem, in which a broad interest is turned into a precise answerable question; laying down the hypothesis, a testable tentative proposition; defining the variables, including their operational definitions; choosing the tools of data collection suited to the question, population and resources; the phase of data collection itself, with its permissions, training, fieldwork and supervision; and data analysis, comprising editing, coding, tabulation, statistical treatment, interpretation and reporting.
2. What makes a good research problem, and why does this stage matter most? It is clear and unambiguous; researchable, in that evidence capable of answering it can actually be obtained; significant, adding something worth knowing; feasible within the available time, money and access; and ethical. The stage matters most because a defect here cannot be repaired later: a study whose question is vague will produce data that answer nothing, however carefully those data are collected and however sophisticated the analysis.
3. What is a hypothesis, and what is the null hypothesis? A hypothesis is a tentative proposition about the relationship between two or more variables, stated so that evidence could support or contradict it. The null hypothesis states that there is no relationship between the variables, and it is the proposition actually tested, the researcher seeking to reject it in favour of the alternative hypothesis, which asserts a relationship. Exploratory and descriptive studies frequently proceed without a hypothesis at all.
4. What is an operational definition and why is it necessary? A statement of a variable in terms of the operations by which it will be measured, so that any investigator would classify the same case identically. Delay, for example, cannot be measured until it is defined as the number of days between institution and final disposal, counted from the court's register and excluding transferred matters. Without operational definitions each investigator applies a private understanding, the returns are not comparable, and the study measures its investigators as much as the world.
5. Give two cautions about data analysis. That correlation is not causation: two variables that move together may be causally related in either direction, may both be produced by a third factor, or may be associated by chance, and distinguishing these requires either an experiment or careful reasoning that excludes the alternatives. And that statistical significance is not importance: a large sample can make a trivial difference statistically reliable, so the size of a relationship must be reported and judged as well as its significance.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.