What is secondary data?
Secondary data is existing information reused for a purpose different from the original collection or by someone other than the original collector. Examples include census tables, administrative records, published studies and earlier surveys.
The distinction depends on the use. A survey is primary data for the team that collected it for that study; it may become secondary data for a later analysis. Secondary does not mean inferior. Fitness for the question matters more than the label.
Examples and where to find them
| Source | Possible use | Check before use |
|---|---|---|
| National statistics or census | Describe population and local conditions. | Geography, release date and definitions. |
| World Bank indicators | Compare economic or development context. | Country coverage, indicator metadata and revisions. |
| Administrative records | Examine service volumes or historical activity. | Why records were collected and who is missing. |
| Earlier surveys | Establish context or investigate a new question. | Sampling, questionnaire and reuse permissions. |
| Partner reports | Review delivery across periods. | Reporting boundaries, units and version changes. |
Start with the original publisher. The World Bank data portal and UN Statistics Division provide entry points; national statistical offices often provide more relevant local detail. Retain the dataset name and version, not just a dashboard screenshot.
Four checks before analysis
- Relevance: does the source measure the concept, population, location and period in your question?
- Quality: how was it collected, what is missing and what known errors or limitations exist?
- Comparability: are units, boundaries, classifications and methods consistent across the records you want to compare?
- Permission: does the license or agreement allow the intended reuse, linking and publication?
A recent publication can contain older observations. Record the observation period separately from the release date. If the publisher revises earlier values, retain the version used so the result can be reproduced.
How to analyze secondary data
- Write the question before selecting convenient data.
- Read the metadata and original collection method.
- Inspect coverage, missing values, duplicates, units and unusual values.
- Define any exclusions and transformations. Keep the original source intact.
- Analyze at the level supported by the data.
- Report methods, limitations and the source version with the finding.
Treat coded missing values carefully. A blank, “not applicable” and zero can have different meanings. Do not convert them all to zero merely to simplify a chart.
Worked example: a regional benchmark is not a participant result
Suppose an illustrative training cohort reports 62% employment at follow-up. A public table reports 70% employment for the region. It is tempting to describe an eight-point performance gap, but first check age range, eligibility, employment definition and observation period.
If the public table includes a different population, it may be useful context rather than a fair benchmark. It also does not show what the trainees would have experienced without the program. Explain that limitation instead of treating the regional rate as a ready-made control group.
Combine secondary and primary evidence carefully
Public data may describe the setting, while your survey or interview describes participant experience. Keep the levels distinct. A neighborhood-level statistic does not establish the condition of every person living there.
Where records can appropriately be linked, use reliable keys and check unmatched and duplicate records. Keep the join rules, permissions and source dates. Matching records technically does not establish that the measures are conceptually comparable.
Advantages and limitations
| Advantage | Corresponding limitation |
|---|---|
| Can reduce new collection work. | Cleaning and interpretation still require time. |
| May cover long periods or large populations. | Definitions and methods may change over time. |
| Can provide independent context. | May not describe the people your program serves. |
| Can reveal questions for new research. | Often lacks the detail needed to answer those questions fully. |
Keep documents and data definitions together
A partner spreadsheet may depend on definitions in a separate reporting guide. Keep that guide, the reporting period and any revisions connected to the data. Document extraction or AI-assisted classification can help locate relevant information, but verify units, tables and exceptions against the source.
Food4Education’s published story describes a starting point in operational supply-chain data and a developing wider evidence foundation. It illustrates why existing files need context; it does not establish causal school or livelihood outcomes. Read the story.
Frequently asked questions
Is a literature review secondary data?
It uses existing sources. Whether a specific project is described as secondary data analysis depends on whether it reanalyzes data or synthesizes published findings.
Can secondary data be qualitative?
Yes. Existing interviews, documents and written records can be reused when permissions and context support the new purpose.
Is publicly available data unrestricted?
No. Check the license, terms and sensitivity of the information. Public access does not settle every reuse question.
Can secondary data replace a baseline?
Sometimes an earlier source is suitable, but only if it represents the relevant measure, population and period. Label a reconstructed baseline and its limitations.

