What is the difference between primary and secondary data?
Primary data is collected for the question or study at hand. Secondary data is existing information reused for a new analysis or purpose. The distinction depends on how the data relates to the current study, not simply whether it comes from inside or outside your organization.
A survey designed for your current evaluation is primary collection. A prior survey, administrative record or public dataset reused for a new question is secondary data for that analysis. Your own organization can therefore be both the original collector and a later secondary user.
Neither category guarantees quality, relevance or causal evidence. Primary collection can be poorly designed; secondary records can be rich, timely and longitudinal. Start with the question and assess whether each source can support the claim you want to make.
Primary vs. secondary data at a glance
| Dimension | Primary data | Secondary data |
|---|---|---|
| Relationship to the current question | Collected specifically for it | Reused from existing collection |
| Examples | New surveys, interviews, observations or assessments | Prior surveys, administrative records, published datasets or research archives |
| Control over collection | You can design questions, timing and procedures | You work with the original definitions and available documentation |
| Typical work required | Design, recruit, collect, validate and follow up | Find, obtain access, understand, validate and prepare for reuse |
| Potential advantage | Can target a specific information gap | May offer existing coverage, history or scale |
| Main risk | Collection may miss the population or measure the wrong thing | The original purpose, definitions or coverage may not fit |
There is no fixed rule that primary data is expensive or secondary data is inexpensive. A short new survey can be modest; access to a complex existing dataset can require substantial preparation and expertise. Estimate the work needed to make either source useful.
Examples of each kind of data
Primary collection for a current question
A training provider interviews participants to understand barriers to applying a new skill. A service team asks customers about a newly changed onboarding step. A network asks local organizations for information needed in this year's shared analysis. These are primary collection when they are designed for the current purpose.
Secondary use of existing information
A team reuses last year's customer surveys, analyzes routine service records for a new question, consults official population statistics or studies an archived panel dataset. Existing information may include quantitative measures, interview transcripts, documents or repeated observations.
The UK Data Service describes secondary analysis as reanalysis of previously collected qualitative or quantitative data. Reuse can answer valuable questions, provided the original collection and the new purpose are understood.
When should you collect new data?
Collect new data when existing sources leave a material gap: the right population is missing, the measure does not answer the question, the time period is unsuitable or the necessary context was never collected. Design the smallest justified collection that can fill that gap.
For example, service records may show repeated contacts but not why people found a process difficult. A focused interview or open-ended survey can add that perspective. Do not ask respondents to repeat information already available and appropriate to reuse merely because a new form makes it convenient.
Plan eligibility, recruitment, wording, accessibility, languages, review and respondent burden. Being the original collector gives you responsibility for those choices, not automatic confidence in the result.
When should you use existing data?
Begin by checking what already exists. Prior surveys may provide history; administrative records may describe repeated events; official statistics may provide population context. Reuse can avoid unnecessary collection and allow questions that a small new study could not address alone.
Validate the original purpose and the process that produced the records. A field entered for billing may not have the meaning an evaluator expects. A missing entry may mean not asked, not applicable, unavailable or simply not recorded. These distinctions affect interpretation.
Confirm access and permitted use, available identifiers, documentation and whether the data can be updated or reproduced. Do not assume public availability means every reuse is appropriate or that the dataset contains the detail you need.
How to assess whether a source is fit for purpose
| Check | What to establish |
|---|---|
| Origin | Who collected it, why and through which process? |
| Population | Who or what is covered, excluded or underrepresented? |
| Definition | What does each measure actually mean? |
| Time | When was it collected, and what period does it describe? |
| Unit | Does a row represent a person, account, event, site or aggregate? |
| Quality | What is known about errors, missingness, duplication and uncertainty? |
| Access and use | Who may use it, for what purpose, and with what restrictions? |
| Version | Can the source, extraction date and later revisions be traced? |
Apply these checks to your own records as well as outside sources. Internal data can lose context as staff, definitions and systems change. Preserve a source note with the analysis so the next reviewer does not have to reconstruct those decisions.
How to combine primary and secondary data
Define the role of each source before joining it. One source may describe the population, another track an individual's experience, and another provide context for interpretation. Putting them in the same table does not make them equivalent evidence.
Use documented keys such as a compatible organization identifier, geography or period where appropriate. Test duplicates, unmatched records and one-to-many relationships. A join that duplicates each participant three times can inflate a count while leaving the table looking plausible.
Keep the original fields and source identifiers. Document transformations and mappings rather than replacing unlike measures with one label. If categories or geographic boundaries differ, determine whether a defensible conversion is possible and record its limitations.
Do not join person-level records when an aggregate comparison is sufficient. Retain only the context needed for the purpose and apply the relevant access restrictions to the combined result.
A worked example: training and employment
A fictional training provider wants to understand participants' transition into work. It has attendance records collected during delivery, collects a new follow-up survey about employment experience and consults regional labor statistics.
The attendance records are reused operational data for this analysis. The new survey is primary collection for the current question. The labor statistics provide broader context. None of those labels tells you whether the records are accurate or comparable; each source needs validation.
The team checks who answered the follow-up, how employment is defined, the reference dates and whether the regional statistic concerns a comparable population. An adult regional employment rate is not automatically an appropriate benchmark for a specific age group completing a training program.
The combined evidence can describe participant outcomes and their context. It does not establish that the program caused employment merely because the records are connected. A causal claim requires a suitable evaluation design and consideration of alternative explanations.
Can secondary data measure change?
Yes. Existing panel surveys, repeated assessments or administrative histories may contain observations of the same units over time. Secondary analysis can use those records to study change when linkage, measurement, timing and coverage support it.
Primary collection can also be a single snapshot, which does not directly establish individual change. Whether data supports change analysis depends on its design and structure, not whether it is primary or secondary. See longitudinal versus cross-sectional studies.
Keep observation of change distinct from attribution. Even a reliable repeated record may not show what would have happened without a program, service or intervention.
Keep definitions and source history with the analysis
Use a shared data dictionary for the core fields needed across teams. Local sources can retain additional detail. The purpose is to make the common comparison meaningful, not force every contributor to use one form for every question.
Retain dates for changing attributes such as location or employment status. Preserve the version of an external benchmark used in a report, and decide how later revisions will be handled. Replacing a source silently can make an earlier published number impossible to reproduce.
Link quantitative results to the relevant narrative evidence where useful. An interview may explain a pattern, challenge it or identify a missing question. Do not treat a quotation as a substitute for a population estimate, or assume a contextual statistic explains an individual's experience.
How Sopact helps manage the combined workflow
Sopact brings new collection, existing records, qualitative evidence and quantitative analysis into a connected workflow. Teams can define shared fields, retain context and inspect the records behind findings rather than rebuilding separate files for every reporting cycle.
The practical benefit is less repeated reconciliation and a clearer review process as sources accumulate. Test the needed import or integration, field mappings, permissions, exception handling and source links. Do not assume an outside source will connect automatically or that software can resolve incompatible definitions.
Estimate total ownership effort: source discovery, access, collection, cleaning, mapping, review, maintenance and reporting. Reuse may reduce collection effort but increase preparation work. A sensible plan makes those tradeoffs visible before committing to another survey or integration.
Frequently asked questions
Is our own previous survey secondary data?
It is secondary use when you reuse it for a new analysis beyond its original collection purpose. Check whether the original questions, population, permissions and period fit the new question.
Are primary data always more accurate?
No. Accuracy depends on collection and measurement quality. Primary collection can be biased or incomplete, while a well-documented existing source may be highly useful for the intended analysis.
Can both types be qualitative?
Yes. A new interview for the current study is primary collection; an archived interview reused for a new question is secondary evidence. Consider context, access and the limitations of interpreting material collected for another purpose.
Should we always start with secondary data?
Checking existing evidence is usually a useful planning step, but it does not mean every existing source should be used. Collect new information when the material gap justifies the effort and respondent burden.
Where can we go deeper?
Read primary data for collection examples, secondary data for reuse guidance, and data collection methods to choose how new evidence should be gathered.
Watch the design comparison
This Sopact video explains primary and secondary data and how to combine them. Assess each source against the current question before reuse.

