play icon for videos

Secondary Data Analysis: Sources, Validation & Integration

Secondary data analysis: where to find quality public data, how to validate before reuse, and how to integrate it with primary data via persistent IDs.

Updated
July 21, 2026
360 feedback training evaluation
Use Case

What is secondary data?

Secondary data is information that someone else collected, for their own purpose, that you reuse to answer your question: census tables, labor statistics, administrative records, published research, open data portals. You did not design the questions or run the collection. That saves months of fieldwork and introduces one obligation: validating that someone else's definitions actually fit your use.

This page covers the two things people come here for: concrete examples of secondary data, organized by source, and a validation routine that keeps reused data defensible. If you are weighing collecting your own data against reusing existing data, that decision lives on the primary vs secondary data comparison. If you want the firsthand side, start with primary data.

Key takeaways

  • Secondary data is reused data: collected by another organization, for another purpose, under definitions you did not choose.
  • The five source families are government statistical agencies, open data portals, academic repositories, commercial providers, and your own historical records.
  • Every secondary source needs the same four-point check before reuse: origin, definitions, time period, and sample exclusions.
  • Secondary data describes context; it becomes evidence when joined to your own participant records. Sopact calls that persistent record the Outcome Thread.
  • Sopact's Loop methodology treats secondary sources as a standing layer that is re-validated each cycle, not a one-time download pasted into a report.

Examples of secondary data

Common examples of secondary data include census records, government labor statistics, administrative program records, published research findings, industry reports, and city open-data portals. Each already exists because another organization collected it for its own operating or statistical purpose; your analysis is the second use, which is where the name comes from.

Government statistics are the most used examples. The US Census American Community Survey publishes median household income by tract (table B19013), poverty rates, and educational attainment. The Bureau of Labor Statistics publishes employment, wage, and occupation series by metro area. CDC and state health departments publish disease prevalence and vital statistics. These are professionally sampled and documented, which is exactly why they anchor so many needs assessments and grant applications.

Administrative and open data are the fastest growing examples: city portals such as the Chicago Data Portal publish permits, service requests, and program participation; HUD publishes housing assistance and geographic crosswalk files; school districts publish enrollment and attendance. Published research is the academic example: datasets archived alongside peer-reviewed studies, and repositories such as ICPSR or World Bank Open Data. Industry examples include market research reports and benchmark surveys sold by commercial providers.

The examples shift emphasis by field, and the query variants track it. In research, secondary data usually means archived study datasets and official statistics analyzed for a new hypothesis. In statistics teaching, it is any data the investigator did not gather firsthand. In business and marketing, it is industry reports, competitor filings, and syndicated consumer panels. For mission-driven programs, the highest-value examples are the ones that describe the communities served: tract-level income and poverty, regional employment, school performance, health prevalence, and housing cost burden, because those frame need and benchmark outcomes.

One distinction resolves many classroom and exam variants of this question. External secondary data comes from outside your organization, like everything above. Internal secondary data is your own organization's existing records reused for a new question: last year's intake forms, CRM history, prior cohort surveys. The same person's survey answers are primary data for the team that collected them and secondary data for the team that reuses them. Origin is about who collected it and why, not where the file sits.

Sources of secondary data, and how to actually access them

The main sources of secondary data are government statistical agencies, administrative and open-data portals, academic repositories, commercial data providers, and your organization's own historical records. Listing sources is easy. What most guides skip is access: how the data arrives in your analysis, at what granularity, and on what update cadence.

Access has changed materially since 2025 because major statistical agencies now expose data to AI tools directly. The Census Bureau maintains an official MCP server for its data API, so an analyst can pull ACS tables by tract from inside an AI workflow instead of hand-downloading CSVs. The open-source OpenGov MCP server does the same for the hundreds of city and state portals built on Socrata, including Chicago's. The practical effect: joining a public table to your own data went from a day of manual work to a single prompted step.

Five named sources cover most program questions. The Census Bureau's ACS gives demographics and income down to tract level, on one-year and five-year cycles. BLS gives employment, wages, and occupations, monthly, at national to metro grain. HUD gives housing assistance, fair market rents, and the geographic crosswalk files that translate between zip codes and tracts, though its API works on registered tokens with daily limits worth planning around. City open-data portals give the local operational layer: permits, inspections, service requests, program rosters. The World Bank gives country-level development indicators for internationally comparative work. Start with these five and you will rarely need a sixth before the analysis outgrows the question.

Cadence and granularity decide fitness. ACS five-year estimates are stable at tract level but lag two to three years; one-year estimates are current but only for large geographies. BLS series update monthly at metro level. City portals update daily to quarterly, dataset by dataset. A source is not good or bad in general; it is fit or unfit for the specific question, which is what validation establishes.

The four-point check every secondary source must pass

Reused data is only as defensible as its documentation. Before any secondary source enters an analysis, Sopact's practice is a four-point check, recorded where the analysis can cite it. One: origin. Who collected it, for what purpose, and with what method? A convenience sample published as a slick dashboard is still a convenience sample. Two: definitions. What exactly does each field measure? "Unemployment" alone spans several distinct BLS definitions; "household income" may or may not include transfers.

Three: time period. When was it collected, and is that still current for your claim? Citing a pre-2023 workforce statistic as the current labor market is the stale-as-current mistake, and reviewers catch it. Four: sample and exclusions. Who is missing? Administrative data covers only people who reached the service; surveys undercount people without stable addresses. The gaps often sit exactly where mission-driven programs work.

The common failure modes mirror the four points: definition mismatch between the source's field and your indicator, stale data presented as current, overgeneralization that assigns a tract-level average to an individual person, and missing provenance, where a number appears in a report with no citation trail. All four are preventable with a documented check, and none are detectable by an AI summarizer after the fact.

Context is not evidence: where secondary data joins your own

Here is the shift most guides never make. Secondary data describes the environment around your program: the median income of the tracts you serve, the regional employment rate, the district's graduation baseline. It cannot say what changed for the people you actually served. That takes primary data, collected firsthand, on records that persist. Sopact calls that record the Outcome Thread: one participant record under a persistent ID, holding the baseline, every follow-up wave, and the person's own words.

Secondary data earns its place when it joins that thread. Keep a participant's geography on the record, and tract-level ACS context attaches to every individual outcome; a wage gain reads differently in a tract with 28 percent poverty than in one with 6 percent. The join runs on the persistent ID and a documented geographic key, not on a copy-paste into a slide. That is the difference between quoting the census and analyzing with it.

This is a data-model property, not an analysis trick. Form-centric survey tools produce disconnected response files that public data cannot reliably join; record-centric collection, with clean-at-the-source IDs and fields, gives secondary sources something stable to attach to. The architecture is the same one described on the stakeholder intelligence pillar and applied across methods in mixed-mode data collection.

Primary vs secondary data: the comparison at a glance

Across ten dimensions, the pattern is consistent: primary data trades cost and time for fit and control, while secondary data trades fit for speed and coverage. The table compresses the working differences; every row is expanded somewhere on this page or its two sibling guides.

Ten dimensions, two kinds of data
DimensionPrimary dataSecondary data
OriginCollected by you, firsthandCollected by someone else, reused
Purpose fitDesigned for your exact questionDesigned for another purpose; fit must be checked
ExamplesYour surveys, interviews, observations, assessmentsCensus tables, BLS series, portals, published research
CostHigh: instruments, fielding, follow-upLow: often free or already owned
TimeWeeks to months to collectAvailable immediately
Control of definitionsYou lock wording, scales, timingFixed by the original collector
CurrencyAs fresh as your last waveOften lagged one to three years
CoverageYour cohort onlyPopulations no single team could reach
Can it prove change in your participants?Yes, with waves on persistent recordsNo; it benchmarks the environment
Main riskCollection quality is entirely on youDefinition mismatch, staleness, unknown exclusions

For worked examples of each row, see primary data for the collection side and secondary data for source-by-source validation.

Primary and Secondary Data Alignment for Compliance Example

Sopact owns clean primary evidence with persistent IDs; the operational system owns secondary quantities; the dictionary layer is where reconciliation, coding, and audit traceability happen

Primary vs. secondary data
Two feeder systems, one compliance output
The compliance anchor is IRS Form 990 Schedule H — community-benefit reporting for nonprofit hospitals under ACA §501(r). Workday and Sopact are the two feeder systems, one operational and one primary-evidence. The data dictionary is the layer that reconciles both into Schedule H line items.
Sources
Primary data · Sopact
CHNA community surveys / interviews
Patient & participant outcomes
Program feedback, open-text stories
Focus groups, equity / access voice
clean-at-source · persistent contact IDs
Secondary data · Workday + systems
Workday: workforce, payroll, labor cost
Workday: health-professions education / residency hrs
Finance/GL: charity care $, Medicaid shortfall, bad debt
EHR: utilization, subsidized service volumes
Public/census: community need denominators
Transformation
Data dictionary — mapping & normalization layer
source_field → standardized metric → Schedule H line + §501(r) dimension
reconcile units/cost basis · dedupe on contact ID · qual → coded themes (IRIS+ / 5 Dimensions)
audit trail: every output value traces to a source row
Compliance output
990 Schedule H
Part I community benefit at cost
Part II–III building, bad debt, Medicare
Part V facility / Part VI narrative
§501(r) evidence
CHNA report (every 3 yrs)
Financial Assistance Policy proof
Implementation strategy tracking
Board / auditor pack
Decision Brief per program
Impact + outcome dashboard
Reviewer-ready traceability
The transformation pattern
What the data dictionary transformation actually looks like. Every row is: source field → normalization → target compliance field. Four transformation types cover almost everything.
Source field (system) Transformation Target (Schedule H / §501r)
Labor cost by cost center (Workday) Allocate to subsidized services, convert to net cost Part I, community health improvement at cost
Residency / CE hours (Workday) × loaded cost rate Part I, health professions education
Charity care charges (GL/EHR) Apply cost-to-charge ratio → net cost Part I, financial assistance at cost
CHNA survey open-text (Sopact) Code to theme → link to prioritized need Part V CHNA + §501(r) implementation strategy
Participant outcome + contact ID (Sopact) Dedupe, pre/post delta, roll to program Part VI narrative + board evidence
Adjacent use cases
The same primary + secondary + dictionary pattern, tuned to a different compliance or funder anchor.
Program Primary (Sopact) Secondary Output
Nonprofit hospital CHNA surveys, patient voice Workday, GL, EHR 990 Schedule H, §501(r)
Workforce training Exit + 6-month trainee survey BLS regional data, internal LMS Funder outcome report
Education program Classroom teacher / student survey District + state assessment data Funder / district report
Funder portfolio Cross-cutting grantee survey Census/ACS, grantee 990s Portfolio rollup report

Similar use cases — same primary/secondary/output shape:

  • CSRD / ESRS double materiality — primary: stakeholder materiality surveys (Sopact); secondary: Workday workforce + emissions/finance systems; output: ESRS datapoints (this is the Eric Darrisaw pattern already in your notes).
  • Workforce / WIOA program compliance — primary: participant outcomes & follow-up (Sopact); secondary: HRIS enrollment, wage records; output: WIOA performance measures.
  • Foundation grant compliance — primary: grantee reporting (Sopact); secondary: financial disbursements; output: funder reports + logic-model attainment.
  • CDFI / community investment — primary: borrower impact surveys; secondary: loan/servicing system; output: CDFI Fund transaction-level reporting.
  • Hospital DEI / community equity — primary: staff & patient experience; secondary: Workday demographics; output: board equity scorecard.

Two eras of working with public data

In the survey-platform era, secondary data lived in reports and primary data lived in survey tools, and the two met once a year in a slide deck. An analyst downloaded a census table, screenshotted a chart, and wrote "compared to the county average" next to a number from a different universe, a different year, and a different definition. Nobody could audit the comparison because it was never a join, only a juxtaposition.

In the record-centric era, the two layers connect. Your own collection stays clean at the source, so every record carries the IDs and geography a public table can join on; agency MCP servers make the public side promptable; and validation is written down once per source, then re-checked on a cadence instead of assumed forever. The evaluation test for any platform claiming this: ask to see one participant's outcome next to the public benchmark for that participant's own geography, with the source, vintage, and definition of the benchmark cited on the same screen. A tool that can only show the benchmark in a separate dashboard is still in era one.

Validation is not a one-time event. The Loop keeps it current.

A secondary source validated in January is an assumption by September: series get revised, definitions change, portals deprecate datasets. That is why reuse belongs inside the Loop, Sopact's method for continuous impact intelligence: collect clean at the source, analyze the moment data arrives, improve while you can still act. Secondary sources sit in the collect layer as standing context, re-validated each cycle rather than re-discovered each reporting season.

The Loop also supplies the audit trail reuse demands. When a funder asks where a benchmark came from, the answer should be the source, table, vintage, and pull date, attached to the analysis itself. That standard is the subject of Loop traceability, and the same-definition-every-wave discipline that makes joins hold is covered in Loop reliability.

One method, three moves that never stop

1 · CollectClean at the source; secondary sources documented and joined by persistent keys.
2 · AnalyzeOn arrival; your records read against public context, every benchmark cited.
3 · ImproveIn time to act; revalidate sources on a cadence, not at report deadline.

Then the cycle runs again, a little sharper each time. Read the method: the Loop methodology →

Put secondary data to work this week

The fastest way to internalize the validation habit is to run it on a source you already cite. Each prompt below is written to paste into Sopact Sense's Assistant, or to reason through with your team; the arrow above each one links the Academy walkthrough with the expected output and tips.

Academy walkthrough → How to build a data dictionary

Here are the secondary sources my report currently cites: [PASTE SOURCES]. For each, create a data dictionary entry: who collected it and why, the exact field definitions I rely on, the vintage and update cadence, and known sample exclusions. Flag any source that fails the four-point check for the claim I am using it to support.

Academy walkthrough → Connect quantitative and qualitative survey data

My participant records include [FIELDS, e.g. zip code, employment status, open-ended reflections]. Propose how to join tract-level ACS context (income, poverty rate) to each record, state the geographic key and its limits, and show how the joined context should change how I read the open-ended answers.

Academy walkthrough → Clean open-ended survey responses

Before I benchmark my cohort against public statistics, audit my own data for join-readiness: [PASTE OR ATTACH SAMPLE]. Identify records missing the keys a secondary join needs (geography, dates, consistent IDs), duplicates that would double-count against a benchmark, and the cleanup order that fixes collection rather than patching this file once.

Academy walkthrough → The Loop methodology: continuous, not annual

We refresh our secondary sources [CURRENT CADENCE, e.g. once a year at report time]. Given these sources: [PASTE LIST], design a revalidation calendar: which series revise monthly vs annually, what to re-check at each touch, and which single source, if stale, would most distort our next funder report.

Learn the how-to in the Academy

Each walkthrough is practical and short: what to do, the prompt to run, the output to expect, and the tips that make it reliable.

Watch: bringing public data into an impact analysis without losing the audit trail.

Frequently asked questions

What is secondary data?

Secondary data is information collected by someone else, for another purpose, that you reuse for your own question: census tables, labor statistics, administrative records, published research. In Sopact's framing it supplies context, while primary data on a persistent Outcome Thread supplies the change evidence.

What are examples of secondary data?

Examples of secondary data include Census ACS income and poverty tables, Bureau of Labor Statistics employment series, CDC health statistics, city open-data portals like Chicago's, HUD housing datasets, published research archives, industry benchmark reports, and your own prior-year program records. Sopact's rule: every example passes a four-point check before reuse.

What are the sources of secondary data?

The main sources of secondary data are government statistical agencies, administrative and open-data portals, academic repositories, commercial data providers, and internal historical records. Sopact treats each source as a documented entry in a data dictionary, with origin, definitions, vintage, and exclusions recorded before first use.

What is the difference between internal and external secondary data?

Internal secondary data is your own organization's existing records reused for a new question, such as last year's intake forms or CRM history. External secondary data comes from outside bodies like the Census Bureau or BLS. Sopact applies the same four-point validation to both, because internal data drifts in definition too.

What is secondary data in research?

In research, secondary data means analyzing data you did not collect: existing datasets, archives, or administrative records examined for a new question. Sopact's guidance for practitioners is identical to the academic standard: document origin, definitions, time period, and sampling before the source touches a finding.

What are the advantages and disadvantages of secondary data?

Advantages: speed, low cost, coverage no single organization could collect, and professional sampling in official statistics. Disadvantages: definitions you did not choose, time lag, unknown exclusions, and no access to the people behind the numbers. Sopact's resolution is to pair it with primary data on the Outcome Thread, so context and change evidence travel together.

How do you validate a secondary data source?

Run the four-point check Sopact uses: origin (who collected it, why, and how), definitions (what each field actually measures), time period (is the vintage current for your claim), and sample exclusions (who is missing). Record the answers with the analysis so every benchmark stays citable.

Can secondary data prove program outcomes?

No. Secondary data benchmarks the environment; it cannot show change in the specific people you served. Outcome claims need primary data, collected across waves on the same participant record. Sopact's Outcome Thread joins the two, so your measured outcomes sit beside the public baseline for the same geography.

How does Sopact Sense handle secondary data?

Sopact Sense keeps primary collection clean at the source, with persistent IDs and geography on every record, so public tables join reliably. Analysts pull agency data through tools like the Census MCP server, and the Loop revalidates each source on a cadence, keeping benchmarks current instead of frozen at report time.

Next: see the firsthand side on the primary data guide, weigh the two on primary vs secondary data, or step back to how methods fit together in data collection methods.