Final Deadline for Fall 2026 Programme is October 10. Apply Now →
HomeOCAR AcademyOxford Scholar ProgrammeOxford Publication FellowshipOCAR Student Research SymposiumOxford Essay PrizeOur AlumniOutreachPodcastsResearch FellowsBlogAboutFAQsFor Parents & CounselorsContactAdmissions Portal
A spreadsheet of survey data beside a notebook of handwritten variable names
Research Guides

Where to find a free dataset you can actually use

In short

Narrow questions come from data, not from brainstorming — so find the dataset first. Twenty-two free sources are listed here by field, with what each holds. Before committing to one, run four checks: can you actually download it, is the documentation readable, is there a variable pairing nobody has reported, and can you open it with the software you have.

The standard advice is to pick a question you care about and then find evidence for it. That order is why most student projects stall.

Invert it. Find data you can actually reach, spend an hour in its documentation, and let the question come out of what is there.

The datasets

By field, with what each one actually holds. Everything here is free.

Economics, development and labour

SourceWhat it holdsNotes
IPUMSHarmonised census and survey microdata, US and internationalFree registration. The single richest source on this list
World Bank Open DataCross-country indicators, decades deepNo registration, downloads straight to spreadsheet
ONSUK population, prices, labour market, housingBest starting point for almost any UK question
OECD DataComparable indicators across member countriesGood for cross-national comparison without harmonisation work
Our World in DataCurated long-run series with sources citedCleaned and documented; follow through to the primary source

Public health and epidemiology

SourceWhat it holdsNotes
CDC WONDERUS mortality, natality, disease surveillanceQuery interface; results export cleanly
NHS England statisticsActivity, waiting times, prescribingMonthly releases, spreadsheet format
WHO Global Health ObservatoryCross-country health indicatorsPairs well with World Bank for income controls
UK Biobank showcaseMetadata on a very large cohortFull access is restricted; the showcase is browsable and useful for scoping

Psychology, education and social attitudes

SourceWhat it holdsNotes
Monitoring the FutureAnnual US adolescent behaviour surveyThe best adolescent dataset for a student project
General Social SurveyUS attitudes since 1972Long time series, easy online analysis tool
PISAInternational 15-year-old attainment plus contextEnormous, well documented, widely used
UK Data ServiceUK social surveys and cohort studiesFree registration; the UK equivalent of IPUMS
DfE statisticsEnglish schools: attainment, attendance, intakeGranular to local authority level

Politics and government

SourceWhat it holdsNotes
ANESUS election surveys since 1948The standard source for US voting behaviour
House of Commons LibraryUK constituency data, briefings, election resultsResearch-grade and written to be readable
Electoral CommissionUK turnout, spending, donationsDonation data is underused and genuinely interesting

Environment and ecology

SourceWhat it holdsNotes
DEFRA air qualityUK monitoring station readings, hourlyPairs with health or school location data
BTO Breeding Bird SurveyUK bird abundance by species and siteDecades of consistent methodology
Nature's CalendarUK phenology: first leaf, first flight, by yearGood for climate-signal questions
NOAA Climate Data OnlineGlobal weather station recordsLong series, station-level granularity

Text, code and everything else

SourceWhat it holdsNotes
GitHub REST APIRepositories, contributors, issues, timingsBuild your own dataset about open-source behaviour
Wikipedia APIArticle text, edit histories, pageviewsEdit histories are a dataset almost nobody uses
Google Books NgramsWord frequency across centuries of booksFast way into a language-change question

Four checks before you commit

Each takes minutes and each has killed a project at week six for somebody.

1. Can you actually download it today? Not "is it available" — can you have a file open in the next thirty minutes. Some sources need a free registration that takes a day. Some need an institutional affiliation you do not have. UK Biobank's full data is a formal application, not a download. Find out now.

2. Is the documentation readable by you? Open the codebook and read ten variable definitions. If they make sense, you are fine. If they assume a statistics degree, pick something else — you will be reading that document more than any other in the project.

3. Is there a pairing nobody has reported? Search Google Scholar for your two variables plus the dataset name. Finding prior work is good news: it confirms the question matters and gives you something to position against. Finding your exact analysis means narrow further, not abandon.

4. Can you open it with what you own? A 40MB CSV opens in a spreadsheet. A 4GB fixed-width microdata extract does not. Match the file to your tools before you commit, not after.

How to turn a dataset into a question in one afternoon

Four steps, in order.

Spend a real hour in the codebook. Not skimming — reading variable definitions and noting the ones you did not expect to exist. Surprising variables are where unclaimed questions live.

Write down three pairings. Two variables and a plausible relationship. "Screen time and sleep duration." "Betting shop density and personal insolvency." "Press freedom score and vaccination uptake."

Name the obvious confounder for each. If you cannot think of one, you have not understood the pairing yet. If the confounder obviously explains everything, drop that pairing.

Say out loud what result would surprise you. If no result would, the question is not open and you are about to spend four months confirming something you already believe.

That is the same process a mentor runs in the first two sessions, and there is nothing stopping you running it alone this week.

If you want to publish afterwards

Choose with the venue in mind, because it constrains the question.

Most student research journals want original findings rather than a literature review. A clean analysis of public data counts as original — the Journal of Emerging Investigators wants new measurements or a novel analysis, and a novel analysis of existing data qualifies. If your project ends up as a synthesis of other people's findings instead, NHSJS and the Columbia Junior Science Journal both accept literature reviews, which most of the category does not.

One practical note: check whether the dataset's terms allow you to republish extracts. IPUMS and the UK Data Service both have redistribution conditions, and that affects what you can put in an appendix.

Where we fit

The Oxford Centre for Advanced Research runs 1:1 mentorships with PhD candidates and early-career researchers at Oxford, Cambridge, other leading UK universities and Ivy League institutions. The Oxford Scholar Programme is £2,000 — about $2,600 — for ten contact hours over ten to fourteen weeks; the Oxford Publication Fellowship is £3,400, roughly $4,400.

Every dataset on this page is free and none of it requires us. What a mentor adds is someone who has read a codebook before and will tell you in week one that your pairing is confounded, rather than in week nine. If you can get that from a teacher with a doctorate, get it there first.

The one-hour version

Pick the source on this page closest to your subject. Open its documentation. Read until you find a variable you did not expect.

Most good student projects start exactly there, and the hour costs you nothing.

Frequently asked questions

What is the best dataset to start with?

For most students, a national survey with good documentation — Monitoring the Future, the General Social Survey, or a World Bank indicator set. They are free, instantly downloadable, documented for non-specialists, and large enough that unchecked questions genuinely remain. The barrier to the fancier sources is usually registration, not ability.

Do I need to code?

Not always. A spreadsheet handles aggregate data — World Bank indicators, ONS releases, CDC WONDER query results — and plenty of good projects never leave one. Individual-level microdata, like an IPUMS extract or a raw survey wave, is where a few weeks of basic Python or R starts paying for itself.

How do I know if a question has already been answered?

Search Google Scholar for your two variables plus the dataset name. If you find the exact analysis, narrow further — a subgroup, a different period, a control nobody applied. Finding prior work on your topic is good news, not bad; it tells you the question matters and gives you something to position against.

Is using someone else's data still original research?

Yes, and most published research works this way. Originality lives in the question and the analysis, not in who collected the numbers. An economist analysing census data did not run the census. What is not original is reproducing an analysis that already exists without adding anything.

What size dataset do I need?

Big enough that a difference you find is unlikely to be noise, which in practice usually means hundreds of observations at minimum and thousands if you want to control for anything. A survey of your own year group — perhaps ninety people — is rarely enough to support the claims students want to make from it.

Do I need permission to use these?

The sources here are public and free. Some require a free registration and an agreement about redistribution — IPUMS and the UK Data Service both do. Read the terms, because a few prohibit republishing raw extracts, which affects what you can put in an appendix.