Narrow questions come from data, not from brainstorming — so find the dataset first. Twenty-two free sources are listed here by field, with what each holds. Before committing to one, run four checks: can you actually download it, is the documentation readable, is there a variable pairing nobody has reported, and can you open it with the software you have.
The standard advice is to pick a question you care about and then find evidence for it. That order is why most student projects stall.
Invert it. Find data you can actually reach, spend an hour in its documentation, and let the question come out of what is there.
The datasets
By field, with what each one actually holds. Everything here is free.
Economics, development and labour
| Source | What it holds | Notes |
|---|---|---|
| IPUMS | Harmonised census and survey microdata, US and international | Free registration. The single richest source on this list |
| World Bank Open Data | Cross-country indicators, decades deep | No registration, downloads straight to spreadsheet |
| ONS | UK population, prices, labour market, housing | Best starting point for almost any UK question |
| OECD Data | Comparable indicators across member countries | Good for cross-national comparison without harmonisation work |
| Our World in Data | Curated long-run series with sources cited | Cleaned and documented; follow through to the primary source |
Public health and epidemiology
| Source | What it holds | Notes |
|---|---|---|
| CDC WONDER | US mortality, natality, disease surveillance | Query interface; results export cleanly |
| NHS England statistics | Activity, waiting times, prescribing | Monthly releases, spreadsheet format |
| WHO Global Health Observatory | Cross-country health indicators | Pairs well with World Bank for income controls |
| UK Biobank showcase | Metadata on a very large cohort | Full access is restricted; the showcase is browsable and useful for scoping |
Psychology, education and social attitudes
| Source | What it holds | Notes |
|---|---|---|
| Monitoring the Future | Annual US adolescent behaviour survey | The best adolescent dataset for a student project |
| General Social Survey | US attitudes since 1972 | Long time series, easy online analysis tool |
| PISA | International 15-year-old attainment plus context | Enormous, well documented, widely used |
| UK Data Service | UK social surveys and cohort studies | Free registration; the UK equivalent of IPUMS |
| DfE statistics | English schools: attainment, attendance, intake | Granular to local authority level |
Politics and government
| Source | What it holds | Notes |
|---|---|---|
| ANES | US election surveys since 1948 | The standard source for US voting behaviour |
| House of Commons Library | UK constituency data, briefings, election results | Research-grade and written to be readable |
| Electoral Commission | UK turnout, spending, donations | Donation data is underused and genuinely interesting |
Environment and ecology
| Source | What it holds | Notes |
|---|---|---|
| DEFRA air quality | UK monitoring station readings, hourly | Pairs with health or school location data |
| BTO Breeding Bird Survey | UK bird abundance by species and site | Decades of consistent methodology |
| Nature's Calendar | UK phenology: first leaf, first flight, by year | Good for climate-signal questions |
| NOAA Climate Data Online | Global weather station records | Long series, station-level granularity |
Text, code and everything else
| Source | What it holds | Notes |
|---|---|---|
| GitHub REST API | Repositories, contributors, issues, timings | Build your own dataset about open-source behaviour |
| Wikipedia API | Article text, edit histories, pageviews | Edit histories are a dataset almost nobody uses |
| Google Books Ngrams | Word frequency across centuries of books | Fast way into a language-change question |
Four checks before you commit
Each takes minutes and each has killed a project at week six for somebody.
1. Can you actually download it today? Not "is it available" — can you have a file open in the next thirty minutes. Some sources need a free registration that takes a day. Some need an institutional affiliation you do not have. UK Biobank's full data is a formal application, not a download. Find out now.
2. Is the documentation readable by you? Open the codebook and read ten variable definitions. If they make sense, you are fine. If they assume a statistics degree, pick something else — you will be reading that document more than any other in the project.
3. Is there a pairing nobody has reported? Search Google Scholar for your two variables plus the dataset name. Finding prior work is good news: it confirms the question matters and gives you something to position against. Finding your exact analysis means narrow further, not abandon.
4. Can you open it with what you own? A 40MB CSV opens in a spreadsheet. A 4GB fixed-width microdata extract does not. Match the file to your tools before you commit, not after.
How to turn a dataset into a question in one afternoon
Four steps, in order.
Spend a real hour in the codebook. Not skimming — reading variable definitions and noting the ones you did not expect to exist. Surprising variables are where unclaimed questions live.
Write down three pairings. Two variables and a plausible relationship. "Screen time and sleep duration." "Betting shop density and personal insolvency." "Press freedom score and vaccination uptake."
Name the obvious confounder for each. If you cannot think of one, you have not understood the pairing yet. If the confounder obviously explains everything, drop that pairing.
Say out loud what result would surprise you. If no result would, the question is not open and you are about to spend four months confirming something you already believe.
That is the same process a mentor runs in the first two sessions, and there is nothing stopping you running it alone this week.
If you want to publish afterwards
Choose with the venue in mind, because it constrains the question.
Most student research journals want original findings rather than a literature review. A clean analysis of public data counts as original — the Journal of Emerging Investigators wants new measurements or a novel analysis, and a novel analysis of existing data qualifies. If your project ends up as a synthesis of other people's findings instead, NHSJS and the Columbia Junior Science Journal both accept literature reviews, which most of the category does not.
One practical note: check whether the dataset's terms allow you to republish extracts. IPUMS and the UK Data Service both have redistribution conditions, and that affects what you can put in an appendix.
Where we fit
The Oxford Centre for Advanced Research runs 1:1 mentorships with PhD candidates and early-career researchers at Oxford, Cambridge, other leading UK universities and Ivy League institutions. The Oxford Scholar Programme is £2,000 — about $2,600 — for ten contact hours over ten to fourteen weeks; the Oxford Publication Fellowship is £3,400, roughly $4,400.
Every dataset on this page is free and none of it requires us. What a mentor adds is someone who has read a codebook before and will tell you in week one that your pairing is confounded, rather than in week nine. If you can get that from a teacher with a doctorate, get it there first.
The one-hour version
Pick the source on this page closest to your subject. Open its documentation. Read until you find a variable you did not expect.
Most good student projects start exactly there, and the hour costs you nothing.
Frequently asked questions
What is the best dataset to start with?
For most students, a national survey with good documentation — Monitoring the Future, the General Social Survey, or a World Bank indicator set. They are free, instantly downloadable, documented for non-specialists, and large enough that unchecked questions genuinely remain. The barrier to the fancier sources is usually registration, not ability.
Do I need to code?
Not always. A spreadsheet handles aggregate data — World Bank indicators, ONS releases, CDC WONDER query results — and plenty of good projects never leave one. Individual-level microdata, like an IPUMS extract or a raw survey wave, is where a few weeks of basic Python or R starts paying for itself.
How do I know if a question has already been answered?
Search Google Scholar for your two variables plus the dataset name. If you find the exact analysis, narrow further — a subgroup, a different period, a control nobody applied. Finding prior work on your topic is good news, not bad; it tells you the question matters and gives you something to position against.
Is using someone else's data still original research?
Yes, and most published research works this way. Originality lives in the question and the analysis, not in who collected the numbers. An economist analysing census data did not run the census. What is not original is reproducing an analysis that already exists without adding anything.
What size dataset do I need?
Big enough that a difference you find is unlikely to be noise, which in practice usually means hundreds of observations at minimum and thousands if you want to control for anything. A survey of your own year group — perhaps ninety people — is rarely enough to support the claims students want to make from it.
Do I need permission to use these?
The sources here are public and free. Some require a free registration and an agreement about redistribution — IPUMS and the UK Data Service both do. Read the terms, because a few prohibit republishing raw extracts, which affects what you can put in an appendix.