Where our data comes from, how it was collected, and what it can and cannot tell you
This site tracks the full pathway from international graduate enrollment to permanent residency, using exclusively official US government data:
| Stage | Source Agency | Coverage | Universities |
|---|---|---|---|
| Thesis/Dissertation Data | University repositories (OAI-PMH) | 2010-2026 | 105 |
| OPT Students Employed | ICE/SEVP SEVIS Reports | CY 2017-2024 | 59 of 105 (Top 100 list) |
| H-1B Petition Approvals | USCIS H-1B Employer Data Hub | FY 2010-2026 | 101 |
| H-1B Job Titles & Salaries | DOL LCA Disclosure Data | CY 2014-2026 (both ends part-years) | 100 |
| Green Card Sponsorship | DOL PERM Disclosure Data | FY 2024 | 71 |
Note on employer entity names: Universities file H-1B and PERM petitions under their legal entity name, which may differ from their common name and can change over time. Multi-campus university systems (e.g., University of Missouri, SUNY, University of California) often file under a single systemwide entity, making per-campus attribution imprecise. We match employer names to universities using pattern matching with known name variants — 296 patterns, tested on word boundaries, not as bare substrings (see Corrections) — but some filings may be attributed to the wrong campus within a system.
This site draws on ten major data sources, each with different strengths and coverage. Nine are federal; the last, SHEEO SHEF, is the state-finance series used for state appropriations.
IPEDS is the primary federal database for higher education statistics. Every Title IV institution reports annually. We use four IPEDS surveys:
An annual census of every research doctorate recipient from a U.S. institution. Unlike IPEDS, the SED collects individual-level data including source country for temporary visa holders, postgraduation plans, and financial support. The SED is the source for our national doctorate trend charts, stay-rate analysis, and country-of-origin rankings (China, India, etc.). Table 4-1 — primary source of financial support, by citizenship — is the basis for the funding analysis on this site. It is the only federal collection that crosses how a doctorate was paid for with the recipient’s citizenship status. Because the SED is a census administered by institutions and completed before graduation, its coverage is effectively complete rather than a sample.
Every chart here is built from the federal collections above by a script in the repository, and each writes a JSON file anyone can download and check:
/data/concentration_cases.json — the 49 sustained
majority-non-resident departments, with year-by-year detail, the 4-digit
programme breakdown and a reporting-consistency check
(scripts/build_concentration_cases.py)./data/distribution.json — national department
histogram, 2010/2015/2020/2025
(scripts/build_distribution.py)./data/programme_shares.json — national share by
programme, 108 programmes with 300+ doctorates
(scripts/build_programme_shares.py)./data/university_distribution.json — per-university
department spread, 103 institutions
(scripts/build_university_distribution.py)./data/support_by_citizenship.json — SED Table 4-1
across 19 fields
(scripts/build_support_by_citizenship.py)./data/long_series.json — core-STEM doctorates
1995–2025 (scripts/build_long_series.py)./data/taxpayer_cost.json — federal obligations per
STEM doctorate (scripts/build_taxpayer_cost.py).
IPEDS split award level 9 (“doctor’s degree”) into 17
(research/scholarship) and 18 (professional practice) in 2010, but
institutions migrated one at a time across 2008 and 2009. In those years
each institution reports one coding or the other — verified:
495 institutions on level 9 and 122 on level 17 in 2008; 350 and 242 in
2009; zero report both. A complete figure therefore
requires summing both bases. Using one alone drops whoever
had already migrated, understating 2008 and 2009 by roughly a fifth and
putting a false trough exactly where the reporting changed. Every series
here that crosses the transition uses
pipeline.models.bases_for().
The SED records citizenship and how a doctorate was funded, but not whether a given assistantship came from a federal grant or from institutional money. The GSS records the federal share of assistantships but not citizenship. The two cannot be multiplied. Combining them to produce a count of federally funded international students would assume the federal share is identical for both groups, which nothing measures. This site does not make that calculation, and any figure of that kind found elsewhere was invented.
Institution-reported data on enrollment, financial support, and demographics of graduate students in science and engineering departments. Provides field-level enrollment counts with foreign/domestic separation at each institution — more granular than IPEDS enrollment for S&E fields.
The Survey of Doctorate Recipients (SDR) surveys ~80,000 PhD holders in the U.S. workforce. The National Survey of College Graduates (NSCG) surveys ~95,000 college graduates. These are the only federal surveys that separate native-born citizens from naturalized citizens and permanent residents — a distinction that IPEDS and the SED cannot make. The NSCG also records specific birth country (via the BTHST_TOGA variable), enabling China-specific analysis.
Reports total R&D expenditures by institution and funding source (federal, state, industry, etc.). Used to show how much research funding flows to universities with the highest NRA concentrations.
Student and Exchange Visitor Information System data tracking all F-1 and M-1 visa holders. Provides current counts by country of citizenship, education level, and U.S. state. Also includes STEM-specific breakdowns at the state level. As of March 2026, China has 229,463 active student records.
H-1B LCA data from the Department of Labor covers every Labor Condition Application for H-1B visas, including employer, job title, wage, and work location. Used to understand the post-graduation employment pipeline. USAspending federal grant data shows total federal awards to each university.
DOL OFLC data · USAspending.gov
USAspending figures exist for only 26 of the 106 universities — the extract is a top-recipients pull, not a full one — so absence there is not a low total.
Grant-level obligations to each university, from the two largest federal science funders. NSF Award Search gives award count and dollars obligated; NIH RePORTER gives grant count, dollars and the awarding department. Both are joined to the 106 universities tracked before the August 2026 additions. Across those 101, 2020–2024 obligations total $17.9B from NSF and $64.9B from NIH; 2025 adds $2.10B and $13.72B respectively.
These are obligations recorded against an institution name, so the same multi-campus attribution caveat applies as for H-1B filings. The four universities added in August 2026 are not yet joined (NSF covers CU Boulder; NIH carries no dollars for CU Boulder or Montana State).
Table 17 of the Federal Support Survey reports obligations to each institution from all federal agencies, not only NSF and NIH, split into R&D, R&D plant, facilities and equipment, fellowships/traineeships/training grants, and other general support. 106 universities on this site match a row. Their FY2023 obligations total $22.8B, of which $0.89B is S&E fellowships, traineeships and training grants — the line item that pays graduate students directly.
One record is a system office rather than a campus: the table has no Baton Rouge row for LSU, only "Louisiana State U., system office." Read LSU's figure as the system's.
State appropriations, net tuition revenue and FTE enrolment, from the SHEF FY25 report data. Used to show the state-taxpayer side of university funding alongside the federal side.
This is a state figure, not an institutional one. SHEEO does not publish institution-level appropriations. Each public university is joined to its state's SHEF series as a proxy. Never read it as "this university received $X"; read it as "this state appropriated $X across its public system."
Different federal surveys define citizenship groups differently. Understanding these definitions is critical to interpreting the data correctly.
The IPEDS Blind Spot: IPEDS — the most widely-cited source for higher education demographics — cannot distinguish between U.S. citizens and permanent residents (green card holders). Both are grouped into a single "U.S. citizen or permanent resident" bucket, then broken down by race/ethnicity. Only "Non-Resident Alien" (temporary visa holders) is reported separately.
This means that when IPEDS reports 60% of PhDs go to "domestic" students, that 60% includes an unknown number of green card holders and naturalized citizens who were born abroad.
| Survey | Categories | Can separate citizens from green card holders? |
|---|---|---|
| IPEDS (all surveys) |
NRA = temporary visa only. Race/ethnicity categories (Asian, White, Black, Hispanic, etc.) include BOTH citizens AND permanent residents combined. |
No |
| NSF SED | "U.S. citizen or permanent resident" vs. "Temporary visa holder." Same grouping as IPEDS. | No |
| NSF SDR |
CTZN field has 5 values: 1 = Native-born U.S. citizen 2 = Naturalized U.S. citizen (foreign-born) 3 = Permanent resident (green card) 4 = Temporary visa holder 5 = Living outside the U.S. |
Yes |
| NSF NSCG |
Same CTZN coding as SDR, plus: BTHST_TOGA = specific birth country code BTHRGN = birth region CTZDUAL = dual citizenship flag |
Yes, plus birth country |
| SEVIS / SEVP | Country of citizenship for each active F-1/M-1 record. | N/A (visa holders only) |
The SDR finding: 28% of "domestic" PhDs are foreign-born.
The 2023 Survey of Doctorate Recipients (80,143 respondents) reveals the composition of what IPEDS calls "U.S. citizen or permanent resident":
That means roughly 28% of the "domestic" PhD workforce that IPEDS counts as American (the 17.9% naturalized + 7.0% permanent resident + additional foreign-born among the native-citizen category) are actually foreign-born. Only about two-thirds of PhD holders working in the U.S. were born here.
The name-list pages on this site are built from dissertation and thesis metadata harvested directly from university digital repositories. No individual student records are accessed — only publicly available metadata from institutional repositories.
| Method | Description | Universities |
|---|---|---|
| OAI-PMH | Open Archives Initiative Protocol for Metadata Harvesting. A standard protocol supported by most DSpace, Fedora, and bepress repositories. We send ListRecords requests to each university's OAI endpoint and collect Dublin Core metadata (title, creator, date, subject, description, type). | Most of the 105 (Arizona, BYU, Cornell, CMU, Columbia, Portland State, etc.) |
| REST API | Stanford Digital Repository exposes a Searchworks/PURL API. We query for genre:Thesis records and collect structured metadata including department and advisor. | Stanford |
| eScholarship API | The University of California system's eScholarship platform supports OAI-PMH with campus-specific sets. We harvest all 10 UC campuses via their OAI endpoint. | UC Berkeley, UCLA, UCSD, UC Davis, UC Irvine, UCSB, UCSC, UCR, UC Merced, UCSF |
| HTML scraping | For repositories without API access, structured HTML parsing of thesis listing pages. | Caltech, UW Seattle |
After harvesting, deduplication and field filtering, the processed corpus holds 517,386 thesis and dissertation records across 106 universities:
| Split | Records | Notes |
|---|---|---|
| STEM | 369,319 | Drives the STEM pages and department reports |
| Non-STEM | 148,067 | Processed separately, same pipeline |
| Labelled PhD | 312,248 | Both splits combined |
| Labelled Master's | 145,379 | Both splits combined |
| Degree label inferred | 88.4% | Ranges from 24.8% to 100% by university |
Every record counted here is a real thesis or dissertation. Where a degree label could not be inferred, the record is still counted in the total and still name-classified — only the PhD/Master's split is unknown for it. So a PhD count from a university with a low labelling rate must be quoted with that rate attached, and a PhD share must never be computed by dividing labelled PhDs by the total.
Author names from dissertation metadata are classified by likely origin: a surname (and, for ambiguous surnames, a given name) is read against curated surname lists and returns one of chinese, korean, vietnamese, indian, iranian, turkish, arabic, excluded_ambiguous or other, with a confidence. The classifier has no citizenship, visa or nationality input, and no such output.
Name-origin classification is not a measure of visa status, and it is not IPEDS NRA. A Chinese-origin name does not make someone a Chinese citizen, an international student or a visa holder. They may be a naturalised U.S. citizen, a permanent resident, born here, or a citizen of a third country.
Rather than assert this, we measured it. A calibration study
(scripts/name_signal_study.py) asked one question: does the distribution of
name features across a cohort predict that cohort's reported NRA share, as the
institution itself filed it to IPEDS? Cohorts are university × CIP 2-digit family cells
(1,153 of them, pooled over 2010–2025), and models were held out by whole university
across 5 folds, so a model is always scored on institutions it never saw.
| Model | Out-of-sample R² | RMSE |
|---|---|---|
| Field of study alone (CIP-family fixed effects, no names) | 0.639 | 0.131 |
| Field + all name features | 0.698 | 0.120 |
| Field + the institution's own published NRA share, no names | 0.713 | 0.117 |
| Field + institution NRA share + all name features | 0.739 | 0.112 |
The two increments are the whole result:
Two honest points against a purely negative reading. The raw association is strong: the share of a cohort with any non-Anglo-origin name correlates with reported NRA share at r = 0.729 (95% CI 0.677–0.772, bootstrapped by resampling whole universities). And it does not vanish under controls: absorbing university, field and year fixed effects leaves a partial r = 0.382. So the correct statement is not "there is no correlation." It is that essentially all of the apparent signal is field composition — engineering and computing cohorts have both more non-Anglo names and more visa holders — and what survives is most parsimoniously explained by finer-grained field composition than the 2-digit CIP controls could absorb.
The error is the disqualifying part. The best model is off by about 11 percentage points RMSE on a cell's NRA share, on a quantity whose spread is roughly 22 points. A department estimated at 40% is routinely 29% or 51% on the reported measure — and RMSE is a typical error, not a bound.
It is also weakest exactly where this project's interest is strongest. By field, correlation with reported NRA share runs:
| Field (CIP 2-digit) | Cells | Mean NRA | r |
|---|---|---|---|
| Psychology (42) | 80 | 9.5% | 0.711 |
| Math & Statistics (27) | 84 | 47.4% | 0.593 |
| Education (13) | 75 | 10.4% | 0.583 |
| Physical Sciences (40) | 91 | 42.8% | 0.524 |
| Computer Science (11) | 84 | 60.9% | 0.422 |
| Engineering (14) | 90 | 57.2% | 0.311 |
Engineering — the field this project cares most about — has the weakest correlation of any large STEM family, because NRA share is high nearly everywhere and the range is compressed. The method looks best in Psychology and Education, fields where almost nobody holds a visa and the correct prediction is trivially "low."
Why name classification is not on the public pages. The measurement above is the reason. Name composition predicts NRA share about as well as knowing which department it is; once you know the department, it adds almost nothing; the residual error is too large for any department-level claim; and IPEDS already publishes the target at institution × CIP × year. The one legitimate use of the study is as a negative control — it quantifies that name classification, used in aggregate and at its best, reproduces a reported NRA share only to within about 11 points.
This is the most consequential limitation on the page, and it affects anything with a year axis.
The OAI harvester takes the first dc:date element in a record. The
oai_dc format flattens dc.date.accessioned,
dc.date.available and dc.date.issued into repeated
dc:date elements, and DSpace emits the accession timestamp first. So for many
repositories the stored year is when the repository ingested the file, not when
the degree was awarded.
Measured, not assumed: 407,462 of 1,035,726 in-window rows (39.3%) carry a date with a clock time, which only an ingest timestamp has, and 72 of the 106 universities have more than 90% of their in-window rows timestamped that way. Iowa State and MIT are effectively 100%: an Iowa State record stores 2018-08-22T18:45:17 for a 1934 degree, and an MIT record stores 2005-08-18T12:00:00Z for a 1996 degree.
Do not read a record-year distribution as a graduation-year distribution, and do not build a trend over time out of these years. Per-university totals and name-origin compositions are unaffected. Counts by year are safe only where the repository publishes a real issued date. Two universities are known-bad in a specific way: Northwestern's years are deposit timestamps from a bulk load, and UVA's post-2025 records carry a deposit-workflow timestamp with no degree date available upstream at all.
Two figures published on this site were wrong. Both have been fixed, and both are recorded here rather than quietly replaced, because a reader who saw the old numbers deserves to know which ones moved and why.
The employer matcher tested entity-map patterns as bare substrings. The entry
MIT therefore matched every employer name containing those three letters —
WIPRO LIMITED, INFOSYS LIMITED, TATA CONSULTANCY SERVICES
LIMITED among them — and credited their filings to MIT.
The scale of it was the tell: MIT's LCA total read 565,883 against a corpus median near 1,200, and its top job titles were "Technology Lead - Us", "Developer" and "Consultant - Us". Matching now requires word boundaries, and after a re-parse MIT's total reads 3,487 — in line with comparable private research universities.
Only MIT was affected. It was the only one of the 296 entity patterns shorter than six characters, and every other university's totals are unchanged, the highest now being Penn at 5,935. Any MIT LCA figure still reading near 565,883 — job titles, wage percentiles or yearly counts — predates this fix and is not MIT's.
Every chart, university, department, year and data panel on this site has its own address. The # control that appears beside a heading copies a link straight to it, and the scheme is simple enough to type by hand.
A link is a page, then a fragment naming the thing you want. Levels are separated by a double hyphen, narrowing from left to right:
| Fragment | What it addresses |
|---|---|
stem.html#missouri_st | a university |
stem.html#missouri_st--chemistry | a department within it |
stem.html#missouri_st--chemistry--2024 | one year of that department |
stem.html#missouri_st--2024 | that year's master's theses |
stem.html#missouri_st--funding | the funding panel, opened |
stem.html#missouri_st--enrollment | the enrollment panel, opened |
stem.html#missouri_st--h1b | the H-1B and OPT panel, opened |
stem.html#nat-visas | a national chart |
china.html#return-chart-section | a named section on any other page |
The university token is the same one used
throughout the site: in the sidebar links, in
/universities/<key>.html, and in
/data/stem/<key>.json. The department token is the
department name in lower case with every run of non-alphanumeric characters
replaced by a single hyphen — Computer Science becomes
computer-science. Section names on the other pages are the
heading text under the same rule.
Resolution is deliberately forgiving: capitalisation never matters, and
neither does the difference between - and _ in a
university key, so #missouri-st--chemistry and
#missouri_st--chemistry reach the same place. Older links keep
working.
Which data panels are open travels after the anchor, so a link can reproduce a whole view rather than just a scroll position:
stem.html#missouri_st--chemistry?funding=1&h1b=1&ms=1
The flags are funding, enrollment,
h1b (H-1B and OPT) and ms (show master's theses).
Each is =1 to open. Naming a panel directly — for example
#missouri_st--h1b — opens it without needing the flag.
immigration.html#mit lands on that university's card.
department_reports.html#missouri_st--trend-chart selects a
university and a tab. query.html#stem_records loads the SQL
workspace with that table selected.
ipeds_dashboard.html#panel-states opens a dashboard tab.
Machine-readable equivalents are listed in
catalog.json and
llms.txt.
| Source | Agency | URL | Data Updated |
|---|---|---|---|
| IPEDS Data Center | NCES / Dept. of Education | nces.ed.gov | 2024 academic year |
| Survey of Earned Doctorates | NSF / NCSES | ncses.nsf.gov | 2024 survey year |
| Graduate Student Survey | NSF / NCSES | ncses.nsf.gov | 2024 survey year |
| Survey of Doctorate Recipients | NSF / NCSES | ncses.nsf.gov | 2023 survey |
| National Survey of College Graduates | NSF / NCSES | ncses.nsf.gov | 2023 survey |
| HERD Survey | NSF / NCSES | ncses.nsf.gov | FY 2024 |
| SEVIS Data Mapping Tool | DHS / ICE / SEVP | studyinthestates.dhs.gov | March 2026 |
| H-1B LCA Disclosure | DOL / OFLC | dol.gov | FY2026 Q3 (through June 2026) |
| USAspending | U.S. Treasury | usaspending.gov | 2024 |
| NSF Award Search | NSF | nsf.gov/awardsearch | 2020–2025 |
| NIH RePORTER | NIH | reporter.nih.gov | 2020–2025 |
| Federal Support Survey, Table 17 | NSF / NCSES | NSF 25-339 | FY 2023 |
| SHEEO SHEF | SHEEO (non-federal) | shef.sheeo.org | FY 2025 |
| American Community Survey | U.S. Census Bureau | data.census.gov | 2022 ACS |
| S&E Indicators (NSB) | NSF / NSB | ncses.nsf.gov/indicators | 2024 and 2026 editions |
This site uses a lightweight, self-hosted analytics system. No third-party tracking (no Google Analytics, no Facebook pixel, no ad networks).
| Data Point | Purpose |
|---|---|
| Page views | Understand which universities and pages are most viewed |
| Click events | Track which departments and toggles users interact with |
| Share link clicks | Measure how often content is shared (# button clicks) |
| Country & city | Geographic distribution of visitors (from CloudFront headers) |
| IP address | Geo-analysis and unique visitor estimation |
| Screen size & user agent | Device type analysis |
Events are sent to analytics.andy-barr.com/collect via the Beacon API.
Data is stored in AWS DynamoDB with a 90-day TTL (auto-deleted after 90 days).
The analytics dashboard is at analytics.andy-barr.com (password-protected).
When you click any # copy-link control, we record which anchor was shared (e.g., "purdue--computer-science--2024") so we can understand which content is most shared. The URL you copy goes to your clipboard — we don't track where you paste it.
Page last updated: August 18, 2026. Data freshness varies by source; see individual entries above for the most recent data year available in each dataset.