Data Sources & Methodology

Where our data comes from, how it was collected, and what it can and cannot tell you

Contents

  1. Federal Datasets
  2. Understanding Citizenship Categories
  3. Dissertation Metadata
  4. Name Classification
  5. Limitations & Caveats
  6. Corrections
  7. Linking to This Data
  8. Links to Original Sources
  9. Analytics & Tracking

Immigration via Education Pipeline

This site tracks the full pathway from international graduate enrollment to permanent residency, using exclusively official US government data:

PhD/MS Production
University repositories
OPT Work Auth
ICE/SEVP SEVIS
H-1B Visa
USCIS + DOL LCA
Green Card
DOL PERM
Stage Source Agency Coverage Universities
Thesis/Dissertation Data University repositories (OAI-PMH) 2010-2026 105
OPT Students Employed ICE/SEVP SEVIS Reports CY 2017-2024 59 of 105 (Top 100 list)
H-1B Petition Approvals USCIS H-1B Employer Data Hub FY 2010-2026 101
H-1B Job Titles & Salaries DOL LCA Disclosure Data CY 2014-2026 (both ends part-years) 100
Green Card Sponsorship DOL PERM Disclosure Data FY 2024 71

Note on employer entity names: Universities file H-1B and PERM petitions under their legal entity name, which may differ from their common name and can change over time. Multi-campus university systems (e.g., University of Missouri, SUNY, University of California) often file under a single systemwide entity, making per-campus attribution imprecise. We match employer names to universities using pattern matching with known name variants — 296 patterns, tested on word boundaries, not as bare substrings (see Corrections) — but some filings may be attributed to the wrong campus within a system.

1. Federal Datasets

This site draws on ten major data sources, each with different strengths and coverage. Nine are federal; the last, SHEEO SHEF, is the state-finance series used for state appropriations.

IPEDS (Integrated Postsecondary Education Data System)

Agency: NCES / U.S. Dept. of Education Coverage: 1984 – 2024

IPEDS is the primary federal database for higher education statistics. Every Title IV institution reports annually. We use four IPEDS surveys:

IPEDS Data Center

NSF Survey of Earned Doctorates (SED)

Agency: NSF / NCSES Coverage: 1979 – 2024

An annual census of every research doctorate recipient from a U.S. institution. Unlike IPEDS, the SED collects individual-level data including source country for temporary visa holders, postgraduation plans, and financial support. The SED is the source for our national doctorate trend charts, stay-rate analysis, and country-of-origin rankings (China, India, etc.). Table 4-1 — primary source of financial support, by citizenship — is the basis for the funding analysis on this site. It is the only federal collection that crosses how a doctorate was paid for with the recipient’s citizenship status. Because the SED is a census administered by institutions and completed before graduation, its coverage is effectively complete rather than a sample.

NSF SED homepage · 2024 tables

Derived datasets published on this site

Every chart here is built from the federal collections above by a script in the repository, and each writes a JSON file anyone can download and check:

The 2008–09 award-level transition

IPEDS split award level 9 (“doctor’s degree”) into 17 (research/scholarship) and 18 (professional practice) in 2010, but institutions migrated one at a time across 2008 and 2009. In those years each institution reports one coding or the other — verified: 495 institutions on level 9 and 122 on level 17 in 2008; 350 and 242 in 2009; zero report both. A complete figure therefore requires summing both bases. Using one alone drops whoever had already migrated, understating 2008 and 2009 by roughly a fifth and putting a false trough exactly where the reporting changed. Every series here that crosses the transition uses pipeline.models.bases_for().

What the funding data cannot be made to say

The SED records citizenship and how a doctorate was funded, but not whether a given assistantship came from a federal grant or from institutional money. The GSS records the federal share of assistantships but not citizenship. The two cannot be multiplied. Combining them to produce a count of federally funded international students would assume the federal share is identical for both groups, which nothing measures. This site does not make that calculation, and any figure of that kind found elsewhere was invented.

NSF Graduate Student Survey (GSS)

Agency: NSF / NCSES Coverage: 2023 – 2024

Institution-reported data on enrollment, financial support, and demographics of graduate students in science and engineering departments. Provides field-level enrollment counts with foreign/domestic separation at each institution — more granular than IPEDS enrollment for S&E fields.

NSF GSS homepage

NSF Workforce Surveys (SDR & NSCG)

Agency: NSF / NCSES Coverage: 2021 – 2023

The Survey of Doctorate Recipients (SDR) surveys ~80,000 PhD holders in the U.S. workforce. The National Survey of College Graduates (NSCG) surveys ~95,000 college graduates. These are the only federal surveys that separate native-born citizens from naturalized citizens and permanent residents — a distinction that IPEDS and the SED cannot make. The NSCG also records specific birth country (via the BTHST_TOGA variable), enabling China-specific analysis.

SDR microdata · NSCG microdata

NSF Higher Education R&D Survey (HERD)

Agency: NSF / NCSES Coverage: FY 2024

Reports total R&D expenditures by institution and funding source (federal, state, industry, etc.). Used to show how much research funding flows to universities with the highest NRA concentrations.

NSF HERD homepage

ICE SEVIS Data (SEVP)

Agency: DHS / ICE / SEVP Coverage: Jan 2024 – Mar 2026 (monthly snapshots)

Student and Exchange Visitor Information System data tracking all F-1 and M-1 visa holders. Provides current counts by country of citizenship, education level, and U.S. state. Also includes STEM-specific breakdowns at the state level. As of March 2026, China has 229,463 active student records.

SEVIS Data Mapping Tool

DOL H-1B Disclosure & USAspending

Agencies: DOL / Treasury Coverage: FY2024 – FY2025

H-1B LCA data from the Department of Labor covers every Labor Condition Application for H-1B visas, including employer, job title, wage, and work location. Used to understand the post-graduation employment pipeline. USAspending federal grant data shows total federal awards to each university.

DOL OFLC data · USAspending.gov

USAspending figures exist for only 26 of the 106 universities — the extract is a top-recipients pull, not a full one — so absence there is not a low total.

NSF Awards & NIH RePORTER

Agencies: NSF / NIH Coverage: 2020 – 2025

Grant-level obligations to each university, from the two largest federal science funders. NSF Award Search gives award count and dollars obligated; NIH RePORTER gives grant count, dollars and the awarding department. Both are joined to the 106 universities tracked before the August 2026 additions. Across those 101, 2020–2024 obligations total $17.9B from NSF and $64.9B from NIH; 2025 adds $2.10B and $13.72B respectively.

These are obligations recorded against an institution name, so the same multi-campus attribution caveat applies as for H-1B filings. The four universities added in August 2026 are not yet joined (NSF covers CU Boulder; NIH carries no dollars for CU Boulder or Montana State).

NSF Award Search · NIH RePORTER

NCSES Federal Support Survey (NSF 25-339)

Agency: NSF / NCSES Coverage: FY 2023

Table 17 of the Federal Support Survey reports obligations to each institution from all federal agencies, not only NSF and NIH, split into R&D, R&D plant, facilities and equipment, fellowships/traineeships/training grants, and other general support. 106 universities on this site match a row. Their FY2023 obligations total $22.8B, of which $0.89B is S&E fellowships, traineeships and training grants — the line item that pays graduate students directly.

One record is a system office rather than a campus: the table has no Baton Rouge row for LSU, only "Louisiana State U., system office." Read LSU's figure as the system's.

NSF 25-339

SHEEO SHEF (State Higher Education Finance)

Agency: SHEEO (non-federal) Coverage: FY 2010 – FY 2025

State appropriations, net tuition revenue and FTE enrolment, from the SHEF FY25 report data. Used to show the state-taxpayer side of university funding alongside the federal side.

This is a state figure, not an institutional one. SHEEO does not publish institution-level appropriations. Each public university is joined to its state's SHEF series as a proxy. Never read it as "this university received $X"; read it as "this state appropriated $X across its public system."

SHEEO SHEF

2. Understanding Citizenship Categories

Different federal surveys define citizenship groups differently. Understanding these definitions is critical to interpreting the data correctly.

The IPEDS Blind Spot: IPEDS — the most widely-cited source for higher education demographics — cannot distinguish between U.S. citizens and permanent residents (green card holders). Both are grouped into a single "U.S. citizen or permanent resident" bucket, then broken down by race/ethnicity. Only "Non-Resident Alien" (temporary visa holders) is reported separately.

This means that when IPEDS reports 60% of PhDs go to "domestic" students, that 60% includes an unknown number of green card holders and naturalized citizens who were born abroad.

How each survey defines groups

Survey Categories Can separate citizens from green card holders?
IPEDS (all surveys) NRA = temporary visa only.
Race/ethnicity categories (Asian, White, Black, Hispanic, etc.) include BOTH citizens AND permanent residents combined.
No
NSF SED "U.S. citizen or permanent resident" vs. "Temporary visa holder." Same grouping as IPEDS. No
NSF SDR CTZN field has 5 values:
1 = Native-born U.S. citizen
2 = Naturalized U.S. citizen (foreign-born)
3 = Permanent resident (green card)
4 = Temporary visa holder
5 = Living outside the U.S.
Yes
NSF NSCG Same CTZN coding as SDR, plus:
BTHST_TOGA = specific birth country code
BTHRGN = birth region
CTZDUAL = dual citizenship flag
Yes, plus birth country
SEVIS / SEVP Country of citizenship for each active F-1/M-1 record. N/A (visa holders only)

The SDR finding: 28% of "domestic" PhDs are foreign-born.

The 2023 Survey of Doctorate Recipients (80,143 respondents) reveals the composition of what IPEDS calls "U.S. citizen or permanent resident":

That means roughly 28% of the "domestic" PhD workforce that IPEDS counts as American (the 17.9% naturalized + 7.0% permanent resident + additional foreign-born among the native-citizen category) are actually foreign-born. Only about two-thirds of PhD holders working in the U.S. were born here.

Key terms

3. Dissertation Metadata Collection

The name-list pages on this site are built from dissertation and thesis metadata harvested directly from university digital repositories. No individual student records are accessed — only publicly available metadata from institutional repositories.

Collection methods

Method Description Universities
OAI-PMH Open Archives Initiative Protocol for Metadata Harvesting. A standard protocol supported by most DSpace, Fedora, and bepress repositories. We send ListRecords requests to each university's OAI endpoint and collect Dublin Core metadata (title, creator, date, subject, description, type). Most of the 105 (Arizona, BYU, Cornell, CMU, Columbia, Portland State, etc.)
REST API Stanford Digital Repository exposes a Searchworks/PURL API. We query for genre:Thesis records and collect structured metadata including department and advisor. Stanford
eScholarship API The University of California system's eScholarship platform supports OAI-PMH with campus-specific sets. We harvest all 10 UC campuses via their OAI endpoint. UC Berkeley, UCLA, UCSD, UC Davis, UC Irvine, UCSB, UCSC, UCR, UC Merced, UCSF
HTML scraping For repositories without API access, structured HTML parsing of thesis listing pages. Caltech, UW Seattle

What the corpus contains

After harvesting, deduplication and field filtering, the processed corpus holds 517,386 thesis and dissertation records across 106 universities:

SplitRecordsNotes
STEM369,319Drives the STEM pages and department reports
Non-STEM148,067Processed separately, same pipeline
Labelled PhD312,248Both splits combined
Labelled Master's145,379Both splits combined
Degree label inferred88.4%Ranges from 24.8% to 100% by university

Every record counted here is a real thesis or dissertation. Where a degree label could not be inferred, the record is still counted in the total and still name-classified — only the PhD/Master's split is unknown for it. So a PhD count from a university with a low labelling rate must be quoted with that rate attached, and a PhD share must never be computed by dividing labelled PhDs by the total.

Metadata fields collected

4. Name Classification

Author names from dissertation metadata are classified by likely origin: a surname (and, for ambiguous surnames, a given name) is read against curated surname lists and returns one of chinese, korean, vietnamese, indian, iranian, turkish, arabic, excluded_ambiguous or other, with a confidence. The classifier has no citizenship, visa or nationality input, and no such output.

Name-origin classification is not a measure of visa status, and it is not IPEDS NRA. A Chinese-origin name does not make someone a Chinese citizen, an international student or a visa holder. They may be a naturalised U.S. citizen, a permanent resident, born here, or a citizen of a third country.

How much signal names actually carry

Rather than assert this, we measured it. A calibration study (scripts/name_signal_study.py) asked one question: does the distribution of name features across a cohort predict that cohort's reported NRA share, as the institution itself filed it to IPEDS? Cohorts are university × CIP 2-digit family cells (1,153 of them, pooled over 2010–2025), and models were held out by whole university across 5 folds, so a model is always scored on institutions it never saw.

ModelOut-of-sample R²RMSE
Field of study alone (CIP-family fixed effects, no names)0.6390.131
Field + all name features0.6980.120
Field + the institution's own published NRA share, no names0.7130.117
Field + institution NRA share + all name features0.7390.112

The two increments are the whole result:

Two honest points against a purely negative reading. The raw association is strong: the share of a cohort with any non-Anglo-origin name correlates with reported NRA share at r = 0.729 (95% CI 0.677–0.772, bootstrapped by resampling whole universities). And it does not vanish under controls: absorbing university, field and year fixed effects leaves a partial r = 0.382. So the correct statement is not "there is no correlation." It is that essentially all of the apparent signal is field composition — engineering and computing cohorts have both more non-Anglo names and more visa holders — and what survives is most parsimoniously explained by finer-grained field composition than the 2-digit CIP controls could absorb.

Where the signal fails

The error is the disqualifying part. The best model is off by about 11 percentage points RMSE on a cell's NRA share, on a quantity whose spread is roughly 22 points. A department estimated at 40% is routinely 29% or 51% on the reported measure — and RMSE is a typical error, not a bound.

It is also weakest exactly where this project's interest is strongest. By field, correlation with reported NRA share runs:

Field (CIP 2-digit)CellsMean NRAr
Psychology (42)809.5%0.711
Math & Statistics (27)8447.4%0.593
Education (13)7510.4%0.583
Physical Sciences (40)9142.8%0.524
Computer Science (11)8460.9%0.422
Engineering (14)9057.2%0.311

Engineering — the field this project cares most about — has the weakest correlation of any large STEM family, because NRA share is high nearly everywhere and the range is compressed. The method looks best in Psychology and Education, fields where almost nobody holds a visa and the correct prediction is trivially "low."

Why name classification is not on the public pages. The measurement above is the reason. Name composition predicts NRA share about as well as knowing which department it is; once you know the department, it adds almost nothing; the residual error is too large for any department-level claim; and IPEDS already publishes the target at institution × CIP × year. The one legitimate use of the study is as a negative control — it quantifies that name classification, used in aggregate and at its best, reproduces a reported NRA share only to within about 11 points.

5. Limitations & Caveats

Record years are often deposit dates, not degree dates

This is the most consequential limitation on the page, and it affects anything with a year axis. The OAI harvester takes the first dc:date element in a record. The oai_dc format flattens dc.date.accessioned, dc.date.available and dc.date.issued into repeated dc:date elements, and DSpace emits the accession timestamp first. So for many repositories the stored year is when the repository ingested the file, not when the degree was awarded.

Measured, not assumed: 407,462 of 1,035,726 in-window rows (39.3%) carry a date with a clock time, which only an ingest timestamp has, and 72 of the 106 universities have more than 90% of their in-window rows timestamped that way. Iowa State and MIT are effectively 100%: an Iowa State record stores 2018-08-22T18:45:17 for a 1934 degree, and an MIT record stores 2005-08-18T12:00:00Z for a 1996 degree.

Do not read a record-year distribution as a graduation-year distribution, and do not build a trend over time out of these years. Per-university totals and name-origin compositions are unaffected. Counts by year are safe only where the repository publishes a real issued date. Two universities are known-bad in a specific way: Northwestern's years are deposit timestamps from a bulk load, and UVA's post-2025 records carry a deposit-workflow timestamp with no degree date available upstream at all.

Federal data limitations

Dissertation & name classification limitations

6. Corrections

Two figures published on this site were wrong. Both have been fixed, and both are recorded here rather than quietly replaced, because a reader who saw the old numbers deserves to know which ones moved and why.

Headline NRA share, recomputed for every university (August 2026)

The per-university NRA percentage shown beside each institution had been a stored constant with no recorded basis. Recomputing it directly from IPEDS found 37 of 103 universities out by more than 5 percentage points, in both directions (21 too low, 16 too high) and by as much as 32. The worst case was Michigan State, published at 16.7% against an IPEDS figure of 49.0%. Louisiana Tech (22.5% → 54.3%), Washington State (18.0% → 47.2%) and Florida (18.4% → 45.2%) were next; Caltech moved the other way (64.9% → 41.2%).

The basis is now pinned, stored per university, and stated so it can be checked:

ComponentValue
SourceIPEDS Completions
FieldsCIP 2-digit families 11, 14, 15, 26, 27, 30, 40, 42
Years2015–2022 inclusive
Award level17 (research/scholarship doctorate)

Michigan State's 49.0% is 1,045 non-resident doctorates of 2,132 on that basis. The method was never wrong — MIT reproduced exactly and Missouri S&T to within 1.1 points — most of the stored values simply had not come from it. Each of the three components is argued in scripts/compute_nra_shares.py, including why the window ends at 2022 rather than 2024. Two universities (CUNY and Colorado School of Mines) do not resolve to a single IPEDS unit on this basis and keep their previous values.

H-1B/LCA employer matching (August 2026)

The employer matcher tested entity-map patterns as bare substrings. The entry MIT therefore matched every employer name containing those three letters — WIPRO LIMITED, INFOSYS LIMITED, TATA CONSULTANCY SERVICES LIMITED among them — and credited their filings to MIT.

The scale of it was the tell: MIT's LCA total read 565,883 against a corpus median near 1,200, and its top job titles were "Technology Lead - Us", "Developer" and "Consultant - Us". Matching now requires word boundaries, and after a re-parse MIT's total reads 3,487 — in line with comparable private research universities.

Only MIT was affected. It was the only one of the 296 entity patterns shorter than six characters, and every other university's totals are unchanged, the highest now being Penn at 5,935. Any MIT LCA figure still reading near 565,883 — job titles, wage percentiles or yearly counts — predates this fix and is not MIT's.

7. Linking to This Data

Every chart, university, department, year and data panel on this site has its own address. The # control that appears beside a heading copies a link straight to it, and the scheme is simple enough to type by hand.

Source Agency URL Data Updated
IPEDS Data Center NCES / Dept. of Education nces.ed.gov 2024 academic year
Survey of Earned Doctorates NSF / NCSES ncses.nsf.gov 2024 survey year
Graduate Student Survey NSF / NCSES ncses.nsf.gov 2024 survey year
Survey of Doctorate Recipients NSF / NCSES ncses.nsf.gov 2023 survey
National Survey of College Graduates NSF / NCSES ncses.nsf.gov 2023 survey
HERD Survey NSF / NCSES ncses.nsf.gov FY 2024
SEVIS Data Mapping Tool DHS / ICE / SEVP studyinthestates.dhs.gov March 2026
H-1B LCA Disclosure DOL / OFLC dol.gov FY2026 Q3 (through June 2026)
USAspending U.S. Treasury usaspending.gov 2024
NSF Award Search NSF nsf.gov/awardsearch 2020–2025
NIH RePORTER NIH reporter.nih.gov 2020–2025
Federal Support Survey, Table 17 NSF / NCSES NSF 25-339 FY 2023
SHEEO SHEF SHEEO (non-federal) shef.sheeo.org FY 2025
American Community Survey U.S. Census Bureau data.census.gov 2022 ACS
S&E Indicators (NSB) NSF / NSB ncses.nsf.gov/indicators 2024 and 2026 editions

9. Analytics & Tracking

Site Analytics

This site uses a lightweight, self-hosted analytics system. No third-party tracking (no Google Analytics, no Facebook pixel, no ad networks).

What we collect

Data PointPurpose
Page viewsUnderstand which universities and pages are most viewed
Click eventsTrack which departments and toggles users interact with
Share link clicksMeasure how often content is shared (# button clicks)
Country & cityGeographic distribution of visitors (from CloudFront headers)
IP addressGeo-analysis and unique visitor estimation
Screen size & user agentDevice type analysis

What we don't collect

Infrastructure

Events are sent to analytics.andy-barr.com/collect via the Beacon API. Data is stored in AWS DynamoDB with a 90-day TTL (auto-deleted after 90 days). The analytics dashboard is at analytics.andy-barr.com (password-protected).

Share tracking

When you click any # copy-link control, we record which anchor was shared (e.g., "purdue--computer-science--2024") so we can understand which content is most shared. The URL you copy goes to your clipboard — we don't track where you paste it.

Page last updated: August 18, 2026. Data freshness varies by source; see individual entries above for the most recent data year available in each dataset.