NTRN Project Data Sets

Other IPEDS-related data resource posts introduce new measures and methods of data preparation. This post shares prepared data sets—along with the STATA code that produced them—that incorporate these advancements. All files are available for download from Penn State Scholarsphere.

I hope to periodically update these data sets to reflect

  • Newly released IPEDS data

  • New measures that can be constructed from the data

  • Improvements to existing measures

The IPEDS survey has evolved over time, and upcoming changes may be substantial given recent policy and personnel shifts at the Department of Education.

Because researchers have different needs, I provide three distinct data sets. Below, you’ll find a description of each—followed by two brief requests for users of the data.

Data Set #1: Title IV Institution

Due to the unit of observation challenge, the structure of IPEDS data presents a challenge for analysis—especially when working with finance data. A relatively simple solution to this challenge is to restructure the data around Title IV institutions, the organizational units examined during accreditation reviews and Title IV funding agreements.

Creating a Title IV institution data set involves two steps

1.     Allocate dollars from administrative units

For IPEDS observations reported at the administrative unit level, allocate expenditures to each associated IPEDS observation.

  • I typically allocate using a weighted average: 75% enrollment and 25% expenditures.

  • A few administrative units with atypical contexts require special handling.

  • See this article for more details on the allocation method. 

2.     Collapse campus-level observations

Sum the data for all IPEDS observations associated with the same Title IV institution.

  • These can be linked using digits 2-6 of the Office of Postsecondary Education Identifier (OPEID).

  • This step—referred to as collapsing the data by Jaquette and Parra—ensures that each Title IV institution is represented by a single observation. 

Data Set #2: Title IV Institution Data Set (with some collapsing of TITle IV Institutions) 

Allocating all dollars from administrative units—as done in Data Set #1—allows all information to be reported at the level of a Title IV Institution, providing researchers with a consistent unit of observation. However, this approach assumes that system-level resources relate to enrollment and expenditure levels, which not be true in all contexts. When a large share of the system’s resources are reported within the administrative unit observation, the resulting allocation errors could be substantial.

The best way to avoid these potential errors is to combine all data associated with the administrative unit, including both the administrative unit observation and all IPEDS observations linked to it. This solution ensures that each observation contains accurate information, but it comes with a treadoff: some observations will include multiple Title IV institutions. This tradeoff may be minor for a community college district housing a few colleges, but much larger for systems housing many institutions.

Data Set #2 balances these considerations and collapses Title IV institutions when:

·      The administrative unit includes a small number of similar institutions, and

·      A substantial share of finances is reported at the administrative unit level.

In a journal article, I provided more detailed guidance on when to combine all data associated with an administrative unit. See this post for updated guidance that incorporates additional years of data as well as institutions residing in U.S. territories. This data set employs this guidance.

Data Set #3: IPEDS Observations Data Set

Researchers who want to add additional variables to the data may find the previous two data sets difficult to work with. That’s because some data sets housing other variables are structured similarly to IPEDS and designate institutions using UNITID—an identifier that distinguishes between IPEDS observations, not Title IV institutions.

 This UNITID-level data set is designed to support such merges. It retains the original IPEDS structure, making it easier to integrate additional variables. After merging, researchers can collapse the data to the Title IV Institution level using the provided STATA code. That code includes explanations for how to modify it to incorporate additional variables. The included explanations can also help researchers create similar routines using software other than STATA.

Note: This data set does not include administrative unit observations. The revenues and expenditures from those observations have already been allocated to the appropriate IPEDS observations using the same methods employed in the Title IV Institution data set. After allocation, the administrative unit records were removed. Since most external data sets do not include these administrative units, their absence should not complicate merges.

Supplementary Data SEts

The NTRN project will contain additional data sets that supplement Data Set #1. These data sets are structured as the level of a Title IV institution, allowing them to merged using the OPEID5 variable. At present, the NTRN project includes one supplementary data set: The FY 2023 MSI-Related Revenue Data Set. This data set contains information on federal spending for MSI-related programs during the 2023 fiscal year. It was developed in partnership with the Urban Institute to support the Institute’s report on Minority Serving Institutions.

Two Requests

1. Please cite these data if you use them in your work. This helps other researchers discover and benefit from these resources.

2. Reach out with feedback. If you have questions about the data sets or ideas for improving them, I’d be glad to hear from you.

Previous
Previous

Introduction to IPEDS

Next
Next

Measuring Revenues Using IPEDS