diff --git a/.codespell-ignore.txt b/.codespell-ignore.txt index cf794152..e5d0a245 100644 --- a/.codespell-ignore.txt +++ b/.codespell-ignore.txt @@ -6,3 +6,4 @@ RECOVR Goin ttest pre-selected +aer \ No newline at end of file diff --git a/data-collection/admin-data.qmd b/data-collection/admin-data.qmd index 68be925f..efcf5e63 100644 --- a/data-collection/admin-data.qmd +++ b/data-collection/admin-data.qmd @@ -35,71 +35,35 @@ license: url: https://creativecommons.org/licenses/by-sa/4.0/ --- -:::{.callout-tip appearance="simple"} - -## Key Takeaways - -* Administrative data refers to information collected, used, and stored primarily for operational purposes rather than research. -* While IPA projects usually rely on primary data, administrative data can also be a valuable input resource for research and impact evaluation. -* Accessing and using Administrative Data requires careful planning across four stages: information gathering, implementation, legal compliance, and data validation. -::: ## What is Administrative Data? -Administrative data is typically collected by government agencies and organizations for registration, transaction, and record-keeping purposes. Examples of administrative data can include credit card transactions, electronic medical records, insurance claims, educational records, arrest records, and mortality records. - -### Administrative Data vs. Survey Data +Administrative data refers to data collected for the administration of programs. It should be systematically collected, stored and used for program operation and management decisions. -Some challenges are common to both survey and administrative data. However, administrative data also comes with its own strengths and limitations. The table below compares the two approaches across key dimensions. +Administrative data is typically collected by government agencies and organizations for registration, transaction, and record-keeping purposes and it is designed to track a program's implementation -primarily the project's activities and expenses- it can also include indicators on program outcomes. Examples of administrative data can include call detail records (CDR), app usage logs, credit card transactions, client information from financial institutions, electronic medical records, insurance claims, hospital records of patient visits, educational records, arrest records, and mortality records. -| Dimension | Survey data | Administrative data | -|-----------|------------|-------------------| -| Recall and social desirability bias | High risk | Low risk | -| Attrition | Higher risk | Lower risk | -| Costs | Higher cost | Lower cost | -| Logistics | Easier to manage | More complex | -| Time to access | Faster access | Slower access | -| Documentation quality | Usually better documented | Often less documented | -| Control over study population | High control | Limited control | -| Control over data processing | High control | Limited control | -| Treatment affects measurement | Lower risk | Higher risk | -| Incentives to misreport | Lower risk | Higher risk | - -: {tbl-colwidths="[40,30,30]"} - -### Real-World Example: IPA's Experience with Administrative Data - -IPA's research on using administrative data for Monitoring and Evaluation (2016) highlights several key advantages: - -* **Cost-Effectiveness**: Administrative data reduces or eliminates the need for additional monitoring activities or surveys. -* **Timely Response**: Regular updates in management information systems enable faster analysis of key indicators. -* **Large Sample Size**: Coverage of entire beneficiary populations provides robust sample sizes. -* **Improved Accuracy**: Reduces social desirability bias and recall issues common in self-reported data. +:::{.callout-tip appearance="simple"} -This research emphasizes that data accuracy and reliability should take precedence over cost savings. Organizations must balance data quality, actionability, and resource allocation when incorporating administrative data into their research design. +## Key Takeaways - -

Unable to display PDF file. Download instead.

-
+* Administrative data refers to information collected, used, and stored primarily for operational purposes rather than research. +* While IPA projects usually rely on primary data, administrative data can also be a valuable input resource for research and impact evaluation. +* Accessing and using administrative data requires careful planning across four stages: information gathering, implementation, legal compliance, and data validation. +::: ## Standard Processes for Accessing Administrative Data Accessing administrative data involves four main steps: -::: {.callout-tip collapse="false"} - -## Finding Administrative Data +### 1. Finding Administrative Data Implementing partners and government agencies are valuable sources of administrative data. Common sources include: * **Health Data**: Regional/national health departments, hospitals, health insurance records * **Financial Data**: Banks, credit unions, credit reporting agencies * **Education Data**: Schools, ministries of education, standardized testing agencies -::: -::: {.callout-tip collapse="false"} - -## Formulating a Data Request +### 2. Preparing a Data Request When requesting administrative data, researchers should: @@ -107,29 +71,48 @@ When requesting administrative data, researchers should: * **List specific variables of interest**: Student ID, school level, Teacher ID, attendance, test scores * **Specify whether you need identified or de-identified data**: Request de-identified data when possible to reduce ethical complexity * **Avoid broad requests**: Instead of "all student data," specify exact variables, time frames needed, and frequency of updates -::: -::: {.callout-tip collapse="false"} +::: {.callout-tip collapse="true"} -## Implementing Data Flow Strategies +#### Data Use Agreements -A well-planned data flow strategy ensures smooth integration of administrative data: +A Data Use Agreement (DUA) is a legal document that outlines the terms under which a data provider shares data with a research institution. DUAs are also referred to as Data Sharing Agreements (DSAs), Memoranda of Understanding (MOUs), or Non-Disclosure Agreements (NDAs), depending on the context and the parties involved. -1. **Gather Identifying Information**: Determine what identifiers are available in the study sample such as national ID numbers, phone numbers, email addresses -2. **Link Datasets**: Use pre-existing identifiers to match study data with administrative data. For example, match student IDs with test records using national ID numbers -3. **Choose a Matching Strategy**: - * **Exact matching**: When identifiers are identical such as national ID numbers - * **Probabilistic matching**: When using combinations of name, date of birth, and location information -4. **Determine Who Performs the Link**: Clarify whether the data provider, researcher, or a third party will conduct the linkage and deindentification to maintain data security and privacy +A well-structured DUA protects both the data provider and the research team by clarifying responsibilities, limiting liability, and ensuring that data are used only for agreed-upon purposes. The table below describes the elements that a DUA should include. + +* **Project description**: Summary of the research purpose and scope +* **Authorized users and analysts**: Names or roles of individuals permitted to access the data +* **Data security procedures**: Encryption, storage, and access control requirements +* **Data to be shared**: Variables, time range, and format of the dataset +* **Timeframe**: Duration of the agreement and data access period +* **Data destruction**: Procedures for deleting or returning data at the end of the project +* **Publication review**: Whether the data provider has the right to review outputs before publication +* **Data publication**: Conditions under which derived data or results may be made public + +##### DUAs must be signed by an institutional representative + +A DUA should be executed by an authorized institutional representative, not by an individual researcher or staff member. This protects the researcher from personal liability and ensures that the institution assumes responsibility for compliance. + +##### Tips for negotiating a DUA + +Negotiating a DUA can be a lengthy process. The following practices help move negotiations forward: + +* **Understand legal constraints from the start**: Some provisions that appear as "requirements" from the data provider may be preferences that can be modified. Understanding what is legally mandated versus negotiable saves time. +* **Build the relationship early**: Establishing trust with the data provider before formal negotiations often makes the process smoother. +* **Frame the research in terms of the provider's mission**: If the research is relevant to the agency's own goals, emphasize that relevance to increase their willingness to share. +* **Consider using an intermediary**: In complex or sensitive cases, a trusted third party can facilitate negotiations between the research team and the data provider. ::: +### 3. Implementing Data Flow Strategies -## Data Flow Options +#### What is Data Flow? + +A data flow describes the path that data travels from the administrative source to the research team, specifying who collects identifying information, who conducts the match between study records and administrative records, and who receives the final de-identified dataset. The structure of the data flow determines which parties have access to personally identifiable information at each stage of the process. The structure of a data flow determines who has access to personally identifiable information (PII) and who conducts the match between study records and administrative data. Choosing the right option depends on the sensitivity of the data, the data provider's legal constraints, and the research team's security infrastructure. The table below summarizes the most common configurations. | Who conducts the match | Who has access to identified data | Notes | -|----------------------|----------------------------------|-------| +| ---------------------- | ---------------------------------- | ------- | | Researcher | Researcher receives identified data from agency | Simplest, but requires strong data security on the researcher's end | | Researcher, on-site at the agency | Researcher brings encrypted finder file; leaves with de-identified file | Useful when the agency restricts data from leaving its premises | | Researcher, on agency device | Researcher conducts match and analysis on an agency-monitored computer | Agency retains oversight of the process | @@ -140,58 +123,68 @@ The structure of a data flow determines who has access to personally identifiabl : {tbl-colwidths="[25,35,40]"} ::: {.callout-warning} -## Matching errors affect statistical power -Exact matching virtually eliminates false positives but increases false negatives, which attenuates impact estimates and reduces statistical power. Probabilistic matching reduces false negatives but introduces the risk of false positives, which can be especially harmful if match quality correlates with treatment or control status. Researchers should document their matching strategy and assess whether matching errors are likely to be unrelated to treatment assignment. +##### Matching errors affect statistical power + +Exact matching virtually eliminates false positives but increases false negatives, which attenuates impact estimates and reduces statistical power. Probabilistic matching reduces false negatives but introduces the risk of false positives, which can be especially harmful if match quality correlates with treatment or control status. Researchers should document their matching strategy and assess whether matching errors are likely to be unrelated to treatment assignment. ::: -### Match rates in practice +#### How to implement a Data Flow process? -Match rates vary considerably across studies depending on the quality of identifiers available, the data source, and the matching strategy used. The table below summarizes match rates from published studies that used administrative data linkage. +A well-planned data flow strategy ensures smooth integration of administrative data: -| Study | Records being matched | Match rate | -|-------|-----------------------|------------| -| Health Care Hotspotting | Hospital discharge records for enrolled participants | ~95% | -| Oregon Health Insurance Experiment (Finkelstein et al., 2012) | Credit reports for lottery participants | 68.5% | -| Effect of pre-trial detention on conviction (Dobbie, Goldin & Yang, 2018) | Tax data for defendants | 73–81% | -| Impact of kindergarten class on life outcomes (Chetty et al., 2011) | Parents' tax records for kindergarten students | 86% | +1. **Gather Identifying Information**: Determine what identifiers are available in the study sample such as national ID numbers, phone numbers, email addresses +2. **Link Datasets**: Use pre-existing identifiers to match study data with administrative data. For example, match student IDs with test records using national ID numbers +3. **Choose a Matching Strategy**: + * **Exact matching**: When identifiers are identical such as national ID numbers + * **Probabilistic matching**: When using combinations of name, date of birth, and location information +4. **Determine Who Performs the Link**: Clarify whether the data provider, researcher, or a third party will conduct the linkage and deindentification to maintain data security and privacy -: {tbl-colwidths="[40,35,25]"} +### 4. Data Validation -These benchmarks illustrate that even well-resourced studies rarely achieve perfect linkage. Researchers should account for expected match rates when designing studies and assessing statistical power. +Data validation is the process of verifying that administrative data received from a provider has the expected characteristics before it enters a project's analysis pipeline. Unlike data cleaning, which modifies records to correct errors, data validation checks whether the data as delivered matches what was agreed upon and what the analysis assumes and needs. -## Data Use Agreements +Administrative data validation focuses on two dimensions: -A Data Use Agreement (DUA) is a legal document that outlines the terms under which a data provider shares data with a research institution. DUAs are also referred to as Data Sharing Agreements (DSAs), Memoranda of Understanding (MOUs), or Non-Disclosure Agreements (NDAs), depending on the context and the parties involved. +* **Correctness**: The reported value reflects the true observed value, accounting for allowable measurement error. +* **Consistency**: Each record captures the same underlying construct across deliveries and across rows. -A well-structured DUA protects both the data provider and the research team by clarifying responsibilities, limiting liability, and ensuring that data are used only for agreed-upon purposes. The table below describes the elements that a DUA should include. +Data validation is different from data cleaning and from high-frequency checks used in survey data collection. The goal is to identify and triage errors in the data as received, not to alter the raw data. -* **Project description**: Summary of the research purpose and scope -* **Authorized users and analysts**: Names or roles of individuals permitted to access the data -* **Data security procedures**: Encryption, storage, and access control requirements -* **Data to be shared**: Variables, time range, and format of the dataset -* **Timeframe**: Duration of the agreement and data access period -* **Data destruction**: Procedures for deleting or returning data at the end of the project -* **Publication review**: Whether the data provider has the right to review outputs before publication -* **Data publication**: Conditions under which derived data or results may be made public +#### When to validate -::: {.callout-warning} -## DUAs must be signed by an institutional representative +:::{.callout-tip appearance="simple"} + +##### Validate as soon as data arrive -A DUA should be executed by an authorized institutional representative, not by an individual researcher or staff member. This protects the researcher from personal liability and ensures that the institution assumes responsibility for compliance. +The first validation should occur immediately upon receiving a new data delivery. Running checks promptly allows the research team to identify problems while there is still time to contact the data provider for corrections. Automating validation scripts reduces processing time and ensures checks are applied consistently across multiple deliveries. ::: -::: {.callout-tip collapse="true"} -## Tips for negotiating a DUA +:::{.callout-tip appearance="simple"} -Negotiating a DUA can be a lengthy process. The following practices help move negotiations forward: +##### Validate each new delivery on the raw file -* **Understand legal constraints from the start**: Some provisions that appear as "requirements" from the data provider may be preferences that can be modified. Understanding what is legally mandated versus negotiable saves time. -* **Build the relationship early**: Establishing trust with the data provider before formal negotiations often makes the process smoother. -* **Frame the research in terms of the provider's mission**: If the research is relevant to the agency's own goals, emphasize that relevance to increase their willingness to share. -* **Consider using an intermediary**: In complex or sensitive cases, a trusted third party can facilitate negotiations between the research team and the data provider. +Validation scripts should run on the most recent file as received, using the variable names and formats the provider uses. This approach ensures that any errors identified are attributable to the source data, not to transformations the research team has applied. Validate before converting files to project-specific formats (e.g., .dta, .RData, or a relational database). ::: +:::{.callout-tip appearance="simple"} + +##### Keep validation and cleaning as separate workflows + +Validation and cleaning serve different purposes and should be maintained as distinct scripts or processes. Validation documents what errors exist in the raw data. Cleaning modifies the data for analysis. Mixing the two makes it harder to trace the origin of any given value in the final analysis file and complicates future deliveries. +::: + +#### Common sources of error in administrative data + +| Source | Examples | +| -------- | --------- | +| Poorly designed forms or systems | No input constraints on data type or format; ambiguous questions that respondents interpret differently | +| Data entry errors | Misspellings, keystroke errors, misunderstanding of field definitions (e.g., annual income entered as monthly) | +| Misreporting or reporting bias | Incentives to over- or under-report; outcomes reported by program staff who have an interest in results | +| Data extraction errors | Incorrect queries, truncated exports, or format changes introduced during extraction | +| Data pre-preparation by the provider | Aggregations, recodes, or transformations applied before delivery that are not documented | + +: {tbl-colwidths="[30,70]"} ## Ethical Considerations for Using Administrative Data in RCTs @@ -303,10 +296,18 @@ Administrative datasets vary in cost, depending on: Administrative data is a valuable tool for randomized evaluations, offering cost-effective, accurate, and comprehensive insights. Researchers must navigate ethical, legal, and logistical challenges to ensure data quality and validity. By following standardized processes for information gathering, implementing appropriate data flow structures, executing well-designed data use agreements, and validating data at each delivery, research teams can produce reliable evidence from administrative sources. -## References +::: {.callout-tip collapse="true"} + +### Real-World Example: IPA's Experience with Administrative Data + +IPA's guide to Using Administrative Data for Monitoring and Evaluation (2016) highlights several key advantages: -Chetty, R., Friedman, J. N., Hilger, N., Saez, E., Schanzenbach, D. W., & Yagan, D. (2011). How does your kindergarten classroom affect your earnings? Evidence from Project STAR. *The Quarterly Journal of Economics*, 126(4), 1593–1660. https://doi.org/10.1093/qje/qjr041 +* **Cost-Effectiveness**: Administrative data reduces or eliminates the need for additional monitoring activities or surveys. +* **Timely Response**: Regular updates in management information systems enable faster analysis of key indicators. +* **Large Sample Size**: Coverage of entire beneficiary populations provides robust sample sizes. +* **Improved Accuracy**: Reduces social desirability bias and recall issues common in self-reported data. -Dobbie, W., Goldin, J., & Yang, C. S. (2018). The effects of pretrial detention on conviction, future crime, and employment: Evidence from randomly assigned judges. *American Economic Review*, 108(2), 201–240. https://doi.org/10.1257/aer.20161503 # codespell:ignore aer +This research emphasizes that data accuracy and reliability should take precedence over cost savings. Organizations must balance data quality, actionability, and resource allocation when incorporating administrative data into their research design. -Finkelstein, A., Taubman, S., Wright, B., Bernstein, M., Gruber, J., Newhouse, J. P., Allen, H., & Baicker, K. (2012). The Oregon Health Insurance Experiment: Evidence from the first year. *The Quarterly Journal of Economics*, 127(3), 1057–1106. https://doi.org/10.1093/qje/qjs020 +[Download: Goldilocks Deep Dive - Using Administrative Data for Monitoring and Evaluation](/assets/files/Goldilocks-Deep-Dive-Using-Administrative-Data-for-Monitoring-and-Evaluation.pdf) +:::