Can AI reproduce an empirical research paper? The American Economic Association’s Data Editor team, responsible for data and code checks, offers a concrete example in its public working materials. Its replication template contains instructions for AI-assisted code review, running replication packages, and preparing reports. The assigned work is specific: read the author’s materials, examine whether programs execute, connect outputs to the paper, and document the findings.
The template has a direct connection to the journal’s data and code checks. The AEA’s official FAQ links to it when explaining the checking process and states that code is run and outputs verified within reasonable time and computational limits. Instructions for AI assistance do not establish that every paper is already reproduced automatically. The report-finalization instructions, in particular, leave final approval to a person. AI can participate in the work while people remain responsible for reviewing its conclusions.
This example makes the question testable. Useful assistance means finding the correct computational entry point, establishing what actually ran, investigating discrepancies, and producing records another researcher can examine. Generating plausible code or announcing “reproduction successful” is not sufficient evidence for any of those tasks.
What Does It Mean to Reproduce a Paper?
An empirical paper combines a research question, data processing, statistical calculations, and interpretation. Using the author’s data and code to regenerate reported tables and figures is one task. Collecting new data to examine whether the conclusion holds is another. A National Academies report distinguishes computational reproducibility from replication using new data. A discussion of AI-assisted reproduction needs to establish which task is being attempted.
Execution and output checking in the AEA template primarily concern the first task. An author supplies materials, and a checker uses them to investigate whether the reported outputs can be obtained. Even if every coefficient agrees, that establishes a relationship between the supplied materials and the reported calculations. Whether variables measure the intended concepts, identifying assumptions are credible, or findings generalize requires additional reasoning. Computational reproduction provides a basis for that discussion without settling it.
A discrepancy also does not establish misconduct. Software versions, random computation, missing files, execution order, or undocumented processing can affect results. Some problems concern incomplete records; others require clarification or reveal substantive errors. The immediate task is to locate the discrepancy and understand its consequences. A judgment about the entire paper would exceed the evidence if its cause remains unknown.
Finding the Right Materials and Entry Point
The presence of code does not establish which program should run first. A project may contain cleaning scripts, exploratory analyses, final regressions, appendix checks, and plotting files. An AI assistant that chooses a file arbitrarily might produce an output unrelated to the final analysis. Entry points, execution order, and links between programs and exhibits give subsequent checks a defined object.
The AEA’s replication-run instructions call for reading the author’s documentation, inventorying code, and locating the author’s driver program. AI can help assemble information scattered across documentation and scripts, identify inputs and dependencies, and flag missing steps. Unspecified conditions should remain questions to resolve. A guessed sequence that happens to run does not establish that it matches the author’s procedure.
Data availability requires separate attention. Empirical work may depend on restricted microdata, commercial databases, or materials accessible only in an approved environment. A missing local file might reflect access restrictions or a partial copy of a larger deposit. The assistant needs to assess documentation, inventories, and authorized access before deciding what can be checked. Inaccessible materials should be recorded as a limitation. Substituting simulated data cannot reproduce results obtained from the original observations.
Establishing What Actually Ran
Execution brings the work closer to reproduction, but a finished process is a coarse indicator. An outer script can exit normally even though a statistical command failed inside it. A calculation might stop because of insufficient memory or a time limit. Files left in an output directory may belong to an earlier run. Exit status and file existence alone can therefore give a misleading impression of completion.
The AEA template specifically calls for inspecting logs, completion markers, and the commands actually executed. AI can compare those records with the expected sequence and organize failures by cause. Missing dependencies, inaccessible inputs, failed commands, and inadequate computing resources require different responses. A run needing more memory is a different finding from an incorrectly specified model. Reporting them separately makes further investigation possible.
Repairs also need boundaries. Replacing an absolute path from the author’s computer will often leave the statistical analysis unchanged. Dropping observations, changing standard errors, or substituting an estimator may alter the result. If an assistant makes such changes merely to obtain successful execution, the object being reproduced changes. Modifications need a record, and changes to samples, variables, or estimation require explicit consideration of their rationale. Keeping the original program alongside the modified version makes the resulting calculation traceable.
Comparing Outputs with the Paper
Suppose a paper reports a coefficient in a regression table and the regenerated estimate differs. An instruction to “make it match” invites the assistant to change variables, samples, or specifications until it approaches the target. That searches for a program capable of producing a desired number. It does not test whether the author’s materials support the published analysis.
A useful investigation first establishes whether both sides describe the same calculation. Which sample, controls, and fixed effects belong to that column? How are standard errors computed? Does the program generate the main table or an appendix specification? Similar coefficients accompanied by different standard errors suggest examining inference settings. A different observation count calls for tracing missing-value handling, merges, and filters. Each discrepancy should lead to a condition that can be inspected, rather than an explanation supplied by guesswork.
Differences in the last decimal place also need context. Deterministic calculations, simulation-based estimates, and iterative algorithms have different sensitivities to environments and random states. A comparison should state the relevant precision and conditions. Neither dismissing every small difference nor demanding identical digits from every algorithm in every environment is adequate. AI can organize discrepancies and locate relevant code; researchers must assess their significance for the argument.
| Evidence obtained | What it supports | What still needs checking |
|---|---|---|
| Documented entry point and input files located | The specified workflow can be attempted | Access, versions, and dependencies |
| Logs show the expected steps completed | Those calculations completed in this run | Whether fresh outputs correspond to the paper |
| A table’s sample sizes, coefficients, and standard errors match | That result was reproduced under the recorded conditions | Other tables, figures, and appendix results |
| A modified model produces similar numbers | The modified program produced those outputs | Whether the changes are justified and preserve the original analysis |
| Restricted data cannot be accessed | The check has a documented materials limitation | Whether verification can continue after authorization |
Reporting Findings Someone Else Can Check
A useful report needs more than a success or failure label. It should identify the materials and environment, the entry points executed, the exhibits compared, the differences found, and the work left incomplete. Findings should connect to logs, outputs, or code locations. Another researcher should be able to inspect the conclusions or continue from an unresolved issue.
For example, “Table 2 reproduced” requires an actual comparison. If only the generating program ran, but the table in the paper was unavailable, the supported finding is narrower. Similarly, an appendix figure left unverified could reflect a program error, missing materials, or insufficient time. These require different next steps: repair, additional access, or further execution. Preserving those distinctions is more informative than assigning one label to the entire package.
The AEA report task consolidates preliminary findings for review and preserves human approval. Researchers can use the same division of responsibility in their own work: AI assists with discovery, execution, and documentation; people review the evidence and decide how to address findings. The repository is primarily intended for the AEA Data Editor’s internal work and includes corresponding environment and process assumptions. Reading it offers methods to adapt, but downloading it does not supply the same data access or a complete editorial review capability.
Preparing a Package That Others Can Reproduce
For authors, the central requirement is to explain the computation without making another researcher guess. The AEA’s data and code policy requires applicable papers to provide supporting materials and documentation, with provisions for nonpublic data. Authors need to describe data provenance and access, the environment, execution, and the relationship between programs and reported results. The recommended social science data editors’ README template provides a useful starting point.
When a project already has a clear path from inputs to final outputs, AI can help check consistency between documentation and code, identify omitted dependencies, and organize exhibit mappings. If the project depends on manual steps remembered only by the author, or mixes final outputs with exploratory results, AI does not automatically recover that missing information. Clear documentation reduces the burden on collaborators and establishes conditions for meaningful automated checks.
A bounded reproduction request might ask AI to identify requirements from the documentation, execute the original programs within authorized access, compare specified exhibits, and record discrepancies and modifications. Unverifiable portions should be listed separately. This makes the deliverable concrete: completed calculations, unresolved findings, and unavailable materials remain distinguishable instead of being compressed into an ambiguous completion message.
AI’s contribution to paper reproduction can be assessed through these records. It may reduce the effort involved in organizing materials, executing programs, and comparing outputs. Whether the paper’s conclusions hold still requires relating computational evidence to the research question, methods, and assumptions. The AEA example offers a practical starting point: define reproducibility work as tasks supported by evidence, use AI to assist with execution, and preserve a researcher’s ability to review the result.