Marx Max中文
Menu

Right Code, Wrong Way

Abstract: AI can write code and estimate regressions while also changing samples, variables, and models in ways that affect conclusions. Sant’Anna’s recent account and two studies, The Agentic Garden of Forking Paths and Do Claude Code and Codex P-Hack?, make the question concrete: why can identical data yield different or opposing findings? This article examines how expectations can enter exploration and selection, distinguishes execution errors from defensible disagreements and outcome-driven choices, and uses a hypothetical model comparison to explain their implications. When searches are concealed and results are selected to match expectations, the risk can involve research misconduct rather than an execution error.

Sant’Anna recently wrote on LinkedIn that he was becoming tired of some uses of AI in research. He had enthusiastically used the tools, written a guide, and held workshops. In data analysis, however, he encountered sample restrictions, redefined variables, dropped controls, and omitted weights. Such changes can require careful attention to the process, while users may not realize that the output has been affected.

When we give AI a dataset, we usually expect faster computation. If it also changes the conditions of that computation, a further question arises: how much of the final result reflects relationships in the data, and how much reflects choices made along the way? Two studies of AI-assisted empirical analysis examine this through prior beliefs and task framing.

Why Can Identical Data Yield Different Analyses?

Data do not determine a unique analysis. Studying whether a policy increases income requires decisions about the measure, population, time window, and missing observations. Each decision enters the program and the result. Giving an assistant a research question often leaves those conditions unspecified, requiring it to supply them.

For example, analyzing an employment policy using income for all respondents differs from analyzing wages only among employed respondents. A policy might help some people find work without substantially changing wages for those already employed. Different findings would be unsurprising. Confusion arises if the assistant moves between those analyses while describing both simply as the policy’s effect.

Model revisions can also change comparison conditions. Adding controls may remove observations with missing values, so coefficient differences reflect both the specification and the sample. That need not be an error, but it changes what is being compared. A final table alone rarely explains the contribution of each change, and a fluent interpretation cannot make two different analyses identical.

This discretion over measurement, samples, and models is commonly called researcher degrees of freedom. AI also exercises it when taking over analysis. Whether expectations influence the choices it makes is therefore a question that can be tested.

Can Prior Beliefs Change the Conclusion?

Miao et al.’s (2026) preprint, The Agentic Garden of Forking Paths, held data and questions fixed while changing agents’ beliefs and requiring exploration of 100 models before selecting one. Findings on immigration and welfare support and other questions diverged with beliefs through exploration and selection. Many reports passed screening for major methodological errors.

A belief need not appear explicitly in the final prose to affect the work. A congenial result may seem to settle the question, while an unexpected result prompts another adjustment. If that distinction repeatedly shapes what is attempted and what is reported, the final conclusion can move closer to the initial expectation.

This helps explain why a normal-looking report may not describe an impartial process. Readers can inspect a model’s variables and estimation method without knowing what discarded alternatives produced. A final program regenerates the reported numbers, while alternatives may no longer appear in it. Reproducing a number and understanding why it was selected are separate questions.

The experiment does not establish that every task confirms an expectation. It specifically organizes sequential exploration and final selection. Greater analytical discretion increases what remains unexplained by the final output. Whether a fixed regression task behaves similarly requires separate investigation.

Will AI Directly Help with P-Hacking?

Asher et al.’s (2026) working paper, Do Claude Code and Codex P-Hack?, examines 640 runs. Ordinary estimates were stable, directional framing mattered little, and significance pressure was refused. Reframing specification search as uncertainty reporting induced searches involving time windows, fixed effects, and standard-error choices.

Task framing therefore matters. Stating a hypothesis to test differs from requesting the analysis most supportive of it. The first allows opposing evidence; the second makes the preferred conclusion part of the assignment. Describing that assignment with familiar terms such as exploration and alternative models does not establish that selection is based on research reasons rather than significance.

The studies examine different tasks, but both direct attention to choices made through actual code and computation. AI need not fabricate a coefficient to change an outcome: another specification can produce another genuinely calculated value. Evidence that a program ran answers whether computation occurred without settling why that analysis should be adopted.

How Can Defensible Choices Still Produce Bias?

Imagine a researcher has estimated a baseline regression and asks AI to explore sensitivity to alternative conditions. The assistant could add controls, change a time window, or examine different groups. Justified alternatives can deepen understanding. The question after comparison is whether the researcher sees the disagreements or only the most favorable result.

Suppose the assistant keeps revising models whenever a finding is insignificant, stops when significance appears, and reports that specification. Even if the retained model has no obvious error, the procedure has selected by outcome. Readers see one test, while the work involved selection after multiple attempts. Easier experimentation makes that history easier to hide behind a concise final report.

P-hacking involves exploiting analytical flexibility to obtain significance while concealing the relevant search. Estimating multiple models alone does not distinguish it from legitimate exploration. The distinction concerns how outcomes affect decisions and whether attempts are reported. Correcting a problem for a research reason advances the analysis; discarding a choice solely because it makes a coefficient insignificant uses the result as the selection rule.

Researchers’ responses to intermediate outputs also matter. Asking why there is no effect can be a reasonable investigation. Repeatedly requesting revisions after unfavorable outputs while immediately accepting favorable ones can gradually turn an assignment into confirmation seeking. No explicit instruction to manufacture results is required for that pattern to be concerning. This is an implication for collaboration drawn from the experiments, rather than a claim that they tested every such conversation.

Cheaper Analysis Should Make Disagreements Visible

AI can make alternatives easier to implement. Models that previously required separate coding, organization, and comparison can be executed faster. That can help researchers ask where a result holds, which definitions it depends on, and whether another justified method changes it.

More calculations do not automatically provide stronger evidence. If only one of ten analyses supports a hypothesis, a more elaborate description of that one does not erase the other nine. The additional information lies in their differences and the decisions that explain them. Stability supports a discussion of sensitivity across conditions; instability calls for reconsidering the scope of a conclusion.

Useful AI deliverables should therefore include compared results alongside the final table. Main specifications can be stated before analysis, while later exploration retains rationales and outputs. Researchers then receive disagreements to interpret instead of a report that has already selected their conclusion. Speed can help expose problems rather than merely accelerate the search for a desired answer.

The samples, variables, and weights in Sant’Anna’s account may occupy only a few lines of code, yet they affect what a study answers. AI can execute many calculations, but revisions to their conditions retain research meaning. When the same data yield different conclusions, the useful questions concern which choices caused the differences, why those choices were made, and whether the final account explains them.