返回文章列表
agent memoryreproducibilityresearch methodsprovenanceChatGPT Dot
🧪

Can Two Researchers Reproduce a Result When Their AI Agents Remember Different Pasts?

A proposed preregistered study separates useful agent memory from hidden historical influence, treating persistent state as an experimental input rather than an invisible convenience.

iBuidl Research2026-10-0117 min 阅读

TL;DR

Persistent AI memory introduces a research input that may be missing from a conventional methods section: the agent's prior working history. Two researchers can supply the same question and documents yet obtain different analyses because their assistants retain different assumptions. The September emergence of ongoing personal agents makes this a practical concern. This article proposes a preregistered experiment to distinguish useful continuity from unjustified historical influence. It reports no experimental results. The central recommendation is to record the relevant memory state, vary it deliberately, and evaluate substantive claims rather than identical wording. Reproducibility should describe the analytical process that produced a result, including its persistent context.

The object of study: a past that is not in the prompt

Imagine two researchers comparing vendor proposals. Both have the same current documents and ask the same assistant model to explain which proposal best fits the team's requirements. One assistant previously worked with a researcher who preferred a particular vendor. The other has no such history. If the first assistant gives that vendor more benefit of the doubt, the difference might come from useful organizational knowledge, an irrelevant preference, or an inaccurate retained assumption.

This is a hypothetical problem, not an observed failure of a specific product. Its importance follows from the structure of persistent assistance. A prompt no longer fully describes the material available to the system. Earlier decisions and notes can influence which sources the assistant reads, what it considers relevant, and how it frames uncertainty. Repeating the visible request does not necessarily repeat the analytical conditions.

OpenAI's Dot memory documentation distinguishes selected conversation context from saved notes and explains that those notes can evolve. It also distinguishes the agent's notes from ChatGPT saved memory. That documented separation gives the problem a concrete product example. It does not establish that memory produces biased research; that remains an empirical question.

Our proposal studies historical influence as an input, not as a personality trait of the assistant. The aim is to determine when prior context improves a research task and when it changes the result without adequate support. A suitable experiment must preserve that distinction. Otherwise it could incorrectly label all personalization as contamination or incorrectly treat all historical influence as helpful expertise.

Terminology and the scope of the claim

The National Academies' 2019 report on reproducibility and replicability distinguishes computational reproducibility using the same inputs and methods from replication using newly collected data. We use that distinction here. The proposed study asks whether a research analysis can be reproduced when relevant agent state is supplied, and how outcomes change when that state is deliberately varied.

We define path dependence operationally: differences in prior exposure change later substantive analysis after current task materials are held constant. The phrase describes a measurable relationship, not a conclusion that the relationship is bad. Remembering a verified definition can improve consistency. Remembering an obsolete preference can create an unjustified change. The experiment needs categories for both.

We define persistent state broadly enough to include saved notes and durable task artifacts available to the assistant. We do not assume access to hidden model internals or proprietary retrieval logic. A study of a managed product may observe only a subset of its state. That limitation must be part of the research claim. Researchers should not call a run clean merely because they opened a new conversation.

The unit of interest is the substantive conclusion and its evidential support. Identical prose is unnecessary. Different wording can express the same warranted conclusion, while similar wording can conceal a changed assumption. The design therefore evaluates claims, source use, uncertainty, and decision consequences. Those outcomes are closer to the scientific question than a text-similarity score alone.

Preregistration item 1: choose a question with an inspectable answer

The initial experiment should use a task whose evidence can be reviewed without private domain knowledge. A comparison of documented features, a reconstruction of a dated policy change, or an analysis of a small synthetic dataset could work. Avoid beginning with an open-ended frontier research question whose correct interpretation remains disputed. The experiment already introduces uncertainty about memory; it should not unnecessarily multiply uncertainty about the target task.

A good task contains several separable claims. For example, proposals might differ in supported deployment locations, export formats, and stated access restrictions. A correct analysis can identify the relevant differences and explain which matter to a defined requirement. The experiment can then examine whether historical exposure changes factual extraction, prioritization, or the final recommendation.

Predeclare the intended evidence boundary. If the task uses frozen documents, live web searches should be excluded or recorded as a separate condition. If current web research is essential, capture what was retrieved and when. A changing website should not become an unrecognized rival explanation for a result attributed to memory. The research question concerns history under controlled current evidence.

Write the outcome definitions before inspecting generated answers. Otherwise the researcher may choose whichever dimension produces a striking effect. The intended contribution is not a dramatic example of an assistant being influenced. It is an interpretable account of where influence enters the analysis and whether that influence remains justified by the task.

Preregistration item 2: construct histories with known differences

The study needs prior histories that differ in an intentional, limited way. One condition might supply a verified organizational requirement relevant to the current comparison. Another might supply a preference unrelated to the requirement. A third might supply a once-correct fact that the current documents explicitly supersede. A fourth might contain neutral background activity that resembles the other histories in length and form.

These are proposed experimental conditions, not measurements. Their value is that they separate beneficial continuity from stale or irrelevant influence. If all histories contain obviously misleading statements, the experiment only tests resistance to bad context. If all histories contain helpful facts, it cannot tell whether the assistant distinguishes authority from familiarity.

The histories should avoid wording that simply commands the final answer. Telling an assistant always to recommend one vendor measures obedience to an instruction more directly than ordinary memory effects. A realistic prior history might contain a tentative preference, a documented past decision, or a note about a previous project. The experiment should state how strongly that history is intended to apply to the current task.

Inspect whether the intended information was actually retained where inspection is supported. Exposure is not identical to storage, and storage is not identical to retrieval. If a product does not expose those distinctions, the study can still measure an exposure effect, but its mechanism claim must remain narrower. This is a methodological limitation to report, not an inconvenience to hide.

Preregistration item 3: define what a reset means

Before running the experiment, register the following manipulation table. The final column states an interpretation rule to apply after scoring; it does not predict an outcome.

History conditionControlled prior materialCurrent authorityInterpretation rule
Relevant continuityVerified organizational requirementSame requirement remains validImproved application can be legitimate memory value
Irrelevant preferencePrior preference with no task justificationFrozen current evidencePreference-driven change requires a disclosed rationale
Superseded assumptionOnce-correct requirementExplicit replacement decisionContinued reliance indicates failed contextual revision
Neutral historyBackground material unrelated to the taskFrozen current evidenceBaseline for exposure without relevant content

A new chat may remove visible conversation history while preserving account-level memory, saved files, or connected-app state. A different account may remove some of those inputs while introducing different settings or availability. A reset procedure must therefore describe what was cleared, what remained, and how the researcher verified the boundary. A checkbox labeled reset is not a complete experimental method.

Our proposed design records the observable state before each run. If notes can be inspected, preserve a research copy with appropriate privacy controls. If files are part of the assignment, inventory the relevant files. If the product exposes memory controls separately, record which controls were used. Do not infer that one control removed every form of retained context unless the controlling documentation supports that claim.

Where complete resetting is impossible, use independent controlled instances and acknowledge the remaining uncertainty. The experiment may then compare intentionally exposed instances with instances whose documented history is neutral. Its conclusion should concern that comparison, rather than an absolute claim about a state-free assistant. Precision about the procedure is more important than giving the control group an impressive name.

The reset question also affects reproducibility. A second research team needs to know how the original conditions were established. If the first team cannot describe the state boundary, a failed reproduction cannot identify whether the model behaved differently or whether the new team supplied a different hidden history. The missing method becomes part of the result's uncertainty.

Preregistration item 4: hold present evidence constant

After constructing the histories, supply identical current task material to each condition. Freeze document versions, record identifiers, and preserve the task wording. Use the same stated objective and audience. If tools are involved, keep their relevant permissions and outputs comparable. A prior preference can affect a recommendation, but so can a missing document or a differently configured search tool.

The design should distinguish controlled input from controlled encounter. All conditions may have access to the same documents while choosing different ones to read. That difference can itself be a memory effect. If the study forces every condition to read identical extracts, it examines interpretation under matched evidence. If it permits free source selection, it examines the broader process, including retrieval and attention.

Both designs can be useful, but they answer different questions. A staged experiment could first test matched evidence, then test source selection with a controlled collection. This separates an assistant's tendency to overlook contrary documents from its tendency to reinterpret documents it actually read. Combining those mechanisms into a single final score would lose explanatory value.

Record the evidence encountered in each run when the product makes it observable. A claim supported by a document the assistant never accessed raises a different question from a claim that misinterprets an accessed document. The study should preserve enough material for a reviewer to distinguish unsupported inference, source omission, and direct contradiction.

Preregistration item 5: counterbalance order and account for dependence

Persistent systems create a special problem for repeated trials: a run may become part of the history for the next run. If the researcher repeatedly asks the same instance the same question, later answers can be informed by earlier attempts. Treating those attempts as independent observations would misrepresent the experimental structure.

Our proposed unit of assignment is a controlled history-instance pair. Repetitions within that pair are recorded as related observations unless the state can be restored between runs. The design should vary run order across conditions so that a provider update or a time-of-day tool issue is less likely to align perfectly with one history. The exact implementation depends on what the studied system exposes.

Do not choose a sample size by convenience and then imply statistical power the study does not have. A pilot can estimate variability and reveal whether the intended manipulation is observable. A later confirmatory design should use an explicit power or precision argument appropriate to the outcomes. If the pilot is small, report descriptive patterns and uncertainty without presenting them as settled causal magnitudes.

The study should also record interruptions and retries. A researcher might rerun an unsatisfactory answer while keeping a satisfactory one from another condition. That selection would bias the comparison. Preregister how failures, refusals, incomplete answers, and tool outages are handled. All attempts belong in the accounting, even when some cannot contribute to the intended outcome measure.

Preregistration item 6: evaluate claims before recommendations

The final recommendation is tempting because it looks like a clean endpoint. But a changed recommendation can arise from several different processes. The assistant may extract different facts, assign different importance to the same facts, or apply a retained preference after accurate analysis. These mechanisms require different remedies and should be scored separately.

We propose a claim inventory produced by reviewers who do not know the history condition. Each substantive statement is classified by its relation to the current evidence: supported, contradicted, not established, or dependent on a stated external assumption. The reviewer also records whether the source is identified and whether uncertainty is appropriately expressed. The categories should be defined with examples before reviewers see the study outputs.

Then evaluate the decision. Does the recommendation follow from the declared requirement and supported claims? If a prior organizational preference legitimately matters, does the assistant disclose that role? If a stale assumption overrides current evidence, where does it appear? This approach can detect a recommendation that happens to be right for the wrong reason, as well as one that is wrong despite mostly accurate factual extraction.

Disagreement among reviewers is evidence about the measurement process. Preserve it and adjudicate through a stated procedure. Do not quietly rewrite the rubric until every answer receives the expected score. A study of subtle historical influence needs particularly careful outcome assessment because persuasive language can make unsupported assumptions feel reasonable.

Preregistration item 7: include a correction phase

Memory's value depends on revision as well as retention. After an initial run, introduce a clear correction that supersedes one historical assumption. For instance, state that the organization no longer requires the previously preferred deployment location and provide the current decision record. Then ask for an updated analysis using the same present evidence.

The correction phase examines whether the assistant distinguishes old context from current authority. A useful system might preserve the history while marking it obsolete. Another might remove the old note yet leave its influence in an existing draft. A third might verbally acknowledge the correction but continue using the old criterion. Those outcomes cannot be discovered by checking whether the assistant repeats the corrected fact when asked directly.

Our proposed measures include which claims change, whether the recommendation changes when it should, and whether the assistant explains the revision. No particular direction is always correct. If the corrected assumption was irrelevant to the original recommendation, stability may be appropriate. If it was decisive, unchanged advice needs scrutiny. The evaluation must be tied to the task's logic.

The design should preserve both the initial and corrected artifacts. Overwriting the first answer would erase the path the study is trying to understand. A versioned record can show whether the correction reached only the latest prose or also changed the decision's evidential basis. That distinction matters when research artifacts persist beyond the conversation that created them.

Preregistration item 8: publish a provenance package

The registered package should contain six separately identifiable artifacts:

  1. The frozen task and its decision criteria.
  2. The controlled history and observable retained-state snapshot.
  3. The current evidence collection with version identifiers.
  4. Every attempted run, including incomplete runs and retries.
  5. The blinded claim inventory and reviewer disagreements.
  6. The correction input and both versions of the resulting analysis.

This sequence makes the proposed study directly inspectable. Missing artifacts should be marked unavailable with a reason; an empty slot should not be treated as evidence that an input had no influence.

The W3C PROV data model provides concepts for entities, activities, responsible agents, and derivation relationships. Those concepts offer a vocabulary for describing how a research artifact was produced. The PROV overview explains the family of specifications. Neither standard automatically guarantees that an analysis is correct.

Our proposed package maps the study's materials onto a practical record. The frozen evidence collection is an input entity. A particular run is an activity. Its answer and claim inventory are outputs. The retained-history snapshot is another input when available. The correction links to both the previous output and the new current evidence. The package can use a simple readable structure initially; adopting a formal serialization is a separate implementation choice.

The package should identify versions rather than merely name files. A document called proposal-final can change without changing its name. A reproducible record needs to distinguish the exact material used. It should also describe unavailable inputs honestly. If a managed model's internal retrieval context cannot be exported, record that limitation instead of replacing it with a guessed transcript.

Publishing provenance is useful even when a second team cannot recreate every service detail. It allows reviewers to assess which parts of the process are observable and which remain uncertain. That is a substantive improvement over presenting the final answer and prompt as if they fully specify the experiment.

Privacy constraints are part of the method

Real personal-agent history can contain private preferences, messages, or business records. A reproducibility package should not publish those materials indiscriminately. The initial study can use synthetic histories constructed for the experiment. That makes the manipulation shareable and avoids treating private user context as a convenient research dataset.

Synthetic histories have limits. They may be cleaner, shorter, or more explicit than histories accumulated through months of use. Their effects may therefore differ from real deployment. The study should state that limitation. A subsequent study with consenting participants could examine realistic accumulated context, using a privacy-preserving record of the variables relevant to the hypothesis. The precise governance would depend on the institution and research population.

Redaction also changes evidence. Removing a name may be harmless for the question. Removing a relationship or a sequence of decisions may alter the meaning of the history. Researchers should document the transformation and identify which aspects of reproducibility remain possible after it. A public package can be useful without claiming to reproduce a private context exactly.

The general principle is proportionate disclosure of analytical inputs. Include enough to assess the claim while respecting legitimate restrictions. When a necessary input cannot be shared, explain its role and the resulting limitation. Openness is valuable, but pretending a sanitized package is identical to the original would undermine the method it is meant to support.

How to interpret results without exaggerating them

Suppose the study finds that helpful history improves factual interpretation while irrelevant preferences shift recommendations. That would support a distinction between continuity and unjustified influence within the tested tasks. It would not establish a universal ranking of personal agents or show that all persistent memory is unreliable. The scope follows the histories, tasks, and product versions studied.

Suppose no difference appears. The assistant may have resisted the manipulation, failed to retain the history, or failed to retrieve it for the task. A null outcome can be informative, but its interpretation depends on whether the intended state was observable. The study should separate evidence of robust behavior from absence of a verified manipulation.

Suppose effects vary across tasks. That could indicate that ambiguity gives prior context more room to matter, or that particular source-selection steps are sensitive to history. Such patterns can guide a new hypothesis. They should not be retroactively declared the original hypothesis unless they were preregistered. Exploratory findings deserve reporting as exploratory findings.

The contribution is a better description of analytical conditions. Persistent assistants make history useful, so research methods need to make relevant history visible. Treating memory as an input allows beneficial continuity to be studied alongside its failure modes. That is more precise than either banning memory from research or assuming a prompt fully specifies an ongoing agent's work.

A methods paragraph for future agent-assisted papers

A future paper using a persistent assistant should explain what the assistant could remember, which history was intentionally supplied, how state was reset or preserved, and which outputs were checked. It should identify the model or service version where possible, the evidence cutoff, the tools used, and the analytical claims for which the assistant contributed. These are proposed reporting expectations derived from the experimental problem.

The paragraph need not include every private interaction. It should identify material context that could change the analysis and describe restrictions on sharing it. If the assistant was used only for copyediting a human-authored conclusion, that role is different from selecting evidence or interpreting a dataset. Recording the role makes the relevance of memory assessable.

The September personal-agent launches make this question timely, but the methodological lesson extends beyond a particular product. A system that remembers its past has a research history. When that history influences a result, it belongs in the account of how the result was produced. The proposed experiment gives researchers a way to study that influence before allowing invisible continuity to become an invisible experimental variable.

Sources

更多文章