返回文章列表
DevDay 2026Codexcloud environmentsdeveloper architectureChatGPT Work
🏗️

After DevDay 2026, Where Should an Agent Task Actually Run?

An architectural decision record for placing developer tasks across reusable cloud environments and local machines, with attention to state, artifacts, and failure recovery.

iBuidl Research2026-10-0117 min 阅读

TL;DR

DevDay 2026 makes task placement an immediate developer decision. OpenAI's September 29 release notes introduce reusable cloud environments and expanded continuity across desktop, web, and mobile. The engineering consequence is that the interface used to request work no longer reliably identifies the machine doing it. Put a task where its required state, tools, and authoritative artifacts are available. Prefer a prepared cloud workspace for reproducible repository work; keep genuinely device-dependent steps on the appropriate local machine. This article proposes an architecture for that choice rather than a benchmark or a claim that every account has every announced feature.

Decision record: a project with two kinds of state

Consider a hypothetical team maintaining a website and an iOS application. Website changes require a repository, a package manager, and automated tests. Simulator validation requires a configured Apple development machine. A designer also keeps an unpublished asset file outside the repository. The team wants an agent to prepare changes while its engineers are away from their desks. The assignment sounds singular, but its prerequisites describe several distinct execution environments.

We propose dividing the work by dependency, then reconnecting it through artifacts. A cloud task can prepare a repository change and its test evidence. A local task can inspect the external asset or run platform-specific validation. A review step can evaluate the combined result. This is a proposed design, not a description of an undocumented product orchestrator. Its purpose is to stop teams from confusing a continuous conversation with a continuous filesystem.

The distinction between project state and environment state is central. Project state includes source revisions and deliberate changes. Environment state includes installed tools, running services, credentials, caches, and files that are not committed. A conversation can mention both without transferring either. If an engineer assumes the local task has the cloud task's files because the chat looks continuous, the architecture already contains an invisible handoff failure.

The record below explains the decisions in dependency order. It uses current OpenAI documentation for product behavior and original engineering analysis for the proposed operating design. There are no invented API calls or measured performance results. The relevant question is what must be true before a task can run meaningfully, and which output makes its result reviewable somewhere else.

ADR 1: treat the client as a control surface

The DevDay 2026 release notes describe synced ChatGPT Work tasks across devices and reusable cloud environments that can continue while a user's computer sleeps. They also distinguish new synced tasks from existing tasks retaining their mode. Those statements support a modest architectural conclusion: a user interface and an execution host are separate concepts. They do not imply that every old task migrates automatically.

Our decision is to expose the execution location whenever it affects the user's next action. A task that requires an online local computer should say so. A cloud task should identify the repository revision and environment it uses. This information need not dominate the interface, but it must be available where the engineer evaluates progress and results. Otherwise, opening the same conversation on another device can create false expectations about what is accessible.

This separation is useful even in teams that use only one machine today. It makes dependencies explicit before remote work becomes routine. An engineer can ask whether the task needs a particular local directory, a device, a private service, or merely a repository checkout. That question often reveals that most of the assignment is portable, while one validation step remains tied to a device.

The rejected alternative is placement by convenience: run wherever the request originated. That works until a task encounters an unavailable dependency and improvises around it. The resulting output might be syntactically correct but unvalidated. A better architecture chooses the host deliberately and distinguishes a completed change from a change waiting for the one check that requires another environment.

ADR 2: publish a baseline, then keep task changes separate

OpenAI's cloud-environment documentation describes preparing, reviewing, and publishing a development setup. New tasks start from its prepared filesystem, while existing tasks keep their own saved state. Republishing updates the baseline for new tasks rather than rewriting ongoing tasks. That behavior makes a prepared environment more like a reusable starting point than a shared live working directory.

Our decision is to version the team's understanding of that starting point. Record the intended runtime, repository inputs, setup assumptions, and tests that establish readiness. The team need not invent an elaborate image registry to do this. A small environment manifest in its own development records can explain which setup was intended when a task began. The important distinction is whether a failing task inherited an outdated baseline or introduced its own failure afterward.

Changing the environment should be deliberate work. A developer discovering a missing system tool can fix the task temporarily, but the reusable baseline should also be reviewed if future tasks need it. Otherwise each task repeats the same repair, and the team mistakes repeated setup labor for model unreliability. Conversely, automatically pushing every task's changes into the baseline could contaminate later assignments with unrelated experiments.

The baseline is an input to work, not the output of work. Source changes belong in the repository's change process. Installed dependencies may belong in the environment setup. Temporary debugging files may belong in neither. Making these destinations explicit prevents a successful one-off session from becoming an unexplained prerequisite for the next task.

ADR 3: place work by its narrowest necessary dependency

For the hypothetical website-and-iOS project, a content parser fix is plausibly portable if the repository and required runtime are available. Simulator behavior is tied to the configured platform toolchain. A design review that requires an unpublished file is tied to wherever that file can legitimately be read. These are examples of dependency reasoning, not claims that a hosted product includes a specific simulator or design application.

Our placement rule is to use the least specialized environment that can produce sufficient evidence for the task. This preserves scarce machines for steps that truly need them. It also avoids a misleading cloud-first ideology. If a change's central uncertainty concerns local device behavior, moving its preparation elsewhere does not eliminate the need to inspect that behavior. An architecture should acknowledge the actual dependency instead of hiding it under a generic automated-check label.

The rule can be applied before any code is written. List the sources the agent must read, the tools needed to change them, and the checks needed to support the claimed outcome. If these requirements fit one prepared environment, keep the task together. If they split naturally, define an artifact boundary. If they are too entangled to split safely, choose the specialized host and accept its availability constraint.

This approach helps engineers explain incomplete results. A cloud task can finish the website tests and report that simulator validation remains outstanding. That is a meaningful partial artifact for a subsequent local task, provided the claim is bounded. It becomes misleading only when the task presents the entire application as verified despite never reaching the device-dependent step.

ADR 4: make the artifact the handoff contract

A portable conversation is helpful, but a downstream task needs material it can inspect. For a code change, the handoff should identify the revision, diff, tests performed, generated files that matter, and unresolved concerns. For a research step, it should identify source material and its cutoff. For a design step, it should identify the asset and the decision it supports. These are proposed handoff contents, not a proprietary artifact format.

The important property is independence from the original session. A local validation task should be able to find the cloud-produced change without asking the original agent to narrate every action. A reviewer should be able to inspect the evidence without trusting a sentence that says everything passed. The handoff can be concise when the diff and test results are available directly. Length is less useful than a clear connection between the claimed result and its evidence.

Artifact identity also prevents accidental validation of the wrong state. If a task checks one revision and a later task applies another change, the old test result does not cover the new revision automatically. The same issue occurs on a single workstation, but multiple hosts make it easier to overlook. A reviewable artifact should state which source state the evidence belongs to.

We would therefore make the receiving task acknowledge its input state before beginning consequential validation. This can be a simple statement that it has the expected change and required files. If the handoff is missing something, discovering that at the boundary is cheaper than discovering it after the agent has confidently repaired the wrong version of the project.

ADR 5: share source changes without sharing incidental state

Developers often rely on incidental state without realizing it. A generated file survives in a local directory. A tool was installed months ago. A service is running with an old configuration. A task succeeds because it inherits those conveniences. When the same work runs in a prepared cloud environment, the missing state appears as a failure. The failure can be useful: it exposes a dependency that the repository never declared.

Our decision is to transfer intentional project artifacts while keeping incidental state out of the default handoff. The receiving task should recreate what it can from documented inputs. When a nonrepository file is essential, name it and explain its role. When a local service is essential, make that requirement visible. Avoid copying an entire home directory or workspace as a substitute for understanding the dependency graph.

This separation also improves review. An engineer can ask whether a proposed source change is necessary, or whether the environment simply lacked a prerequisite. Without that distinction, agents may alter application code to compensate for a broken setup. Such a change can make one task pass while reducing the application's correctness elsewhere.

There is a tradeoff. Recreating declared state may initially require more setup than reusing a developer's accumulated environment. The payoff is interpretability. A later failure can be traced to a revision, a setup assumption, or an external dependency instead of an unexplained difference between machines. Teams should value that clarity without inventing a productivity number for it.

ADR 6: reuse preparation, isolate the assignment

A prepared baseline reduces repeated installation work, but tasks still need distinct working state. Imagine two agents editing the same utility file for unrelated features. If they share a writable checkout, one task's test result can depend on another task's unfinished changes. The conversation histories may be separate while the actual work is entangled. This is a coordination problem before it is a model-quality problem.

Our proposed architecture gives each assignment an explicit working revision and an ownership boundary. Agents can share the environment definition and dependency caches where supported, while deliberate source changes remain attributable to a task. The precise mechanism can be a branch, a checkout, or another repository workflow the team already understands. The architecture should not depend on undocumented claims about automatic conflict resolution.

Isolation creates a later integration step, which must not be ignored. Two individually valid changes can conflict semantically without producing a textual merge conflict. One may change an interface the other assumed was stable. Integration evidence should therefore cover the combined source state when the interaction matters. Repeating every possible test is unnecessary; repeating relevant checks after a material combination is sensible.

The design favors explicit integration over accidental collaboration through a shared directory. It also makes progress reports more honest. A task can report that its bounded change passed its checks, while the team still understands that combination with another change remains work. That is a more useful claim than a global green status with unclear ownership.

ADR 7: distinguish background execution from a waiting task

A task can appear quiet for several reasons. It may be running tests, waiting for a tool, blocked on an unavailable host, or paused for a human decision. The interface should not use the same vague in-progress state for all of them when the distinction affects action. A mobile user reviewing the work needs to know whether doing nothing is appropriate or whether the task requires the local computer to become available.

Our decision is to express waiting in terms of the dependency that prevents progress. A device-dependent check waiting for a computer is different from a review waiting for approval. A task with a missing source file is different from a long computation. The user need not see internal scheduler details. They need enough context to resolve the block or judge whether the remaining work can wait.

This becomes especially relevant with an always-on coordinator. OpenAI's Dot guide makes local-computer access optional. A coordinating agent's availability therefore should not be confused with the availability of every machine it could use. A visible assistant can remain reachable while a necessary execution host is offline.

An honest waiting state also protects the agent from pressure to fabricate completion. If the architecture expects an answer even when a required host cannot run, the assistant may substitute an estimate for evidence. Designing a first-class waiting handoff allows it to preserve completed work and explain the exact remaining check. That is an operational advantage of clear task placement.

ADR 8: retry the work that failed, not the whole story

Suppose dependency installation succeeds, a test fails, and a later service call times out. Restarting the entire task can create confusion about which earlier steps are still valid. It may also repeat external side effects. A reliable architecture needs a way to identify the failing boundary and preserve the artifacts that remain meaningful. This is a general distributed-work principle, not a claim about a specific hosted retry implementation.

Our recommendation is to classify steps by whether they can be safely repeated and what evidence they consume. Reading a source file can usually be repeated without changing the project. Editing a file requires knowing whether the prior edit already happened. Posting a result to another system needs a way to recognize the original operation. These differences should guide recovery, even when the user only sees a natural-language task.

Stripe's idempotent-request documentation provides a concrete example of a service recognizing retries through an idempotency key. That is a service-specific guarantee, not something an agent can assume for every integration. The architectural lesson is to investigate retry semantics at each external boundary rather than treating another model invocation as sufficient protection.

A useful recovery report states the last confirmed result and the unconfirmed operation. If a request timed out after submission, the next step may be reconciliation rather than resubmission. For repository work, checking the current diff may be enough to determine whether an edit landed. For external work, an operation record may be required. The common aim is to resume from evidence rather than from the assistant's memory of intent.

ADR 9: separate faster generation from faster delivery

DevDay announcements include model and execution features, but their effects enter different parts of a workflow. Faster model output may shorten drafting. A prepared environment may shorten setup. Neither necessarily shortens a simulator queue, a dependency download, or human review. An architecture that measures only time to the first answer can therefore reward improvements that barely change the time to a usable result.

Our proposed measurement is task completion with sufficient evidence. Break elapsed time into preparation, model work, tool work, waiting, and review. Observe where the team actually spends attention. This does not require a sophisticated telemetry system for an initial trial; a handful of timestamped task records can reveal whether a bottleneck is computation, availability, or a missing handoff.

Avoid comparing local and hosted execution using unlike workloads. A cloud task with a warmed dependency setup and a local task performing installation from scratch are answering different questions. Similarly, a local device check and a cloud unit test are not equivalent validation. State the task's required outcome first, then compare placements that can plausibly produce it.

Cost should be treated with the same discipline. There is no current pricing calculation in this article because plan availability, usage terms, and task composition differ. A team's own comparison should include repeated setup, failed attempts, review effort, and scarce-machine availability alongside direct usage. The cheapest visible invocation can become expensive if it produces artifacts that another engineer must reconstruct.

ADR 10: let the repository define durable work

A persistent hosted filesystem can be convenient, but it should not replace the repository's account of a software change. OpenAI's environment documentation explicitly advises saving important outputs and notes that saved state does not replace source control. Our design adopts that distinction. The task's durable contribution is the intentional change and its supporting evidence, rather than the continued existence of its original machine.

For the hypothetical team, this means the website fix can move through its ordinary review process, while the simulator result is attached to the same proposed revision. A local asset dependency can be recorded through the team's established asset workflow. The agent need not invent a new universal store for every artifact. It should make the relationship among those records explicit.

This also makes abandonment manageable. If a task is stopped or an environment becomes unavailable, the team should still know whether useful changes were preserved. A task that has only produced conversational assurances has a fragile result. A task that has a readable diff, known input revision, and saved relevant evidence can be continued by another person or tool.

Durability is not synonymous with retention forever. Temporary logs may be discarded after their useful facts are captured. Failed exploratory edits may be omitted from the final change. The design question is which material another engineer needs to assess or continue the work. Keep that material in a stable place with a clear identity, and allow incidental task state to remain incidental.

ADR 11: maintain an escape route to ordinary development

An agent workflow should improve the team's ability to complete changes, not make completion dependent on one assistant's continued interpretation. A human developer should be able to check out the proposed revision, run the relevant commands, and understand the unresolved issue. This expectation keeps architecture grounded in the software's actual development process.

The escape route is especially useful when the task's environment differs from a reviewer's machine. A clear record of runtime assumptions and external prerequisites turns a failed reproduction into a diagnosable mismatch. Without it, the reviewer may spend time asking the agent why a test passed in one place and failed in another. Such conversation can be helpful, but it should supplement evidence rather than substitute for it.

Our suggested acceptance exercise is to hand a completed artifact to an engineer who did not participate in the original task. Ask them to identify what changed, which checks support it, and what remains uncertain. Their difficulties reveal missing handoff information. This is a proposed exercise, not a reported experiment with a measured success rate.

The test also distinguishes task quality from output volume. An agent may produce a large explanation while omitting the exact revision. It may produce a concise report with everything the reviewer needs. The architecture should encourage the latter. Clear artifact identity and bounded claims do more for maintainability than a long narrative of every command executed.

Applying the record to the first real assignment

For the hypothetical website-and-iOS team, the proposed placement record would look like this. It describes dependency-based choices, not a promise about available hosted toolchains.

Work unitNecessary inputProposed placementHandoff required
Website parser changeRepository revision and declared runtimePrepared cloud workspace when prerequisites fitDiff and relevant automated-check evidence
Simulator behavior checkMatching source revision and configured Apple toolchainAppropriate local development machineObserved behavior tied to the tested revision
Unpublished asset reviewLegitimately accessible asset and design briefHost holding the required assetReviewable asset version and decision record
Combined acceptanceAll proposed changes and relevant evidenceExisting team review processBounded acceptance statement and unresolved checks

Start with a bounded change whose dependencies are familiar. Identify the source revision and the evidence needed for acceptance. Place portable repository work in a prepared environment if the documented access and tools fit. Keep the device-dependent check on the configured local host. Define the artifact each stage produces before starting the next stage.

The team can then learn from one real handoff. Did the receiving task have the expected files? Did setup reproduce the relevant prerequisites? Did the local check validate the same revision? Did the final report distinguish passed checks from checks that could not run? These observations are concrete enough to improve the architecture without making unsupported claims about the entire product category.

The lasting DevDay question is therefore practical. An agent may be reachable from many surfaces and may continue between conversations, but every action still happens somewhere with particular inputs. Good task placement makes that somewhere explicit. Good artifacts let work cross the boundary. Good recovery preserves what actually happened. Those properties turn a launch feature into a development workflow a team can understand and maintain.

Sources

更多文章