Always-On AI Agents Need an Operating Contract, Not Just a Better Prompt
DevDay’s shift toward ongoing AI responsibilities makes ownership, permissions, quiet notifications and reliable recovery central product features.
On this page
A one-off assistant can disappoint you once. An always-on assistant can make the same mistake every morning, quietly accumulate unnecessary permissions, or notify you so often that its important warning gets ignored. Recurrence changes the product contract.
What changed at DevDay?
OpenAI’s September 29 DevDay recap emphasized agents with ongoing responsibilities, including Dots, and described expanded cloud-based Codex workflows and computer use in its Agents API. The recap also specified eligibility limits for individual offerings. This is evidence of the company’s product direction, not proof that every feature is available to every account. Availability was reviewed October 6, and the framework below is Crashtech’s product analysis rather than a vendor implementation guide.
The broader shift is from asking for an answer to assigning a responsibility. “Summarize this report” has a clear end. “Keep an eye on these reports and tell me when something matters” requires an interpretation of importance, a recurring execution policy and a way to retire the assignment when circumstances change.
What should the operating contract contain?
Start with a concrete objective and a bounded information source. A useful assignment says which report, repository or public page to inspect and what change matters. An instruction to monitor everything creates both noisy output and an unclear access requirement. More context is useful only when it is relevant and authorized.
Then define the permitted outcome. Can the agent only report a finding, prepare a draft, or modify a resource? These are separate levels of responsibility. A monitoring task should not acquire write authority simply because a write would make the issue disappear. If a human approves a particular change, preserve the scope of that approval in the task record.
Finally, define an owner, a review date and a stop condition. Recurring tasks outlive the conversation that created them. Without an owner, they become abandoned automation; without a review date, they can continue serving a goal that no longer exists. A task should be easy to pause, inspect and remove through the ordinary product interface.
How do you make notifications useful?
Notify on meaningful state transitions. If yesterday’s known issue remains unchanged, another identical message may add no value. If a previously healthy service fails, a blocked release becomes ready, or a deadline now requires human action, the transition itself is useful information.
A good notification includes what changed, the evidence, why it matters and the smallest next action. It should not force the user to reconstruct the task from a long transcript. At the same time, the underlying record should remain available for anyone who needs to investigate how the agent reached its conclusion.
Silence also needs a meaning. Distinguish “checked successfully and nothing changed” from “did not run” and “could not access the source.” Otherwise quiet automation can create false reassurance. Show operational health in the task view without turning every successful check into an interruption.
What does a well-bounded assignment look like?
Consider a hypothetical release watcher. Its objective is to identify when a named production deployment becomes ready. It may read the deployment dashboard and the public health endpoint. It should report success only when the expected revision is live and the public check passes. It must not change environment variables or redeploy another revision to make the status green.
Its notification policy can be equally specific: send one completion message, report a failed deployment with the observed error, and remain quiet while an unchanged build is still running. Its stop condition is verified completion or an explicit cancellation. That small contract is more useful than an enthusiastic instruction to “handle the release.”
For a research watcher, the contract would be different. It may collect newly published primary sources and prepare a summary, but publication could remain a separate editorial decision. The right boundary depends on the responsibility, not on how autonomous the model appears.
How should teams handle repeated execution?
Give each run an identity and preserve enough state to compare it with previous runs. If a scheduler retries after a timeout, the system should recognize whether the earlier run already completed its side effect. This matters for notifications as well as writes: sending the same alert five times can damage trust even when no database was corrupted.
Keep business state outside the model’s conversational memory. A model-generated sentence saying that an alert was sent is not an authoritative delivery record. Store the actual outcome from the service responsible for the action. On resumption, reconcile that state before deciding what to do next.
Set a limit on repeated failures. A missing credential or inaccessible source is unlikely to be repaired by endlessly restating the same request. The agent should surface a precise blocker and retain the unfinished task so that the owner can resolve it. That behavior is more useful than either silent abandonment or unlimited retries.
What metrics reveal whether recurring agents help?
Measure useful interventions, missed events, false alarms and human handling time. Run count and generated-word count mainly describe activity. They do not show whether the assigned responsibility is being fulfilled. For a monitoring task, a hundred uneventful checks can be perfectly successful if the agent catches the one event that matters.
Review a sample of silent runs too. Teams that inspect only notifications cannot tell whether the agent missed an important change. Compare the source history with the agent’s decisions, and investigate disagreements. Adjust the task definition before increasing its autonomy.
A practical pilot uses one owner, one source family and a small set of actions. Expand only after the record shows that the agent recognizes meaningful changes and respects its boundaries. The long-running agent design guide explains the recovery side of this problem. Always-on AI becomes valuable when responsibility is explicit enough to verify, not merely persistent enough to keep running.
Frequently asked questions
What is an operating contract for an always-on AI agent?
It is a durable description of the agent’s objective, information sources, permitted actions, notification policy, owner and stop conditions. It should distinguish reporting from writing and define what evidence counts as completion. This makes recurring work inspectable even after the original conversation is no longer fresh.
Should a recurring AI agent report every successful check?
Usually the task should distinguish an operational health record from a user notification. Record successful checks, but notify according to the user’s agreed policy, such as a meaningful change or required action. Also distinguish quiet success from failed access so silence never falsely implies that monitoring worked.
How can teams evaluate whether an autonomous watcher is useful?
Compare its decisions with the actual history of the monitored source. Measure useful interventions, missed events, false alarms and human handling time. Inspect silent runs as well as notifications. Start with bounded read access and expand responsibility only when the evidence supports reliable performance and appropriate restraint.
/* Comments */
Comments are offline right now — we reconnect automatically, nothing is lost.