What changed in September 2026

OpenAI introduced the Agents API in public beta as managed infrastructure for cloud agents. The company describes support for long sessions, tool use, subagent coordination, saved intermediate results and a choice of execution environments.

Microsoft announced Copilot Autopilot and Copilot Managed Runtime. Autopilot is positioned as a proactive agent that keeps working without constant user attention, while Managed Runtime provides an IT-governed environment for running code within Microsoft 365. Together, the announcements point in the same direction: the agent is becoming a participant in an operational workflow rather than a chat window.

A long-running agent is a service, not a prompt

A short chat ends with an answer. A long-running agent must survive a pause, wait for an event, recover context, retry a safe operation after failure and explain what has already happened. Model quality is therefore only one layer of the solution.

The architecture needs explicit task state, allow-listed tools, action logs, time and budget limits, error handling and human approval. When an agent acts for an employee, its permissions should be no broader than that employee's and bounded to the specific purpose.

  • Durable state and controlled recovery
  • Least-privilege access to data and tools
  • Limits on time, cost and actions
  • A complete log of decisions, calls and outcomes

Choose the right first pilot

Start with a repeatable task that has a clear input, visible output and low cost of delay—not with the largest process. Examples include a weekly operational summary, preparation of a document pack, application completeness checks or monitoring defined sources for changes.

A poor first candidate lets an agent transfer funds, change critical records or send external messages without review. Begin with an agent that prepares the action and a person who approves it. Increase autonomy only after collecting evidence about errors and exceptions.

  • One process owner
  • One measurable outcome
  • A bounded tool set
  • A clear human approval point

Agent economics needs its own metric

Microsoft explicitly links long-running agentic work to usage-based billing. This matters because cost is no longer always driven by seat count; it can depend on task duration, model choice, tools, retries and context volume.

A pilot should measure total cost per completed task rather than tokens or runtime alone. Include human review, correction, infrastructure, observability and maintenance of data sources. Compare the result with the baseline process on quality, cycle time and cost per acceptable outcome.

A practical 30-day plan

Use week one to map the workflow and record baseline metrics. In week two, assemble the smallest data and tool boundary in read-only mode. In week three, run the agent in a test environment with approval for every external action. In week four, compare it with the manual process and review every deviation.

Scale only when three things are true: the agent completes the intended work reliably, the team can explain every critical action, and the full task cost is acceptable. If one is missing, improve the operating boundary before adding autonomy.

Quick answers

What is a long-running AI agent?

It is a software agent that performs a multi-step task over an extended period, retains state, uses approved tools and can resume after a pause or failure.

How is it different from a chatbot?

A chatbot mainly produces an answer within a conversation. A long-running agent maintains task state, acts in systems, waits for events and returns an auditable result.

Which process should be automated first?

Choose a repeatable, low-risk workflow with a clear input, measurable output, available data and a defined human approval point.

How should agent cost be controlled?

Set limits for runtime, models, tools and retries, then measure total cost per completed task including human review and corrections.

Primary sources

  1. OpenAI — Introducing the Agents API, 10 September 2026
  2. Microsoft — Introducing the new Copilot with Home, Code and Autopilot, 25 September 2026
  3. Microsoft — Building the system for AI at work, 23 September 2026