The agent cannot rewrite the rules
OpenAI Presence puts job-scoped access, company policy, testing and human approval around enterprise voice and chat agents.

An agent trusted with billing, claims or account systems needs more than a convincing answer. It must know which information it may retrieve, which actions it may take and when the case belongs with a person.
OpenAI presents Presence as the managed system around that work. A deployment begins with one job, limited knowledge and access, company-defined policies, simulations and escalation rules. Its improvement loop is bounded too: production signals expose gaps, Codex proposes changes, and teams test and approve them.
Useful action needs boundaries
The model sits inside a wider control system that determines what it can see, do and change.

OpenAI says each deployment starts with a specific job. The agent receives only the knowledge and system access needed for that work, while the company defines permitted actions, approval points and the conditions for human takeover.
That structure changes the question from whether a model can produce a plausible response to whether the whole deployment can keep its actions inside company boundaries. OpenAI describes the controls; it does not provide independent proof that every deployment will perform reliably.
From one job to controlled rollout
Presence links initial scope, pre-launch testing and production changes into one governed lifecycle.


Before launch, teams can simulate ordinary requests, edge cases and higher-risk scenarios. OpenAI says evaluations check outcomes, policy adherence, tool use and escalation, giving teams a chance to find weak behaviour before live requests arrive.
After launch, sessions, escalations and quality signals reveal where the agent needs attention. Codex can then suggest an update, but the team tests the proposal and approves a controlled rollout. The sequence matters because improvement is presented as managed change, not live self-modification.
What OpenAI says it can do
Presence combines conversation, system access, permitted action and human escalation within company-defined limits.

OpenAI says Presence supports real-time voice and chat across customer and internal workflows. Its example follows a billing request from understanding the issue and verifying identity to retrieving account information, applying company policy and taking an approved action.
The action is only one possible ending. Policies determine when approval is needed and when a person should take over, so escalation remains part of the intended operating model rather than evidence that the agent has failed.
AI agents can answer questions, use company systems, take approved actions, and escalate to people when needed—while improving over time.
OpenAI on X
The proof — and its limits
The available interface views show OpenAI's intended controls and monitoring, not independently verified outcomes.

The source material shows workflow, simulation and production-health interfaces published by OpenAI. These views support a narrower claim: Presence is designed to connect live signals with investigation, proposed updates and controlled rollout.
They do not establish typical reliability, financial return or superiority across enterprise deployments. OpenAI's support figures come from its own English-language phone channel, while BBVA, SoftBank and IAG are described as exploring or testing uses rather than proving broad production adoption.
How a request moves through the gates
Every useful action depends on identity, permitted context, policy and a clear route to a person.

Imagine we call about a duplicate charge. The agent first needs to understand the request and verify who we are. It may then retrieve only the account context allowed for that job and apply the company's billing policy.
If the request fits the permitted path, the agent can take an approved action. If identity remains uncertain, policy requires an exception or the action sits outside its authority, the case moves to a person. Conversation quality never replaces those gates.
Verify
Confirm identity before account context or action becomes available.
Constrain
Apply only the knowledge, permissions and policy assigned to the job.
Escalate
Hand exceptions and out-of-bounds decisions to a person.
How an update reaches production
A weak session can inform a change, but it cannot silently become one.

Find the gap
Sessions, escalations and quality signals show where the agent needs attention.
Test the change
Codex suggests an update that the team can evaluate against production cases.
Control rollout
People approve the change before it moves into a managed deployment.
The phrase “improves over time” could suggest an agent learning freely from live conversations. OpenAI describes a more controlled sequence. Production evidence identifies a problem, and Codex investigates the signals and suggests an update.
The proposal then returns to the team for testing and approval. That keeps the rules, permissions and deployment decision outside the live agent, even as its behaviour is refined.
The bigger shift is controlled improvement
Enterprise agents become more useful when changing their behaviour is governed as carefully as taking an action.
An enterprise agent should not earn trust because it can change. It should earn scrutiny because every meaningful action and update passes through visible boundaries.
Altior synthesis, grounded in OpenAI's Presence announcement
Presence is best understood as a control system around an agent rather than a standalone conversational product. Job scope, access, policy, simulation, escalation and post-launch change management are parts of the same promise.
The remaining questions are substantial. Presence is available through limited general availability for eligible enterprise customers, with deployments led by OpenAI Forward Deployed Engineers and selected systems integrators. OpenAI does not publish pricing or defined eligibility criteria, and its performance evidence is not independent.
Stress-test a controlled agent update
Act as the governance lead for an enterprise voice agent that handles billing enquiries. The agent may verify identity, retrieve permitted account details, apply the published refund policy, take approved account actions and escalate exceptions to a person. A production review has found that callers asking about duplicate charges are escalated too early. Propose one narrowly scoped update without changing the agent's permissions or refund policy. Return five sections: observed gap; proposed behaviour change; policy and permission boundaries preserved; simulation cases covering normal, ambiguous and higher-risk requests; and a rollout decision table with Approve, Revise or Reject criteria. Require human review before rollout, identify any missing evidence and make no claims about reliability that the test results cannot support.Ready to copy
Watch the boundaries in practice
Ask who receives access, what a deployment costs, which outcomes survive independent testing, how broadly Presence is operating and whether people continue to approve material changes.
Try the promptEvidence that could change the assessment
- Access and eligibility
- Pricing and delivery
- Independent evidence
- Deployment breadth
- Human control