>OpenAI Presence Puts Human Approval Inside the Agent Improvement Loop
ALTIOR AI ADVANTAGEWhat to remember
OpenAI Presence

The agent cannot rewrite the rules

OpenAI Presence puts job-scoped access, company policy, testing and human approval around enterprise voice and chat agents.

An enterprise agent passes through scoped access, policy, approved-action and human-escalation gates within a controlled feedback loop.

An agent trusted with billing, claims or account systems needs more than a convincing answer. It must know which information it may retrieve, which actions it may take and when the case belongs with a person.

OpenAI presents Presence as the managed system around that work. A deployment begins with one job, limited knowledge and access, company-defined policies, simulations and escalation rules. Its improvement loop is bounded too: production signals expose gaps, Codex proposes changes, and teams test and approve them.

Why it matters

Useful action needs boundaries

The model sits inside a wider control system that determines what it can see, do and change.

A layered enterprise trust stack places the AI model beneath scoped access, policies, guardrails, simulations, evaluations and human escalation.
Presence places the model beneath job-scoped access, company policies, guardrails, simulations, evaluations and human-escalation rules.

OpenAI says each deployment starts with a specific job. The agent receives only the knowledge and system access needed for that work, while the company defines permitted actions, approval points and the conditions for human takeover.

That structure changes the question from whether a model can produce a plausible response to whether the whole deployment can keep its actions inside company boundaries. OpenAI describes the controls; it does not provide independent proof that every deployment will perform reliably.

The deployment path

From one job to controlled rollout

Presence links initial scope, pre-launch testing and production changes into one governed lifecycle.

The controlled enterprise-agent lifecycle from choosing one job and bounded access through simulation, launch, testing, human approval and rollout.
The source-grounded sequence runs from choosing one job through bounded access, policies, simulation and launch to production signals, proposed updates, testing, human approval and controlled rollout.
Provider image: openai-presence-still-1.webp
OpenAI's simulation view depicts teams testing a new annual-refund policy before a change reaches production.
1 jobThe starting scope for each deployment
75%Inbound issues OpenAI says its English-language phone support resolves without human assistance
15 percentage pointsOpenAI-reported reduction in human handoffs over 10 days

Before launch, teams can simulate ordinary requests, edge cases and higher-risk scenarios. OpenAI says evaluations check outcomes, policy adherence, tool use and escalation, giving teams a chance to find weak behaviour before live requests arrive.

After launch, sessions, escalations and quality signals reveal where the agent needs attention. Codex can then suggest an update, but the team tests the proposal and approves a controlled rollout. The sequence matters because improvement is presented as managed change, not live self-modification.

The published claim

What OpenAI says it can do

Presence combines conversation, system access, permitted action and human escalation within company-defined limits.

Provider image: openai-presence-still-2.webp
OpenAI's customer-support workflow shows a conversation beside a related tool action, illustrating how an agent may move from a request to a permitted system step.

OpenAI says Presence supports real-time voice and chat across customer and internal workflows. Its example follows a billing request from understanding the issue and verifying identity to retrieving account information, applying company policy and taking an approved action.

The action is only one possible ending. Policies determine when approval is needed and when a person should take over, so escalation remains part of the intended operating model rather than evidence that the agent has failed.

AI agents can answer questions, use company systems, take approved actions, and escalate to people when needed—while improving over time.

OpenAI on X
Evidence limits

The proof — and its limits

The available interface views show OpenAI's intended controls and monitoring, not independently verified outcomes.

Provider image: openai-presence-still-3.webp
OpenAI's production-health view depicts quality signals and intent analytics that can help teams identify sessions requiring attention.

The source material shows workflow, simulation and production-health interfaces published by OpenAI. These views support a narrower claim: Presence is designed to connect live signals with investigation, proposed updates and controlled rollout.

They do not establish typical reliability, financial return or superiority across enterprise deployments. OpenAI's support figures come from its own English-language phone channel, while BBVA, SoftBank and IAG are described as exploring or testing uses rather than proving broad production adoption.

A request in practice

How a request moves through the gates

Every useful action depends on identity, permitted context, policy and a clear route to a person.

A controlled request flow from understanding and identity verification through context and policy checks to an approved action or human handoff.
A billing request moves through understanding, identity verification, permitted account context and policy checks before an approved action or human handoff.

Imagine we call about a duplicate charge. The agent first needs to understand the request and verify who we are. It may then retrieve only the account context allowed for that job and apply the company's billing policy.

If the request fits the permitted path, the agent can take an approved action. If identity remains uncertain, policy requires an exception or the action sits outside its authority, the case moves to a person. Conversation quality never replaces those gates.

Verify

Confirm identity before account context or action becomes available.

Constrain

Apply only the knowledge, permissions and policy assigned to the job.

Escalate

Hand exceptions and out-of-bounds decisions to a person.

The improvement loop

How an update reaches production

A weak session can inform a change, but it cannot silently become one.

A five-stage improvement loop moves from production signals through proposal and testing to human approval and controlled rollout.
Production signals expose a gap; Codex proposes an update; the team tests it, approves it and controls its rollout.

Find the gap

Sessions, escalations and quality signals show where the agent needs attention.

Test the change

Codex suggests an update that the team can evaluate against production cases.

Control rollout

People approve the change before it moves into a managed deployment.

The phrase “improves over time” could suggest an agent learning freely from live conversations. OpenAI describes a more controlled sequence. Production evidence identifies a problem, and Codex investigates the signals and suggests an update.

The proposal then returns to the team for testing and approval. That keeps the rules, permissions and deployment decision outside the live agent, even as its behaviour is refined.

The stronger reading

The bigger shift is controlled improvement

Enterprise agents become more useful when changing their behaviour is governed as carefully as taking an action.

An enterprise agent should not earn trust because it can change. It should earn scrutiny because every meaningful action and update passes through visible boundaries.

Altior synthesis, grounded in OpenAI's Presence announcement

Presence is best understood as a control system around an agent rather than a standalone conversational product. Job scope, access, policy, simulation, escalation and post-launch change management are parts of the same promise.

The remaining questions are substantial. Presence is available through limited general availability for eligible enterprise customers, with deployments led by OpenAI Forward Deployed Engineers and selected systems integrators. OpenAI does not publish pricing or defined eligibility criteria, and its performance evidence is not independent.

Stress-test a controlled agent update

Act as the governance lead for an enterprise voice agent that handles billing enquiries. The agent may verify identity, retrieve permitted account details, apply the published refund policy, take approved account actions and escalate exceptions to a person. A production review has found that callers asking about duplicate charges are escalated too early. Propose one narrowly scoped update without changing the agent's permissions or refund policy. Return five sections: observed gap; proposed behaviour change; policy and permission boundaries preserved; simulation cases covering normal, ambiguous and higher-risk requests; and a rollout decision table with Approve, Revise or Reject criteria. Require human review before rollout, identify any missing evidence and make no claims about reliability that the test results cannot support.
Ready to copy
ALTIOR AI ADVANTAGE
Before we trust the promise

Watch the boundaries in practice

Ask who receives access, what a deployment costs, which outcomes survive independent testing, how broadly Presence is operating and whether people continue to approve material changes.

Try the prompt

Evidence that could change the assessment

  • Access and eligibility
  • Pricing and delivery
  • Independent evidence
  • Deployment breadth
  • Human control