OpenAI’s new enterprise product is not another general-purpose chatbot. Presence is a managed system for putting voice and chat agents into specific jobs, with company policies, approved actions, evaluations and human escalation built around the model.
That distinction makes Presence one of OpenAI’s more consequential product announcements. The frontier-model race still dominates headlines, but enterprises rarely fail because a model cannot produce an impressive answer in a demonstration. They fail when an agent enters a live workflow without dependable permissions, clear limits, tested escalation or a way to improve safely after launch.
Presence is OpenAI’s attempt to package those operational layers. It is available through a limited general-availability programme for eligible enterprise customers, led by OpenAI Forward Deployed Engineers and selected systems integrators. It is not yet a self-service product.
That delivery model tells us almost as much as the feature list. OpenAI is not claiming that a company can switch on autonomous customer service with a credit card and a prompt. It is selling a deployment discipline.
The product is the operating envelope
OpenAI describes each Presence deployment as beginning with a specific job: resolving billing issues, supporting insurance claims or handling employee IT requests. The agent receives only the knowledge and system access required for that job. The customer determines what it can do, when approval is needed and when a person should take over.
Presence then combines the less glamorous components that make production agents viable: policies and standard operating procedures, guardrails, approved actions, simulations, evaluation tools and an improvement process powered by Codex. Production sessions and escalations reveal weaknesses; Codex can propose updates; teams test those changes against the live version and approve a controlled rollout.
This architecture matters because a capable model is only one part of an accountable service. The model interprets a request and reasons about a response. The operating envelope determines which data it can see, which tools it can invoke, what counts as acceptable performance and where its authority stops.
Many organisations have tried to build that envelope themselves from prompts, retrieval systems, API connections and dashboards. Presence turns the integration and continuous-evaluation layer into a named OpenAI product. Strategically, it moves the company closer to owning the enterprise agent stack rather than supplying only the intelligence inside it.
OpenAI is selling improvement after launch
Traditional software is usually tested against specified behaviour and then monitored for defects. An agent is less stable. Customer language changes, policies are updated, products evolve and previously rare requests become common. A prompt that worked in a pilot can fail when it meets the ambiguity and adversarial pressure of production.
Presence addresses that problem with an explicit improvement loop. OpenAI says simulations and graders test whether an agent reached the right outcome, followed policy, used tools correctly and escalated when necessary. After launch, sessions, handoffs and quality signals feed further investigation. Proposed changes can be tested against the deployed version before rollout.
This is a stronger model than silently editing a system prompt after complaints arrive. It creates the possibility of versioned behaviour: a proposed change, a test set, a comparison and an approval. But the value depends on the quality of the evaluation regime. A grader cannot protect what it does not measure, and a test set built only from ordinary requests may miss the edge cases that create the greatest harm.
Enterprises should ask what evidence a Presence deployment actually produces. Can a reviewer see the scenarios tested, the policy version applied, the actions taken, the reason for escalation and the before-and-after results of an update? If not, “continuous improvement” risks becoming an attractive phrase rather than an auditable control.
The early performance claims need context
OpenAI says Presence powers its English-language telephone support channel and now resolves 75% of inbound issues without human assistance. It also says a Codex-powered improvement loop reduced human handoffs by 15 percentage points in ten days. Those figures are notable, but they are company-reported results from OpenAI’s own deployment; the announcement does not provide an independent evaluation or enough detail to generalise them across industries.
The launch partners illustrate the intended market. BBVA is exploring voice support for everyday banking needs in Mexico, SoftBank is testing Japanese-language customer conversations, and IAG is exploring support during high-demand insurance events such as severe weather. OpenAI’s wording is careful: these companies are exploring or testing the product, not presenting completed, universally successful transformations.
That caution is appropriate. Resolution rate alone is not a sufficient measure of service quality. An agent can reduce handoffs by discouraging escalation, narrowing the definition of a resolved case or creating downstream rework that the headline metric does not capture. A credible deployment needs a balanced scorecard: correctness, policy compliance, customer effort, repeat contact, complaint rate, inappropriate action, escalation quality and outcomes for vulnerable users.
Least privilege becomes a product requirement
Presence is designed to do more than answer questions. It can use company systems and take approved actions. That is where agent capability turns into delegated authority.
NIST’s current work on software and AI-agent identity asks how organisations can establish least privilege for agents, bind an agent’s authority to a human and retain tamper-resistant records of actions and intent. OWASP’s guidance recommends minimum tool permissions, separate approval for high-risk actions and a clean boundary between deciding and executing irreversible operations.
Presence’s job-specific access model is aligned with those principles, at least at the level of product design. The harder question is how precisely it works in a customer environment. “Access to the billing system” is too broad. An agent might need to read an account, calculate an adjustment and submit a refund below a threshold, while being prohibited from changing identity data or overriding fraud controls. The useful unit of permission is the action, not merely the application.
Human escalation also needs more than a handoff button. The receiving person needs the conversation, relevant evidence, actions already attempted, applicable policy and the reason the agent stopped. Otherwise automation can make the first stage faster while making the difficult cases slower and less intelligible.
Policies must be executable without becoming invisible
OpenAI repeatedly places policies at the centre of Presence. That is welcome, but it creates a governance challenge. A written policy is designed for interpretation by people across many situations. An agent needs a more operational expression: decision rules, thresholds, prohibited actions, approval routes and exception handling.
Turning policy into machine-enforced behaviour can improve consistency. It can also hide consequential interpretations inside configuration. Someone must decide how a phrase such as “reasonable evidence” or “appropriate support” becomes a rule the agent follows. Those decisions should be recorded, reviewed by the right owner and reopened when the model, workflow or policy changes.
That is why our argument that AI governance needs decision logs, not just policies applies directly here. The organisation needs a durable account of who translated policy into agent behaviour, which evidence supported that choice, what residual risk was accepted and which event triggers review.
Presence changes the competitive question
Model quality still matters. Better reasoning can raise task success and reduce the cost of handling difficult cases. Yet Presence suggests that OpenAI believes the next enterprise battle will be fought over deployment systems: evaluations, permissions, workflow integrations, improvement loops and the people who configure them.
That puts the company in closer competition with customer-service platforms, systems integrators and specialist agent vendors. It may also create deeper dependence. If policies, evaluations, integrations and improvement history accumulate in one provider’s operating layer, switching models may become harder even if model APIs remain interchangeable.
Enterprise buyers should therefore assess portability alongside performance. Can the organisation export conversation records, test suites, policy mappings, action definitions and evaluation results? Can it substitute a model or change an integrator without reconstructing the operating model from scratch? The more successful Presence becomes, the more valuable those questions will be.
The real test is controlled usefulness
Presence is a serious acknowledgement that production agents need more than intelligence. Its emphasis on bounded jobs, restricted access, evaluation, escalation and controlled updates is directionally right. The limited, engineer-led launch is also more credible than pretending the system is already a universal self-service solution.
What remains to be proved is whether those controls stay legible under commercial pressure. Enterprises will want higher automation rates and faster expansion into new workflows. The product will earn trust if it can increase useful autonomy while preserving precise permissions, meaningful human intervention and evidence that survives scrutiny.
The most important Presence metric will not be how often the agent avoids a person. It will be how often it resolves the right problem, under the right authority, with a record that lets the organisation explain what happened afterwards.
Sources
- OpenAI: Introducing OpenAI Presence
- NIST: New concept paper on identity and authority of software agents
- OWASP: AI Agent Security Cheat Sheet
How we work: Artificially Confident articles are source-led, AI-assisted and editorially reviewed.

