OpenAI Shipped the Agents API on 10 September — and 4 Things Changed in Every Singapore Engineering Interview We Ran Since

Abstract visualisation of an artificial intelligence system processing connected nodes
Panos Petropoulos

Panos Petropoulos

Web Development Expert · 12 September 2026 · 10 min read

TL;DR

  • • On 10 September 2026 OpenAI released the Agents API in public beta, putting the Codex harness behind a single API call.
  • • Managed by OpenAI: session management, orchestration, context compression, tool calling, subagent coordination, and hosted or self-hosted sandboxes.
  • • No additional fee beyond usage was announced at launch.
  • • The orchestration layer many teams hired for has become a vendor product. That is the hiring news, not the model quality.
  • • The scarce skill moved downstream: evaluation, cost control, blast radius and rollback — not context window management.
  • • “We are building our own agent framework” stopped being a defensible reason to open a requisition this week.
  • • It is a public beta. Put a thin internal interface in front of it and keep the switching cost in days.

A client in Singapore has been recruiting two “agent infrastructure engineers” since July. The role exists because their team spent eight months building an orchestration layer: context management, tool routing, retry logic, a scheduler for long-running sessions. On Thursday, OpenAI shipped most of that as a managed API. The requisition is still open. It should not be, at least not in its current form, and the reason is worth spelling out carefully because the obvious conclusion is the wrong one.

What actually shipped on 10 September

OpenAI released the Agents API in public beta, described in its own announcement as bringing “the harness and infrastructure that powers Codex to developers through a simple, flexible API”. The word doing the work there is harness.

A model is not an agent. The gap between them is a body of unglamorous engineering: managing a context window as a session runs long, deciding which tool to call and parsing what comes back, coordinating subagents, compressing history without losing the thread, persisting intermediate state, and keeping a session alive reliably for hours or days. That layer is what every team building agents has been writing for themselves since 2024.

What OpenAI has done is operate that layer as a service. Session management, orchestration and context compression are handled on their side. Developers choose between OpenAI-hosted and self-hosted sandbox execution environments, with multi-agent configurations and tool calling available through the same interface. At launch, it was announced with no additional fee beyond usage costs.

Our expert view — this is a commoditisation event, not a capability event

Most coverage of this release has focused on what the agents can now do. For a hiring manager that is the less interesting half. Nothing became possible on Thursday that was impossible on Wednesday — teams with the engineering capacity were already building this. What changed is the price of having it, which fell to roughly zero for anyone willing to accept a vendor dependency. Commoditisation events do not reduce the number of engineers you need; they relocate where those engineers add value. Every team that has treated its orchestration layer as a differentiator now has to answer, out loud, what that layer does that a managed harness does not. In our experience about one team in five has a real answer.

The four things that changed in our interviews

1. “How would you manage the context window?” stopped discriminating

It was a reasonable screening question a month ago and it is close to useless now, because it tests familiarity with a problem that has a vendor answer. A candidate who describes a sophisticated compression strategy is describing work they may never need to do again, and a candidate who does not is no longer disqualified by that.

What replaced it in our loops: how do you know the agent did the right thing? Evaluation of non-deterministic output at scale, without a human reading every result, is the problem that has not been commoditised and shows no sign of being. The answers separate candidates sharply.

2. Cost questions became the seniority signal

When the harness is managed and billed on usage, an agent that loops is no longer a bug that shows up in your logs — it is a bug that shows up on an invoice. We now ask every candidate how they would bound the cost of a long-running session, and the range of answers is the widest of any question we ask. Weak candidates talk about setting a token limit. Strong candidates talk about budget ceilings per task, circuit breakers on repeated tool calls, alerting on cost per successful outcome rather than cost in aggregate, and the difference between an agent that is expensive and an agent that is stuck.

3. Blast radius replaced architecture as the security question

The hosted sandbox option changes the shape of this conversation for Singapore teams in particular, because a meaningful share of our clients operate under data-residency expectations from financial-services customers or public-sector contracts. The relevant question is no longer “how do we secure our agent infrastructure” but “what can this agent reach, what can it write to, and what happens to data that passes through a third-party execution environment”. That is a procurement and architecture question combined, and we now ask it of senior candidates directly.

4. Rollback became a first-class requirement

Agents that take actions rather than produce text need an undo path, and almost nobody designs one until an incident forces it. We ask candidates to walk through what happens when an agent has already sent the email, already moved the record, already merged the branch. The good answers involve staging actions for confirmation above a risk threshold, writing an action log before execution rather than after, and designing compensating actions at the same time as the action itself.

THE LINE MOVED. YOUR REQUISITIONS PROBABLY DID NOT.NOW A MANAGED SERVICE — Agents API, public beta, 10 September 2026• Session management and orchestration• Context compression across long-running work• Tool calling, subagent coordination, hosted or self-hosted sandboxesEight months of in-houseengineering for many teamsNow: one API callSTILL YOURS — and this is what you should be interviewing for• Evaluating non-deterministic output without a human reading everything• Bounding cost per task, not per token — and telling expensive from stuck• Blast radius: what the agent can reach, write to, and where data lands• Rollback and compensating actions designed before the first incident

What this means for an open requisition in Singapore

The instinctive reaction — close the role, the vendor has solved it — is wrong, and the opposite reaction, carry on as before, is worse. What is required is a rewrite.

Start by asking what the role was actually for. If the honest answer is “to build and maintain our agent orchestration layer”, that justification no longer holds on its own, and continuing to recruit against it means hiring someone to rebuild a managed service. If the answer involves constraints the managed harness does not address — data that cannot leave a jurisdiction, deterministic evaluation of output, cost governance at volume, integration with systems that cannot be exposed to an external sandbox — the role is real and it just became easier to describe precisely.

That precision matters more than usual here, because the market for these engineers is thin and badly defined. Our method for hiring AI infrastructure engineers in Singapore starts from exactly this question, and the related trap of buying a capability you should build is covered in our notes on statement-of-work clauses for outsourced AI development.

Have an agent engineering role open that suddenly reads differently?

We will go through the requisition with you and separate the parts a managed harness just absorbed from the parts that are still genuinely yours to build.

Let’s discuss it

Our expert view — the capacity you just freed is the real decision

A team that had two engineers on orchestration has not lost two roles; it has recovered two engineers. What happens to that capacity over the next quarter is a more consequential decision than anything in the release notes, and it is usually made by default. The two outcomes we see are that the capacity flows into the product backlog, which is normally correct, or that it flows into rebuilding parts of the managed harness with the argument that the in-house version was better in some specific way. The second outcome is seductive because the argument is often narrowly true and almost never worth the engineering years. If you take one thing from this week, make the reallocation an explicit decision with a named owner rather than something that resolves itself.

The public beta question, answered honestly

It is a public beta. Interfaces in public beta change, sometimes in ways that break callers, and a team that wires its core product surface directly into one has made a decision it may not realise it made.

The proportionate response is not to wait. It is to put a thin internal interface between your application and the vendor API, so the cost of a breaking change or a provider switch is measured in days. This is ordinary dependency hygiene rather than anything specific to agents, and it is the sort of judgement that distinguishes a senior engineer from a fast one — which makes it, incidentally, a good interview question in its own right.

Our expert view — the resume signal is about to get noisy

Within a month, “built production AI agents” will appear on a very large number of CVs, and it will mean something much weaker than it did in August. Calling a managed harness is not the same accomplishment as building one, and neither is the same as operating one at volume with a cost budget and an evaluation suite. We are already adjusting our screening on the assumption that this claim now carries almost no information. The follow-up that restores the signal is simple and hard to fake: ask what broke, what it cost, and how they found out.

What we are seeing across the region

The reaction has not been uniform. Singapore teams have moved fastest to ask the data-residency question, which is a direct consequence of how many of them sell into financial services and government. Our UAE team reports the same week dominated instead by cost-governance questions, driven by clients running high-volume customer-facing workloads. Our US practice is fielding the opposite call: companies that had not yet started building a harness, relieved that they no longer need to, and now hiring for evaluation instead.

Three markets, one release, three different first questions. All three are downstream of the orchestration layer, which is the point.

THE INTERVIEW SWAP WE MADE THIS WEEKRETIRED — now has a vendor answer“How would you manage thecontext window?”“How do you coordinate subagents?”“Design a retry and scheduling layer”“Walk me through your harness”Tests familiarity with a solved problem.Strong and weak candidates converge.ADOPTED — still separates candidates“How do you know the agent was right,without a human reading everything?”“Bound the cost of a long session.”“What can this agent write to?”“It already sent the email. Now what?”Hard to prepare for, hard to fake.Answer quality spreads immediately.

The bottom line

A large piece of the agent stack became a managed service on 10 September, at no additional cost beyond usage. Teams that were building that piece have lost a differentiator and gained several months of engineering capacity. Teams that were hiring for it are now hiring against a job description that describes a solved problem.

The scarce skills did not disappear; they moved one layer down, to evaluating what the agent produced, bounding what it costs, limiting what it can reach and undoing what it did. Those are harder to interview for and considerably harder to fake, which is the most useful thing about them.

Frequently asked questions

What exactly did OpenAI release on 10 September 2026?

The Agents API, in public beta. In plain terms, OpenAI took the execution harness that powers its own Codex coding agent — the layer that manages context, calls tools, coordinates subagents and keeps long-running sessions alive — and exposed it through a single API that OpenAI operates. Developers can choose between OpenAI-hosted and self-hosted sandbox execution environments, and the announced pricing carries no additional fee beyond normal usage costs. The significant part is not the model. It is that a substantial amount of orchestration engineering that teams have been writing themselves for two years is now a managed service.

Does this mean we no longer need to hire agent infrastructure engineers?

It means the justification for hiring them changed, and most teams never had a good one. Building your own harness was defensible when there was no alternative. Now the question a hiring manager has to answer is what your orchestration layer does that a managed harness does not, and for the majority of product teams the honest answer is nothing. The engineers who remain genuinely necessary are those working on constraints the managed service does not address — data residency, deterministic evaluation of agent output, cost control at volume, and integration with systems that cannot be exposed to a third-party sandbox. Those are real roles. “We are building our own agent framework” is usually not one.

How should this change our technical interviews?

Move the questions downstream of the orchestration layer. Asking a candidate to describe how they would manage context windows or coordinate subagents now tests familiarity with a problem that has a vendor answer. The questions that still discriminate between candidates are about what happens when the agent is wrong: how do you evaluate output quality without a human reading everything, how do you bound the cost of a long-running session, what is the blast radius when an agent has write access to production, and how do you roll back an action the agent already took. Very few candidates have good answers to those, which is exactly what makes them useful interview questions.

Is it too early for a Singapore team to build on a public beta?

It depends on what you are willing to rewrite. A public beta carries a genuine risk of interface changes, and building your core product surface directly against one is a decision that should be made deliberately rather than by default. The pragmatic approach most of our clients take is to put a thin internal interface between their application and the vendor API, so that the cost of a breaking change or a provider switch is measured in days rather than months. That is ordinary engineering discipline rather than anything specific to agents, and it is the same reasoning that applies to any dependency on a single vendor for a core capability.

Rewriting an AI engineering requisition this week?

We screen for the layer that did not get commoditised — evaluation, cost governance, blast radius and rollback — and we can tell you quickly whether your role still needs the seniority it was scoped for.

Let’s discuss it