The updated OpenAI Agents SDK introduces a model-native harness and native sandbox execution. Build agents that safely inspect files, run commands, and execute code for long-horizon tasks across built-in providers like E2B, Modal, Daytona, and Vercel.
GPT-5.3 Instant delivers more accurate answers, better web synthesis, fewer unnecessary refusals, and a more natural tone without the cringe, caveats, or dead ends. It writes with more range, responds with better judgment, and stays focused on what you actually asked. Same speed. Sharper results. Better conversations by default.
The official Codex desktop app by OpenAI brings parallel coding agents natively to Windows. It isolates tasks in OS-level sandboxes and dedicated worktrees so agents can write, test, and propose code without trashing your local environment.
Codex just got faster, more reliable, and better at real-time collaboration and tackling tasks independently anywhere you develop—whether via the terminal, IDE, web, or even your phone.
FrontierScience is a new benchmark for evaluating AI’s expert-level scientific reasoning across physics, chemistry, and biology. It measures both Olympiad-style problem solving and real research tasks, helping track how well advanced models can support and accelerate scientific work.
On their livestream today, OpenAI just released a bunch of new tools for reliably building and using AI agents. From what I can tell, this is what's new-
New APIs:
Responses API - a new multi-modal API that builds on chat completions to allow for the next-generation of tool calling, starting with the new tools announced today.
A new platform that helps enterprises build, deploy, and manage AI agents that can do real work. Frontier gives agents the same skills people need to succeed at work: shared context, onboarding, hands-on learning with feedback, and clear permissions and boundaries. That’s how teams move beyond isolated use cases to AI coworkers that work across the business.
Advances the frontier of coding and computer work. SOTA on SWE-Bench Pro (57%) and OSWorld (64%). Features mid-task steerability (interact while it works), 25% faster speeds, and "High" capabilities in cybersecurity.