Odyssey 3 - A world model that simulates physics in real time
by•
Odyssey-3 is a foundation world model that generates interactive environments from a prompt and predicts in real time how they change as you or an agent act in them. Its Pro version posts the highest reported Physics-IQ Verified video-to-video score (66.1, best-of-8). The same model has been adapted to control robot arms and humanoids, drive a car, and train agents. Try the research preview, or get in touch for API access.x`

Replies
Meet Odyssey-3 a foundation world model from Odyssey that generates interactive environments from a prompt and predicts, in real time, how they change as you or an AI agent act in them. Odyssey says it’s their most powerful world model yet.
You now get:
Environments generated from a prompt, with first-person and third-person navigation and independent camera movement
A model that predicts how objects move and interact and how situations evolve, learned from video
The same model adapted to real machines: robot arms, humanoids and a car
What’s new?
Physical accuracy: Odyssey-3 Pro posts 66.1 on Physics-IQ Verified’s video-to-video benchmark (best-of-8 sampling), which Odyssey says is the highest reported score, and 54.7 on image-to-video. The benchmark tests fluid dynamics, optics, solid mechanics, magnetism and thermodynamics. Odyssey also says the standard model improves the tradeoff between physical accuracy and generation cost, so you get more simulations from the same compute.
Real-time interaction: Odyssey-3 reacts to actions and events as they happen. A distilled few-step version makes that fast enough for real-time use.
It adapts to different bodies. By training a small action decoder or policy on top of the same foundation:
Robot arms completed tasks like “pour the cereal into the bowl” and “close the screwbox” after tens of hours of demonstrations, including recovery behaviors that weren’t in those demos, like re-orienting after a missed grasp
Flexion built humanoid control policies on it that Odyssey says beat the tested VLA baselines when the environment changed, and kept working under lighting changes that made those baselines fail
A driving policy trained on 20 hours of driving data, with the Odyssey-3 backbone frozen, drove a car on real roads in India
Synthetic data and agent training: an early experiment generated three camera views of driving scenes together after only 100 training steps, and agents can pursue natural-language goals inside the generated world.
Availability:
Research preview live now at experience.odyssey.systems, where you can prompt an environment, act in it and watch how it responds
Physical AI developers (robots, humanoids, cars, drones) can get in touch for access to build on it
P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified → @rohanrecommends
Best-of-8 is doing a lot of work in that 66.1. If the scores it's being compared against are single sample then it isn't the same measurement, and real-time use is single sample by definition because you only get one rollout when you're acting in the environment. The distilled few-step model's single-shot score is the number I'd want on the page. Publishing that would be a stronger claim than the best-of-8 headline, not a weaker one.
It’s a cool shift, getting AI to predict how the world reacts rather than just describe what it sees.