Odyssey-3 is Odyssey's most powerful foundation world model yet, described as a learned dynamical system implemented as an autoregressive diffusion transformer that predicts how objects move and interact through space and how situations evolve over time. It generates embodied environments from a prompt and predicts in real time how those environments change as a person or an agent takes actions or introduces events, using previous observations together with the latest inputs. The model is built for physical AI developers — teams building robots, humanoids, self-driving cars, drones, and other autonomous systems — who want to explore how foundation world models can accelerate their work. A research preview is available now, letting users prompt an environment, act within it, and observe how the world model responds, while developers interested in building with Odyssey-3 can get in touch for API access.
The starting point for Odyssey-3 is a problem the founding team lived with for a decade. They spent that decade building driverless cars, where predicting the world was essential to determining what a car should do next. They founded Odyssey to pursue that idea far beyond the roads — to build a general-purpose technology that could bring learned world knowledge to all machines and tasks. With Odyssey-3, they believe this idea is now being realized. Physical accuracy is central to that goal, because a foundation model for physical intelligence must learn to predict how the world actually behaves rather than simply producing plausible-looking video. Odyssey-3 learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations, and developers use that knowledge both to simulate environments and to train policies for different physical systems.
The first capability is real-time environment generation. Odyssey-3 generates embodied environments from a prompt and predicts in real time how they change as a person or an agent takes actions or introduces events, using previous observations and the latest inputs. You can move through the environment or introduce an event during generation and observe how the model responds. Today's preview provides first-person and third-person navigation alongside independent camera movement, giving you different ways to interact with and inspect the model's predictions.
Physical accuracy is measured against public benchmarks. Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified's video-to-video benchmark, achieving 66.1, the highest reported score. Physics-IQ tests physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics, asking models to continue videos of real physical experiments and comparing their predictions with what actually happened. Odyssey-3 Pro also scores 54.7 in image-to-video, while Odyssey-3 improves the measured tradeoff between physical accuracy and generation cost, making it possible to generate more simulations within the same compute budget. On WorldMark, which measures control-following, visual quality, and world memory, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized environments, using the benchmark's own captions and the mean of its 13 reported metric scores.
Odyssey-3 adapts to physical systems. Its learned world knowledge can be applied to different systems by training an action decoder or policy on paired observations and actions. These learned components translate that knowledge into the controls required by a particular machine, letting developers adapt the foundation to a new body or task. With only tens of hours of robot demonstrations, Odyssey-3 completed manipulation tasks and showed recovery behaviors absent from those demonstrations, including reorienting a gripper after a missed grasp and retrieving a dropped object in an unusual position. Flexion has built humanoid control policies on Odyssey-3 whose performance exceeded the tested VLA baselines under environmental changes and continued to perform tasks under lighting changes that caused those baselines to fail. Odyssey-3 was also adapted to drive a car on real roads in India, training a driving policy on just 20 hours of driving data while keeping the Odyssey-3 backbone frozen; the policy uses the model's visual representations to predict waypoints ahead of the car, allowing it to drive in closed loop. The model can additionally be adapted to generate observations for particular sensor arrangements — in an early experiment using the front three cameras of an autonomous-driving dataset, an Odyssey-3 training checkpoint produced driving sequences with three camera views generated together after just 100 training steps.
Odyssey-3 can also help train agents. An agent is an AI system that pursues a goal by observing its surroundings, choosing actions, and using what happens to decide what to do next. A world model can provide the environment in which those decisions are made, giving a way to study how an agent responds to changing conditions and whether it can complete a task inside a world whose behavior is learned. In the task-completion demonstration, an agent receives a natural-language goal and pursues it inside Odyssey-3, observing the generated world as it works toward the task.
How Odyssey-3 was built is described in the launch material. Its training data combines internet video with time-localized, schema-verified event annotations, gameplay recordings with time-aligned keyboard and mouse inputs, and simulated rigid-body interactions with captions and metadata. Together, these sources connect diverse observations with descriptions of what happens and, where available, the actions that produced it. The stated design goal was to enable dynamic, open-ended interactions with an environment that responds as the user acts. Odyssey-3 is built as a multi-step video diffusion transformer, using temporally resolved prompts and controls to guide how the world unfolds. It is then extended autoregressively through teacher forcing and causal masking, training it to continue from preceding observations and predict future states conditioned on action inputs. Finally, a post-training pipeline that combines distribution-matching and adversarial distillation produces a distilled variant of Odyssey-3 — a few-step model capable of real-time interaction.
The measured results point to concrete benefits for users. Because physical accuracy and generation cost are traded off more efficiently, developers can generate more simulations within the same compute budget. Because the model can be adapted with small amounts of data — tens of hours of robot demonstrations, or 20 hours of driving data with a frozen backbone — developers can bring learned world knowledge to a new machine or task without training a world model from scratch. Because the distilled variant is a few-step model capable of real-time interaction, environments respond as the user acts rather than only after generation is complete. And because the model can provide the environment in which agents pursue natural-language goals, it offers a way to study agent behavior in worlds whose behavior is learned.
Odyssey-3 appears in several concrete scenarios in the launch material. For robot arms, prompts such as "Pour the cereal into the bowl" and "Close the screwbox" were used with the model to perform manipulation. For humanoids, Flexion built control policies on Odyssey-3 and ran tasks such as "Open the blue container and take out the cardboard box" and "Move the plate to the center of the table and place the mug on top of it." For vehicles, prompts like "Take the first roundabout exit" and "Drive along the road" were used while the adapted policy drove autonomously on real roads in India. For multi-sensor data, Odyssey-3 generated driving sequences with three camera views produced together for an autonomous-driving dataset. For agents, an agent receives a natural-language goal and pursues it inside the generated world. And in the research preview, a user prompts an environment, navigates it in first or third person, and introduces events to see how the model responds.
Odyssey-3 is aimed at physical AI developers — teams developing a robot, humanoid, self-driving car, drone, or any other autonomous system — as well as researchers studying world models and anyone who wants to try the research preview. Developers who want to build with Odyssey-3 are invited to get in touch for API access. The research preview, offered as "Try Odyssey-3 Flash," is available now through experience.odyssey.systems. Published evaluation details describe two model sizes: Odyssey-3 at 832×480 and Odyssey-3 Pro at 1280×720. The launch material presents Odyssey-3 as available today as a research preview, with API access offered on request; no pricing or plan details are published in the provided content.
Taken together, Odyssey-3's primary value proposition is that it brings learned world knowledge to physical AI: a foundation world model that generates interactive environments from a prompt, predicts in real time how they evolve as someone acts, advances the reported state of the art in physical accuracy on Physics-IQ Verified, and can be adapted to control robot arms, humanoids, and vehicles or to serve as the environment in which agents learn and are evaluated.