Xpeng's X-Foresight predictive world model debuts in vehicles: remembers 30 seconds of road conditions and predicts 6 seconds ahead.

At the XPeng Physical AI Sharing and Second-Generation VLA New Version Experience Day event held today, XPeng announced that the X-Foresight predictive world model has been integrated into vehicles for the first time, capable of predicting what will happen in the next 6 seconds and inferring the potential future behaviors of surrounding traffic participants.
The second-generation VLA new version model has comprehensively enhanced capabilities, with multi-dimensional comprehensive safety performance improved by 20 times.
The second-generation VLA new version brings four core upgrades. The on-device model parameter count has increased by 3.5 times, over 15 times that of mainstream VLA models. The larger parameter count enables stronger generalization capabilities, laying the foundation for rapid entry into global markets.
End-to-end response speed has been accelerated by 300%, achieving millisecond-level reaction and rapid avoidance. It supports effective temporal sequences of up to 30 seconds, introducing long-sequence memory into intelligent driving decision-making for the first time. It can infer scenarios 6 seconds into the future, giving vehicles predictive capabilities.
On real roads, the vehicle ahead may brake suddenly, pedestrians may change direction at any moment, and adjacent vehicles may cut in unexpectedly.
XPeng has introduced the X-Foresight predictive world model, using Flow Matching technology to predict more possible future scenarios, thereby selecting better actions.
The model not only recognizes past and current visual information but also needs to make predictions based on the target's position, velocity, motion trends, road environment, and more.
Through long-sequence driving representations, the model generates multiple driving intentions, transitioning from single deterministic outputs to multiple solutions. Through reinforcement learning, the generated trajectories become more natural and human-like.
The second-generation VLA introduces the Infini-VLA long-sequence architecture, allowing the vehicle to remember the world from the past 30 seconds, enabling driving decisions to incorporate more effective information into the model.
XPeng stated that these technologies do not work in isolation. Infini brings the past in, streaming autoregressive reasoning processes the continuously unfolding present, and X-Foresight with Flow Matching looks toward the future. The three temporal dimensions of past, present, and future are connected, forming a complete spatiotemporal decision-making loop.
From a technical evolution perspective, the second-generation VLA was first unveiled at XPeng Tech Day, and official rollout began on March 2, 2026.
This model is a native multimodal physical world large model that integrates "seeing, hearing, and reading." It innovatively removes the "language translation" step of traditional VLA models, achieving direct end-to-end generation from visual signals to action commands for the first time.
The training data volume approaches 100 million clips, equivalent to tens of thousands of years of human driving experience, requiring no manual data annotation.
Related Articles

Starting at 1.388 million yuan, Zunjie S800 Grand Design Collection officially delivered
about 2 hours ago

6 out of every 10 electric vehicles are BYD! Chinese cars dominate overseas markets: foreigners queue up to take delivery.
2 days ago

BYD released its 2026 semi-annual financial report: exports surged 67.9% to nearly 790,000 vehicles!
3 days ago

