Thursday, 27 August 2026

Aube.

News of progress
PilotSingle source

XPeng’s VLA 6.3.0: 30-sec memory, 6-sec prediction

Languages for this articleCompare with the original

Machine-translated from Chinese — read the original text. 3 languages available; yours is one click away.

On August 27, XPeng’s test vehicle entered a ferry terminal without clear lane markings. It found the entrance, drove onto the ferry, followed the vehicle ahead and stopped, then joined the traffic after reaching the shore. Behind this maneuver is a key change in XPeng’s second-generation VLA 6.3.0: the model looks beyond what is immediately in front of it, attempting to incorporate both the past and the future into driving decisions. The new version will begin rolling out in September, with the G9L Ultra and Ultra SE receiving it first.

Infini-VLA is responsible for retaining context, with the production version set to remember 30 seconds of road conditions. In one demonstration, the vehicle ahead made a U-turn in the middle of the road, briefly revealing open space behind it. The test vehicle did not accelerate immediately, instead waiting for the U-turn to be completed. Streaming autoregressive inference allows the model to reason and output a trajectory as it receives inputs. XPeng says this raises decision speed by 300%, reducing the idle time that would otherwise occur while waiting for a complete analysis.

The other side of the system is anticipating the future. X-Foresight projects the next 6 seconds from historical states, while Flow Matching models a distribution of possible trajectories to help the vehicle choose an action. When faced with a white car occupying the oncoming lane, the system must determine whether the other vehicle will wait, continue forward or make room before deciding whether it would block the passage. In a mountain-road video shown by XPeng, the vehicle also slowed down first to assess the relationship between the overhanging branches and the vehicle, rather than treating them as an ordinary obstacle or driving straight through the gap.

XPeng increased the model’s parameter count by 3.5 times and built a training loop combining real-world data, vectorized retrieval and simulation. The company says the model’s data throughput per training run reached 100 million video clips. From June this year to the two weeks before the presentation, the number of new models undergoing daily simulation validation rose by 290%, while overall safety capability improved by more than 20 times according to its simulation and testing metrics. These figures come from XPeng and have not been independently verified in public trials.

More specifically, this design is aimed first at continuous decision-making in complex scenarios: the vehicle may be less likely to be driven by a single video frame. For an instruction such as “find a place to pull over,” the system is designed to break it down into changing lanes, checking guardrails and bus stops, ruling out unsuitable areas and then searching for a parking spot. The September rollout covers Ultra and Ultra SE models, while the G9L Max will be the first to receive VLA Lite. XPeng also says its Robotaxi internal testing has completed more than 2,000 trips, and that Robotaxis equipped with second-generation VLA have obtained qualifications for remote testing of intelligent connected vehicles in Guangzhou, allowing driverless tests without an onboard safety operator on designated roads.

30 secondsHistorical road-condition memory retained by the production version of second-generation VLA

Sources — read the originals(Paris time)

爱范儿ZH
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in