Thursday, 27 August 2026

Aube.

News of progress
Original and translation

XPeng’s VLA 6.3.0: 30-sec memory, 6-sec prediction

Leave comparison

Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.

Original · Chinese
小鹏新版VLA将于9月推送:记住30秒,预判6秒
Translation · English
XPeng’s VLA 6.3.0: 30-sec memory, 6-sec prediction
Original · Chinese
第二代VLA 6.3.0将向Ultra和Ultra SE车型推送,G9L率先搭载。
Translation · English
Second-generation VLA 6.3.0 will roll out to Ultra and Ultra SE models, with the G9L receiving it first.
Original · Chinese
模型可结合30秒历史路况,并推演未来6秒的多种行驶可能。
Translation · English
The model can use 30 seconds of road history and project multiple driving possibilities over the next 6 seconds.
Original · Chinese
小鹏称流式推理让决策速度提升300%,但数据仍主要来自厂商测试。
Translation · English
XPeng says streaming inference improves decision speed by 300%, but the data still comes mainly from the automaker’s own tests.
Original · Chinese

8月27日,小鹏的测试车驶入没有清晰车道线的轮渡码头:它找到入口,开上渡船,跟随前车停下,靠岸后再驶入车流。这个动作背后,是小鹏第二代VLA 6.3.0的新变化:模型不只看眼前,还试图把过去和未来一起带进驾驶决策。新版本将于9月开始推送,G9L Ultra与Ultra SE首发搭载。

Translation · English

On August 27, XPeng’s test vehicle entered a ferry terminal without clear lane markings. It found the entrance, drove onto the ferry, followed the vehicle ahead and stopped, then joined the traffic after reaching the shore. Behind this maneuver is a key change in XPeng’s second-generation VLA 6.3.0: the model looks beyond what is immediately in front of it, attempting to incorporate both the past and the future into driving decisions. The new version will begin rolling out in September, with the G9L Ultra and Ultra SE receiving it first.

Original · Chinese

Infini-VLA负责记住上文,量产版本的历史记忆设为30秒。演示中,前车在道路中央掉头,车后短暂露出空位,测试车没有立刻加速,而是等待掉头完成。流式自回归推理则让模型在获取输入的同时推理并输出轨迹,小鹏称决策速度因此提升300%,减少等待完整分析结束后才行动的空白。

Translation · English

Infini-VLA is responsible for retaining context, with the production version set to remember 30 seconds of road conditions. In one demonstration, the vehicle ahead made a U-turn in the middle of the road, briefly revealing open space behind it. The test vehicle did not accelerate immediately, instead waiting for the U-turn to be completed. Streaming autoregressive inference allows the model to reason and output a trajectory as it receives inputs. XPeng says this raises decision speed by 300%, reducing the idle time that would otherwise occur while waiting for a complete analysis.

Original · Chinese

另一端是对未来的判断。X-Foresight根据历史状态推演未来6秒,Flow Matching拟合多种可能的轨迹分布,帮助车辆选择动作。面对占据逆向车道的白车,系统需要判断对方会等待、继续前进还是让出空间,再决定自己是否会堵住通路;小鹏展示的山路视频里,车辆也先减速观察悬空树枝与车身的关系,而不是把它当作普通障碍物或直接钻过空隙。

Translation · English

The other side of the system is anticipating the future. X-Foresight projects the next 6 seconds from historical states, while Flow Matching models a distribution of possible trajectories to help the vehicle choose an action. When faced with a white car occupying the oncoming lane, the system must determine whether the other vehicle will wait, continue forward or make room before deciding whether it would block the passage. In a mountain-road video shown by XPeng, the vehicle also slowed down first to assess the relationship between the overhanging branches and the vehicle, rather than treating them as an ordinary obstacle or driving straight through the gap.

Original · Chinese

小鹏把模型参数量扩大了3.5倍,并用真实数据、向量化检索和仿真组成训练闭环。公司称,模型单次训练的数据吞吐量达到1亿个视频片段;从今年6月到分享会前两周,每日仿真验证的新模型数量提升290%,综合安全能力在其仿真和测试口径下提升超过20倍。这些数字来自小鹏,尚未有独立试验公开复核。

Translation · English

XPeng increased the model’s parameter count by 3.5 times and built a training loop combining real-world data, vectorized retrieval and simulation. The company says the model’s data throughput per training run reached 100 million video clips. From June this year to the two weeks before the presentation, the number of new models undergoing daily simulation validation rose by 290%, while overall safety capability improved by more than 20 times according to its simulation and testing metrics. These figures come from XPeng and have not been independently verified in public trials.

Original · Chinese

具体来说,这套设计首先针对复杂场景中的连续判断:车辆可能更少被单帧画面牵着走;对于“找个地方靠边停车”这类指令,系统的目标是将其拆成变道、观察护栏和公交站、排除不合适区域,再寻找停车位置。9月的推送覆盖Ultra和Ultra SE,G9L Max则首发搭载VLA Lite;小鹏还称,其Robotaxi内测已完成超过2000单,搭载第二代VLA的Robotaxi获得了广州市智能网联汽车远程测试资质,可在相关道路开展主驾无安全员测试。

Translation · English

More specifically, this design is aimed first at continuous decision-making in complex scenarios: the vehicle may be less likely to be driven by a single video frame. For an instruction such as “find a place to pull over,” the system is designed to break it down into changing lanes, checking guardrails and bus stops, ruling out unsuitable areas and then searching for a parking spot. The September rollout covers Ultra and Ultra SE models, while the G9L Max will be the first to receive VLA Lite. XPeng also says its Robotaxi internal testing has completed more than 2,000 trips, and that Robotaxis equipped with second-generation VLA have obtained qualifications for remote testing of intelligent connected vehicles in Guangzhou, allowing driverless tests without an onboard safety operator on designated roads.

Back to the article