Research
Itseez3D Unveils AWE 3.5: A Native Embodied AI Brain for Industrial Automation
At the World Artificial Intelligence Conference (WAIC) 2026, Itseez3D (它石智航) has unveiled its groundbreaking AWE 3.5, a native embodied foundation model poised to revolutionize industrial automation. The company showcased the model's prowess by setting up a 1:1 scale replica of a circular wire harness assembly line within the exhibition hall. Multiple robots, orchestrated by AWE 3.5, collaborated seamlessly to assemble different wire harness types at various stations, demonstrating remarkable efficiency and adaptability.
A New Era of Embodied AI in Industry
The automotive wire harness workshop, often described as an 'industrial deep water zone' due to its demanding nature and high employee turnover, has long been a challenge for traditional automation. Itseez3D's demonstration directly confronts this issue, showcasing robots performing complex, multi-tasking operations with unprecedented fluidity. Chief Scientist Ding Wenchao humorously noted that they refrained from letting the robots work continuously for fear of depleting the parts supply due to their high speed.
This advanced capability is powered by AWE 3.5, which Itseez3D claims is the first native model in embodied intelligence to fully bridge the gap between pre-training and post-training paradigms. The naming convention, referencing GPT-3.5, highlights its ambition for enhanced multi-tasking abilities. Unlike previous versions, AWE 3.5 integrates industrial assembly, plugging, and sorting tasks within a single model, eliminating the need for model reloads or parameter switches between different operations. This unified approach allows robots to perform a diverse range of tasks, from tidying desks and packing phones to sorting parts, akin to a highly adaptable human worker.
Simulating Reality: The Power of World Models
Beyond practical industrial applications, Itseez3D also provided a tangible experience of its world model capabilities. Visitors could interact with a "parallel world" generated by AWE 3.5, where actions like dropping a building block or dragging a phone box were simulated with realistic physics, including gravity and friction, without any explicit physics engine code. Even subtle details like a candle's flame were faithfully rendered, all learned directly from real-world data. This interactive demonstration offers a stark contrast to abstract diagrams and numerical benchmarks typically associated with world models.
A particularly impressive, yet often underestimated, capability highlighted is long-term memory and spatial understanding. Traditional video models struggle with predicting future states over extended periods, often leading to object disappearance or deformation. AWE 3.5, however, has been specifically enhanced with human-centric data to explicitly learn long-term memory and sequential spatial understanding, maintaining environmental stability over minutes of continuous interaction. This enables reinforcement learning within the world model itself, a significant leap forward.
Reconstructing the Embodied Brain from First Principles
Itseez3D's approach diverges from the common VLA (Vision-Language-Action) or World Model debates. Instead of integrating action modalities late in the post-training phase, as seen in some VLA architectures, or adapting existing internet video models, Itseez3D opted to train a native embodied model from scratch. This "One Model" structure integrates vision, language, and action from the pre-training stage, allowing the same architecture to output actions and predict future states. This design principle ensures a tighter coupling between perception and action, avoiding the inherent disconnects that can arise from later integration.
The "One Model" architecture offers a unique advantage: by simply altering input-output combinations, the model can adopt different "personalities." The Base Policy dictates strategy and commands, while the Action Condition World Model predicts environmental changes based on actions. Together, they form a complete reinforcement learning loop. This is exemplified by the zipper-pulling task, where the model can simulate various outcomes and optimal strategies internally before executing them on a real robot, drastically reducing trial-and-error time and hardware wear compared to traditional methods.
This self-simulation capability, where the model acts as its own simulator without the need for manually constructed environments, is a core breakthrough. Itseez3D's success is attributed to a long-term strategy combining massive real-world data and robust infrastructure capable of processing it. AWE 3.5 is trained on over a million hours of human-centric data, a scale comparable to datasets used for video generation and autonomous driving. The company emphasizes Data Efficiency, focusing on high-quality, diverse data over sheer volume, which has accelerated the training process to the point where new tasks can be mastered with only hours of data collection.
Scaling Embodied Intelligence: The Path Forward
Ding Wenchao predicts that embodied scaling has been fully unleashed, with a potential watershed moment arriving between mid-2027 and the end of the year. The industry is transitioning from "robot entertainment" to robots that can genuinely perform work. The surge in WAIC exhibitors, with over 200 embodied intelligence companies present, underscores this shift, with the robotics section drawing significant crowds.
In an interview, Ding Wenchao elaborated on the challenges and future direction. He highlighted the critical importance of hardware-software co-design, particularly the role of dexterous hands, which are essential for tasks beyond the 80% achievable with standard grippers. Itseez3D's top-down approach to designing these hands, informed by human actions and data, represents a significant investment. He also confirmed that embodied intelligence scaling has been unlocked, enabling new tasks to be completed with minimal data collection, a testament to the "One Model" architecture's foresight.
The path to Physical AGI is seen as a long tail problem, where the focus will shift from data quantity to identifying and acquiring high-value data for increasingly diverse tasks. Itseez3D is actively seeking "killer app" scenarios that provide strong positive feedback loops, driving the iterative development of data, models, and infrastructure. This mirrors the trajectory of AI coding, which propelled language models forward. The company believes that the next 12-18 months will see significant industrial trends emerge around dexterous hands, with potential for substantial progress in future AWE versions.
Redefining Research in the Scaling Era
Ding Wenchao also addressed the evolving landscape of research. He argued that traditional academic research opportunities are narrowing as industrial scaling takes precedence. The focus is shifting towards complex challenges like large-scale data curation, infrastructure design, and robust post-training iteration. He stated that "research and engineering are tightly integrated" within Itseez3D, with all employees functioning as engineers, reflecting the need for rapid, iterative development in this fast-paced field. For new entrants, differentiation is key, requiring more than just novel model architectures; a strong foundation in data and infrastructure is paramount.
When asked about judging embodied AI companies, Ding Wenchao pointed to the ability to reliably perform multiple tasks across various scenarios as the key differentiator, expected to become clear by mid-2027. He emphasized that robots will become more accessible in the next 12 months, performing real commercial tasks. The ultimate metrics will be task completion, speed of new task acquisition, and cost reduction. He also dismissed concerns about large model inference speed, stating that asynchronous processing of perception and action can achieve high frequencies, making large models viable for real-time deployment. The biggest "mystery" for him remains the potential for new capabilities to emerge from the high-dimensional complexity of dexterous hands.