Business

Real-World Workflows Emerge as the Next Frontier for AI Model Training

AI
AI Hub Feed
July 23, 20265 min read

The landscape of artificial intelligence model training is undergoing a significant shift, with a growing consensus among AI labs and data companies that real-world enterprise workflows are the next critical source of high-quality training data. This emerging trend signals a move away from the heavily mined internet data towards a richer, yet largely untapped, reservoir of information generated within businesses' day-to-day operations. This transition is driven by the inherent limitations of existing data sources and the increasing complexity of AI tasks, particularly in areas requiring long-horizon reasoning and intricate task execution.

The Shifting Data Paradigm

The data market is being likened to a dynamic surfing business, where continuous adaptation is key to success. As AI capabilities advance, the definition of "most valuable data" evolves, creating new opportunities and challenging established players. Companies like Scale AI and Mercor have capitalized on previous waves of data demand, such as data labeling for autonomous vehicles and expert data for early LLMs. Now, as the focus pivots to real-world workflows, new entities like Protege, Sunset, and SF Data are emerging, while established companies like Mercor and Handshake are recalibrating their strategies to address enterprise data use cases. This evolution is fundamentally driven by the pursuit of Long Horizon tasks, which require models to handle increasingly extended sequences of actions, moving from single code modifications to entire project orchestrations.

Building the Future Data Pipeline

Two primary categories of data are shaping this new era: Type 1 data, captured directly from real work processes like session replays and collaboration logs, and Type 2 data, artificially constructed through expert-designed tasks. While Type 2 data has been instrumental for pre-training and early development, its limitations become apparent as models tackle higher-value, longer-sequence tasks. The ideal scenario involves achieving a Type 1.5 state, where artificially constructed data closely mimics real-world scenarios. This requires deep partnerships with experts, refined incentive mechanisms to prevent data distortion, and robust, multi-layered quality assurance processes that go beyond simple manual review. The true challenge lies not just in creating tasks that current models cannot solve, but in designing tasks that sit precisely at the model's capability boundary, offering sufficient challenge for steady progress.

Challenges and Opportunities in Data Provision

Despite the burgeoning interest, the real-world data market faces significant hurdles. Issues such as data distortion, where Type 2 data is misrepresented as Type 1, and evaluation distortion, where benchmarks are manipulated to inflate perceived model improvement, plague the industry. Furthermore, the scalability of quality assurance remains a critical engineering problem, often misconstrued as a purely operational one. The demand transmission process is also fraught with inefficiencies, as researchers struggle to articulate precise data needs, and intermediaries can introduce information loss. To build trust, data vendors must demonstrate data taste, research taste, and scalability, moving beyond generic pitches to provide auditable data lineage and measurable performance gains from their data.

The Evolution of Reinforcement Learning Environments

Reinforcement Learning (RL) environments are evolving from simple sandboxes into complex software spaces where AI agents can operate, utilize tools, and complete intricate tasks. These environments are expanding in complexity across several dimensions: an expanded action space allowing agents to move across interconnected systems, dynamic tools that may include other AI models with uncertain outputs, and organizational complexity where tasks are distributed among multiple sub-agents. The ultimate goal is to shift training objectives from single actions to the orchestration of entire end-to-end workflows. The key to effective RL lies not only in the verifiability of outcomes but also in the environment's ability to be repeatedly reset for sufficient model practice.

Bridging the Gap to Real-World Complexity

While coding environments are highly replicable and resettable, the real world presents a stark contrast. Activities like founding a startup or managing investments have measurable outcomes but cannot be easily replicated for mass trial-and-error learning. This necessitates an improvement in sample efficiency, enabling AI to extract transferable patterns from limited, non-repeatable experiences. Approaches like on-policy self-distillation (OPSD) are being explored, where a "veteran model" that has learned from extensive real-world interaction can guide a base model, effectively compressing experience and transferring judgment frameworks. This marks a critical step towards AI agents that can navigate the nuanced complexities of human professional workflows, moving beyond brute-force learning to sophisticated, efficient problem-solving.

The Future of Data and AI Collaboration

The dichotomy between data-rich enterprises and Machine Learning Engineer-rich but data-poor AI labs is fostering a new ecosystem. Companies are emerging to bridge this gap, either by providing specialized datasets or by offering RL-as-a-service (RLaaS) to convert business workflows into trainable environments. This symbiotic relationship is crucial for unlocking the next generation of AI capabilities. The emphasis on real-world data signifies a maturation of the AI field, moving from theoretical constructs to practical, impactful applications that mirror human professional endeavors. The companies that can effectively navigate the challenges of data quality, scalability, and the intricate demands of real-world workflows will undoubtedly lead this transformative wave.

Related Articles

36Kr AI