Business
Google's Mantis Skills: A Strategic Play in the AI Security Harness Race
The ongoing debate surrounding the practical performance of AI systems, particularly in demanding fields like cybersecurity, often centers on a fundamental question: Is the primary driver the AI model itself, or the sophisticated orchestration and 'harness' that guides it? While many companies argue for the latter, focusing on model-independent products, Google has entered the arena with Mantis Skills, a portable toolkit for building security review harnesses. This initiative not only offers an architecture for continuous, autonomous security evaluation and optimization but also openly shares its core elements as open source, a move that could significantly reshape the competitive landscape.
Details of Mantis Skills
Google's Mantis Skills is presented as a "Toolkit for Building Security Review Harnesses." It provides an architecture designed for continuous, and potentially autonomous, security evaluation and optimization, even extending to auto-patching capabilities. Crucially, Google has committed to documenting the architecture and design openly, making the core components available as open source for free use. This openness extends to competitors and startups, fostering a collaborative environment around the toolkit's development and adoption. While Google's own product built upon this framework remains partially undisclosed, the underlying Mantis architecture is designed to be model-independent.
This model independence is a significant departure from the prevailing trend. Mantis can, in principle, be utilized with competitor LLMs or even locally operated open-weight models. The project's authors explicitly recommend combining different models for various tasks, citing cost-effectiveness as a key driver. This approach directly contrasts with companies like Anthropic, which heavily emphasize the superiority of their proprietary models like Mythos and Fable, suggesting that complex harnesses can hinder top-tier LLMs. Conversely, Microsoft, with its agentic security helper MDASH, also focuses on the harness, stating, "The harness does the work; the model is just an input," reflecting a common strategy among companies without their own frontier models.
Context: The Harness vs. Model Debate
The AI industry has largely seen a division in strategy. Companies that do not possess their own frontier models, such as AISLE, XBOW, and IronCurtain in the security sector, tend to advocate for model-independent products built on sophisticated working environments or "harnesses." This perspective is also adopted by large corporations like Microsoft for their MDASH security helper. Their argument posits that the orchestration layer is paramount, with the AI model serving merely as an input. This business model thrives on providing a robust platform that can integrate with various AI models, offering flexibility and avoiding lock-in to a single provider.
At the extreme opposite pole is Anthropic, which champions the model as the central driving factor. Their communication strategy heavily emphasizes the exceptional capabilities of their models, such as Mythos and Fable, with slogans like "Forbiddenly good." The underlying thesis here is that overly complex harnesses can actually impede the performance of advanced LLMs. Nicholas Carlini of Anthropic has highlighted that their approach often involves connecting an LLM to a repository and allowing it to operate with minimal external orchestration, with the harness providing only initial guidance and basic guardrails. OpenAI occupies a middle ground, possessing top-tier models but employing them within an elaborate harness framework for security tasks, as seen in their Codex Security (formerly Aardvark) project. This system, benefiting from the expertise of offensive security pioneer Dave Aitel, is a closed ecosystem reliant on OpenAI's GPT models.
Google's Strategic Motivation
Google's release of Mantis Skills is not purely altruistic, despite its open-source nature. By maintaining the reference architecture, Google gains the power to shape interfaces and the broader ecosystem. This strategic positioning allows them to potentially monetize production-ready versions of the toolkit, as well as compute and operational services within their cloud infrastructure, mirroring strategies seen with products like Chrome and Android. While Mantis is currently described as a practical engineering project without explicit corporate mandate for it to become the new Google line for AI, its implications are far-reaching.
This initiative fundamentally reframes the "harness or model?" question. It represents a clever strategic move by Google, which has recently appeared to be trailing in the race for AI market leadership. By providing an open, model-agnostic framework, Google aims to become the central player in the AI security tooling ecosystem. Regardless of whether users opt for models like Claude, Fable, GPT, Gemini, or even open-weight alternatives such as DeepSeek, Qwen, and Llama, any system utilizing Mantis or its successors would ultimately benefit Google's ecosystem. The company's future commitment to this project will be closely watched.
Impact and Future Outlook
The release of Mantis Skills has significant implications for the AI security market. It democratizes access to advanced security evaluation tools, empowering startups and smaller organizations that may not have the resources to develop their own proprietary harnesses or access top-tier proprietary models. By promoting model independence and open interfaces, Google encourages innovation and competition, potentially leading to more robust and cost-effective AI security solutions across the board. Developers can now build security review systems that are adaptable to a wide range of AI models, fostering a more dynamic and less fragmented ecosystem.
For end-users and enterprises, this could translate into more secure AI applications and services. The ability to leverage diverse models, including potentially cheaper or more specialized open-weight options, within a standardized, high-quality harness framework, could drive down costs and improve the overall security posture of AI deployments. The emphasis on continuous and autonomous evaluation also promises to accelerate the identification and remediation of vulnerabilities, a critical need in today's rapidly evolving threat landscape. The long-term success of Mantis Skills will depend on Google's continued commitment to the open-source community and its ability to foster a thriving ecosystem around the toolkit.
Addendum: An internal perspective from a DeepMind researcher highlights concerns about the erosion of ethical principles following Google's acquisition and the company's secret contract with the Pentagon, which grants extensive usage rights. This adds a layer of complexity to Google's motivations and the broader ethical considerations surrounding AI development and deployment.