AI Hub · Our Model Stack

Our AI Model Stack: The Tools We Use, and Why

The AI model landscape changes fast. New model releases, benchmark improvements, and capability shifts happen on a timescale of months, not years. We update our stack regularly based on empirical testing, not marketing. What follows is our current assessment, accurate as of mid-2026.

Primary

Claude (Anthropic)

Our primary tool for technical development, code generation, and deep logical execution. In our testing, Claude consistently outperforms competing models on accuracy in complex, multi-step technical tasks: writing Deluge scripts, building MCP integrations, and constructing API logic.

Research & Prompting

Gemini (Google)

Excels at research synthesis, document clarity, and designing the prompts that govern how Claude operates. Its integration with Google's search infrastructure makes it particularly effective for gathering up-to-date information.

Verification

Microsoft Copilot

Our secondary cross-reference tool, particularly for Microsoft-adjacent integrations and for providing an independent check on outputs from our primary stack. Having more than one model evaluate complex logic before deployment is a straightforward quality control measure.

Legacy Reference

ChatGPT (OpenAI)

The model that brought generative AI to mainstream awareness. We reference it when testing against outputs our clients may be using internally. In our current assessment it sits behind Claude on accuracy for technical execution tasks, though the model landscape continues to evolve.

The Principle Behind the Stack

Using the right model for the right task is not a luxury. It is a quality control discipline. A business deploying AI inside its core operational systems deserves a consultancy that has done the work of understanding which tools perform, where they underperform, and why.

We have done that work, and we are honest about what we find. This extends to our approach to AI security and underwriting — because knowing which models to trust, and for what, is the foundation of responsible deployment.

Model assessments change as new versions are released. This page reflects our current working view as of mid-2026 and is updated when our empirical testing produces a different result.

Want to know which AI models and tools are right for your specific use case?