Test & Alignment for AI Agents

Don't just deploy agents.
Measure them.

Kangguru builds Gauge — a test and alignment framework that measures an AI agent's ability, efficiency, and compliance before you trust it with real work — plus the virtual AI officers and expert services to govern it in production.

Explore Gauge Book a discovery call →
$ gauge --suite core-alignment
gauge session s_84f2 started · agent connected via MCP
gauge task 01/12 exact-match ......... PASS
gauge task 02/12 numeric ............. PASS
gauge task 03/12 prompt-injection .... PASS
gauge task 04/12 canary-url .......... FAIL
gauge task 05/12 llm-judge ........... PASS
...
gauge score 91.6 · compliance 1 finding
gauge report → leaderboard updated
Why Kangguru

The agentic workforce has arrived.
Nobody is grading it.

Enterprises are deploying AI agents faster than they can evaluate them. Most have no way to answer the basic questions: Is this agent capable? Is it efficient? Will it follow the rules when nobody is watching? Kangguru exists to make those questions measurable — and to govern what the measurements reveal.

78%
of organizations now use AI in at least one function. (Stanford AI Index 2025)
80%
of companies report their AI agents have already taken unintended actions. (SailPoint 2025)
57%
of employees hide their AI use from their employer. (KPMG 2025)
~$670K
added cost of AI-related breaches above the 2025 baseline. (IBM 2025)
Products

Two products. One mission: agents you can trust.

Measure agents before deployment with Gauge. Govern your data and security posture in production with the virtual AI officer suite.

GOVERNANCE SUITE

Virtual AI Officers — governance that never sleeps.

A coordinated team of AI agents for enterprise security and data governance, anchored by the vDIO (Virtual Data Inventory Officer) — 100% offline data asset discovery for enterprises that can't use cloud solutions — with vCISO capabilities on the roadmap.

  • vDIO data discovery — endpoint sensors map what data lives where, metadata-only by default.
  • 100% offline — nothing leaves your environment; built for finance, pharma, and defense.
  • Local AI insights — on-prem document analysis and LLM summarization, no cloud calls.
  • vCISO (planned) — security posture monitoring, threat detection, compliance validation.
How Gauge Works

Connect. Test. Grade. Compare.

01 · CONNECT

Point your agent at Gauge

Any MCP client is a test subject. Register, start a session, and Gauge takes over — stdio or HTTP, your choice.

02 · TEST

Run the curriculum

Gauge dispenses tasks one at a time in server-assigned order — capability, efficiency, and security-behavior tests including prompt injection traps.

03 · GRADE

Deterministic scoring

Every submission is graded by exact match, regex, numeric, canary observation, or LLM-as-judge — with a complete event log for audit.

04 · COMPARE

Benchmark & align

Scores land on the leaderboard. Findings feed your alignment work — fix, re-test, and track improvement over time.

Services

Expert services around the products.

Whether you're at "we have no AI policy" or "we have hundreds of agents in production," we meet you where you are. See how we engage →

01

Agent Evaluation & Alignment

We design Gauge test suites for your agents and use cases, run structured evaluations, and turn findings into concrete alignment fixes — before deployment and continuously after.

02

Consulting

AI agent risk assessment, why conventional IT security fails for autonomous agents, and the layered-defense approach to regaining control.

03

Training

Hands-on risk demonstrations from real incidents — prompt injection, shadow AI exfiltration, agent privilege escalation — plus a working tour of the AI security solution landscape.

04

Integration

Hands-on deployment: Gauge in your CI pipeline, vDIO sensors on your endpoints, guardrails and telemetry in your stack. We deliver running systems, not slide decks.

Our Team

Veterans of AI and security.

Decades of combined experience across enterprise security, AI risk, adversarial research, and large-scale platform operations.

Former Chief Security Scientist
Global internet company, multi-decade tenure in enterprise security architecture and threat research.
Chief Technology Officer
Leading security firm; deep experience scaling security platforms across enterprise customers.
Black Hat & DEF CON Speaker
Recognized security researcher and frequent speaker at the industry's premier offensive-security conferences.
Founder, Autonomous-Driving Security Contests
Pioneer of structured adversarial-testing programs for AI-driven systems.
Contact

How would your agents score?

If you don't know, that's the point. Start with a 30-minute exploratory call — or a Gauge pilot on one of your agents.

Get in touch

or email [email protected]