inspect-ai
Here are 60 public repositories matching this topic...
Consolidated model evaluation framework for LLM benchmarking with Ollama
-
Updated
Apr 2, 2026 - Python
Benchmark for measuring instrumental-convergence behaviour in tool-using LLM agents
-
Updated
May 9, 2026 - Python
ACL Findings benchmark for measuring LLM sycophancy and correction selectivity
-
Updated
Jul 8, 2026 - Python
Factor(UT): Controlling Untrusted AI Decomposers — AAAI 2026 workshop paper on monitoring untrusted decomposition in code generation workflows.
-
Updated
Jul 31, 2026 - Jupyter Notebook
Release gates for AI agents: replay known incidents and check prompt, model and tool-policy changes before they ship. Runs under Inspect.
-
Updated
Jul 29, 2026 - Python
A lightweight Inspect AI benchmark for obvious public-facing LLM failures.
-
Updated
Jul 29, 2026 - Python
Open-source, MCP-native benchmark for whether AI agents reason reliably over fragmented enterprise knowledge — accuracy, provenance, correct refusal.
-
Updated
Jul 22, 2026 - Python
Real DuckDB Quack infrastructure for multi-agent Werewolf: containerized player nodes, Quack gateway federation, browser runner, and local/hosted LLM evals.
-
Updated
Jul 28, 2026 - JavaScript
Statistically rigorous, batch-first reliability auditing for LLM judges & reward models — clustered-bootstrap CIs, McNemar, BH-FDR, power/MDE, planted-bias validation. No API key required.
-
Updated
Jul 7, 2026 - Python
Long-horizon reliability benchmark for industrial edge agents - do self-improving agents get better or corrupt over 1000+ episodes on constrained hardware?
-
Updated
Jul 27, 2026 - Python
Wellness verification harness for companion AI. Multi-turn adversarial suites grounded in six decades of mental-health research and current clinical standards (988, VERA-MH) and law (SB 243). Point it at any chat endpoint — get an evidence-backed, reproducible report.
-
Updated
Jun 14, 2026 - HTML
Open source pipeline for running evals of chemical security tasks
-
Updated
Jul 28, 2026 - Python
LLM agent that plays the Wikipedia game, built on AISI Inspect
-
Updated
Jul 27, 2026 - Python
Inspect AI task pack for financial agent safety and first public LangGraph adapter for Inspect AI.
-
Updated
May 14, 2026 - Python
Security benchmark for LLM agents & their harnesses — offensive, defensive, secure-coding & prompt-injection-safety challenges on live sandboxed targets, with gradated scoring. Built on Inspect AI.
-
Updated
Jul 4, 2026 - Python
Cross-framework agent regression gate — 9 agent frameworks (OpenAI SDK, PydanticAI, Google ADK, Strands, LangChain, Claude SDK, MAF, CrewAI, Smolagents) evaluated on shared tasks with cross-vendor LLM-as-judge, ordinal pairwise ranking, and multi-axis drift attribution.
-
Updated
Jul 25, 2026 - Python
Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
-
Updated
Jul 30, 2026 - TypeScript
Open-source psychological-safety benchmark for conversational AI: frozen clinical scenarios scored by an LLM-judge panel into a CI pass/fail gate + a clinician-grade diagnostic.
-
Updated
Jul 15, 2026 - Python
Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
-
Updated
Jul 28, 2026 - Python
Improve this page
Add a description, image, and links to the inspect-ai topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the inspect-ai topic, visit your repo's landing page and select "manage topics."