All tags
Catalogue tag

#ai-safety

Every repository in the public record carrying this tag — from its GitHub topics or the keywords its package registries publish. Health is measured under the same versioned methodology as the rest of the record.

42 records
Tagged “ai-safety”Ranked by health index
npm · Go · crates.io
96Exceptionalhealth index
microsoft/agent-governance-toolkit
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
Python★ 6,138↓ 14.7K/moAug 28, 2026
MITAug 28, 2026 · metrics 2.10.0
PyPI · npm
94Exceptionalhealth index
sattyamjjain/agent-audit-kit
Static scanner for MCP-connected AI agent pipelines — 271 rules across 12 categories, 12 compliance frameworks, OWASP Agentic 10/10 + MCP 10/10, GitHub Action, SARIF, public CVE-to-rule ledger.
Python★ 13↓ 2,808/moAug 2, 2026
MITAug 2, 2026 · metrics 2.10.0
PyPI · RubyGems
90Excellenthealth index
OWASP/www-project-agent-memory-guard
OWASP Foundation web repository
Python★ 155↓ 2,902/moAug 25, 2026
Apache-2.0Aug 25, 2026 · metrics 2.10.0
PyPI · Go · npm
90Excellenthealth index
cordum-io/cordum
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.
Go · TypeScript · Python★ 494↓ 17/moAug 9, 2026
Custom licenseAug 9, 2026 · metrics 2.10.0
npm · PyPI
86Excellenthealth index
issdandavis/SCBE-AETHERMOORE
Geometric AI governance and evaluation framework with a 14-layer security pipeline, semantic projection, and reproducible benchmark lanes.
Python · TypeScript★ 6↓ 4,607/moAug 26, 2026
MITAug 26, 2026 · metrics 2.10.0
npm
84Excellenthealth index
JKHeadley/instar
Persistent Claude Code agents with scheduling, sessions, memory, and Telegram.
TypeScript★ 77↓ 157.9K/moSep 5, 2026
MITSep 5, 2026 · metrics 2.10.0
npm
80Excellenthealth index
lua-ai-global/governance
Zero-dependency TypeScript SDK for AI agent governance: policy enforcement, injection detection, tamper-evident audit, and standards mapping (EU AI Act, OWASP, NIST, ISO 42001)
TypeScript★ 25↓ 3,545/moJul 26, 2026
MITJul 26, 2026 · metrics 2.10.0
npm
78Goodhealth index
node9-ai/node9-proxy
The Execution Security Layer for the Agentic Era. Providing deterministic "Sudo" governance and audit logs for autonomous AI agents.
TypeScript★ 209↓ 9,089/moJul 22, 2026
Custom licenseJul 22, 2026 · metrics 2.10.0
npm
77Goodhealth index
IgorGanapolsky/ThumbGate
ThumbGate Pre-Action Checks derive rules from repeated failures, flag risky tool calls, hard-block detected secret leaks, and block matches in strict mode.
JavaScript · HTML★ 24↓ 4,611/moJul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
PyPI
77Goodhealth index
alizahidraja/isnad
Grade every agent, scraper and model in a claim's chain — provenance, trust scoring and audit evidence for LLM pipelines
Python★ 37↓ 4,204/moAug 29, 2026
Apache-2.0Aug 29, 2026 · metrics 2.10.0
npm
75Goodhealth index
bookedsolidtech/rea
Zero-trust governance layer for Claude Code. Policy-enforced MCP gateway with autonomy controls, middleware chain, audit log, HALT kill-switch, and prompt-injection defense.
TypeScript · Shell★ 0↓ 1,950/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
PyPI · npm
75Goodhealth index
gautamvarmadatla/mcpsafetywarden
MCP servers expose tools with no information about what they actually do at runtime. mcpsafetywarden sits between your agent and any MCP server, profiling tool behavior, blocking destructive calls, and running active security audits before you trust them in a workflow.
Python★ 9↓ 2,923/moAug 1, 2026
Custom licenseAug 1, 2026 · metrics 2.10.0
npm
73Goodhealth index
Keesan12/martin-loop
Make AI coding agents safe to scale autonomously: assign work, cap spend, enforce policy, verify output, roll back failures, learn from loops, and prove ROI across every repo.
TypeScript · JavaScript★ 38↓ 3,459/moJul 19, 2026
Apache-2.0Jul 19, 2026 · metrics 2.10.0
PyPI · npm
73Goodhealth index
aiexponenthq/license-compliance-checker
License Compliance Checker — Multi-ecosystem license + AI model scanner for EU AI Act Article 53 GPAI compliance. SBOM, SARIF, training-data risk. Apache 2.0.
Python · TypeScript★ 1↓ 258/moJul 27, 2026
Apache-2.0Jul 27, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
auspexai/worker
AuspexAI volunteer worker for donating compute to AI research
Python★ 0↓ 628/moAug 22, 2026
AGPL-3.0Aug 22, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
b7n0de/proofbundle
Offline cryptographic receipts for AI evaluation results — Ed25519 + RFC 6962 Merkle + optional SD-JWT. Integrity, not truth
Python★ 2↓ 6,574/moJul 23, 2026
MITJul 23, 2026 · metrics 2.10.0
Go
73Goodhealth index
gumieri/nenya
A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstream LLM providers (Gemini, DeepSeek, Zhipu z.ai).
Go★ 25Jul 24, 2026
Apache-2.0Jul 24, 2026 · metrics 2.10.0
PyPI
73Goodhealth index
philpaz/recusal
Deterministic governance for Claude and MCP tool calls. Pin approved capabilities, detect drift, and refuse unsafe or unapproved actions before execution. No model in the decision path.
Python★ 3↓ 2,854/moJul 27, 2026
Apache-2.0Jul 27, 2026 · metrics 2.10.0
npm
71Goodhealth index
alexandriashai/cbrowser
Cognitive Browser: The browser automation that thinks. Constitutional safety • Persona UX testing • Natural language interface • Self-healing selectors • Built for AI agents
TypeScript · JavaScript★ 18↓ 2,900/moJul 17, 2026
Custom licenseJul 17, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
auspexai/tenant-sdk
SDK for authoring research tenants on AuspexAI
Python★ 0↓ 9,632/moJul 16, 2026
Apache-2.0Jul 16, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
crucible-security/crucible
pytest for AI agents - Autonomous red-teaming, behavioral monitoring & security testing for LLM agents
Python★ 45↓ 1,793/moJul 17, 2026
Apache-2.0Jul 17, 2026 · metrics 2.10.0
PyPI
69Goodhealth index
fathom-lab/styxx
Verification for the agent era. Your coding agent's PR summary cannot lie about its diff - one CI line. Plus the research: the first map of which AI minds can read each other. Every claim machine-verified against committed receipts, negatives included. pip install styxx
Python★ 14↓ 2,231/moAug 19, 2026
MITAug 19, 2026 · metrics 2.10.0
PyPI
67Goodhealth index
CognitiveThoughtEngine/constitutional-agent-governance
The WHY layer for AI agents: six constitutional gates + 12 hard constraints, plus cross-session risk composition — catches agents that pass every individual gate but accumulate risk across a sequence (stateless engines can't). pip install constitutional-agent
Python★ 0↓ 331/moJul 27, 2026
MITJul 27, 2026 · metrics 2.10.0
Go
67Goodhealth index
nikicat/secrets-dispatcher
Per-operation approval and audit logging for secret access and git commit signing on Linux
Go · TypeScript★ 8Jul 19, 2026
MITJul 19, 2026 · metrics 2.10.0
crates.io
63Moderatehealth index
kahramanemir/Vallum
Security boundary CLI proxy between AI coding agents and your shell — redacts secrets, neutralizes prompt injection, wraps untrusted terminal output, and audits every command. Single Rust binary.
Rust★ 5↓ 159/moSep 6, 2026
Apache-2.0Sep 6, 2026 · metrics 2.10.0
PyPI
62Moderatehealth index
sunglasses-dev/sunglasses
Sunglasses for AI agents. Protection layer + neighborhood watch.
Python★ 4↓ 2,911/moAug 18, 2026
MITAug 18, 2026 · metrics 2.10.0
Go
62Moderatehealth index
trustabl/trustabl
Static analyzer for agent reliability.
Go★ 22Jul 21, 2026
Apache-2.0Jul 21, 2026 · metrics 2.10.0
PyPI
62Moderatehealth index
veldtlabs/veldt-kya
KYA (Know Your Agents) — Open-source trust, governance, and evidentiary assurance infrastructure for autonomous systems. Built on KYP (Know Your Principal), a unified trust model for human users, AI agents, service accounts, and machine identities.
Python★ 2↓ 5,683/moJul 16, 2026
Custom licenseJul 16, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
YugantM/hvtracker
AI Agent Trust Registry — independent, evidence-based trust scores for 300+ open-source AI agents. Runtime-trust calibrated (MCP, dependencies, provenance drift), with side-by-side comparison and embeddable live badges. Ranked by verifiable signals, not hype.
HTML★ 5Jul 26, 2026
MITJul 26, 2026 · metrics 2.10.0
PyPI
60Moderatehealth index
haqaliz/belay
The agent harness: sandbox any agent, verify each step by replaying it against real state, and keep a deterministic trace.
Python★ 0↓ 2,022/moAug 20, 2026
Apache-2.0Aug 20, 2026 · metrics 2.10.0