Testing middleware
-
agent-reliability-toolkit
Open-source testing framework for AI agents. Test for the 7 failure modes before production deployment.
-
agentprobe
Test, record, and replay AI agents locally and in CI with a private pytest-style framework
-
awesome-android-app-repositories
⚡ Curated & continuously updated catalog of 1,600+ open-source Android apps, GitHub repositories, developer tools, AI platforms, and utilities automatically synchronized from Telegram.
-
ckdn
ckdn (checkdown): deterministic check runner and log digester for AI-assisted development loops
-
claude-computer
Claude Computer demonstrates AI autonomy in a virtual machine with real-time streaming, research, creation, and exploration. Watch Claude navigate, interact, and learn in real time 🐙
-
CyberAI
AI-powered pentest platform
-
Derivative
Software synthesis engine that turns requirements into verified artifacts — or explicit failure evidence — through execution-grounded validation.
-
edgeverdict
A code review gate that verifies changes by executing tests, not by judging diffs. LLM proposes; a deterministic gate decides.
-
Entroping
AI-native quality governance for API and backend systems
-
harnessrouter
HarnessRouter Community Edition: the self-hosted, Apache-2.0 edition of the unified interface for agent harnesses. Run Codex, Claude Code, Hermes, and more through one API, with sessions, streaming, files, cancellation, and failure handling. Implements the Unified Harness Protocol (UHP), an open standard. Your keys, your infrastructure.
-
iBetaBot
⚡ Real-time Apple beta & RC firmware watchdog — iOS, iPadOS, macOS, tvOS & visionOS builds pushed straight to Telegram the moment they drop. Cron-ready, dependency-light, zero manual checking.
-
lager
Hardware test automation: drive real instruments and embedded targets from your laptop or CI
-
leashd
Safety-first agentic coding framework. Three-layer safety pipeline (sandbox, YAML policies, human-in-the-loop approval) for AI coding agents. Pluggable runtimes (Claude Code, Codex), autonomous task orchestrator, full audit trail.
-
litmux
Build unit tests for AI prompts to compare models, track performance, and prevent regressions.
-
livekit_agent_simulator
Black-box LiveKit voice agent testing: real WebRTC/SIP calls with an AI caller. Catch barge-in, noise, quiet-mic, and latency failures that text pytest misses. Forensic reports + MCP + lks CLI — no agent code changes.
-
llm-drift
LLM drift detector — know within 5 min when GPT-4o, Claude, or Gemini silently changes behaviour. Open source, self-hostable.
-
OcyShield-Framework
Audit Android systems with this modular framework for security testing, payload deployment, and session management.
-
pentest-with-LLM
Automate authorized pentesting with LLMs, combining scanning, RAG, and exploit research for faster target analysis, vuln discovery, and reporting
-
periscope-mcp
Web-app QA, testing & analysis for AI agents — 74 Playwright tools: authenticated flows (2FA/SSO), 25-action E2E workflows, hard assertions, session reports, and accessibility / SEO / GEO / Core-Web-Vitals + Lighthouse audits across a page or a full crawled site. Battle-tested by an AI agent on real apps.
-
pwnagotchi-store
🛒 Simplify Pwnagotchi management with PwnStore, a lightweight CLI package manager for seamless plugin browsing, installing, and updating.
-
stillworks
Take a snapshot of what your code does, then see if your edit moved anything. For code with no tests. Zero dependencies, MIT.
-
suitest
Self-hostable, MCP-native testing platform. Manual test management, deterministic runs, optional AI. Your stack, your LLM, your data.
-
TrashDroid
Automate comprehensive Android app security testing with TrashDroid using adb, drozer, and apktool for a full nine-phase dynamic analysis.
-
XHS_Business_Idea_Validator
📊 Analyze 小红书 data to uncover market opportunities and user needs, generating automated reports for business idea validation.
-
API-Header-Spoofer
-
GitHub-Cyber-Scanner-Pro