benchmarking — independent software & tools
-
reticle
Browser-native, GPU-accelerated editor for very large hierarchical 2D IC-layout scenes, written in Rust (native and WebAssembly).
-
Eval-Forge
Production-grade open-source LLM evaluation platform with G-Eval, LLM-as-a-Judge, RAG evaluation, and customizable AI evaluation pipelines.
-
SLMarena
SLMArena is a self-hosted benchmarking & red-teaming platform for local SLMs. Measure TTFT, tok/s, & latency while an automated frontier LLM judge evaluates accuracy, grammar, & prompt injection resilience in real time.
-
llmgauge
Practical local LLM evaluation CLI for reproducible GGUF/llama.cpp testing on real hardware.
-
atlas.bench
🏎️ High-precision, multi-command TUI benchmarking tool with statistical rigor and side-by-side comparisons. Part of the Atlas Suite.