awesome-evals
awesome-evals is A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.. It is ranked #10 on the Stars Radar, in Learning & awesome lists, first seen 1 h ago and shared in 1 post (934 views).
Alternatives to awesome-evals
- awesome-context-engineering — A curated collection of resources, papers, tools, and best practices for Context Engineering in AI agents and Large…
- awesome-ai-guardrails — A curated list of materials on AI guardrails
- ai-engineering-hub — In-depth tutorials on LLMs, RAGs and real-world AI agent applications.
- awesome-agent-observability — Curated list of tools, standards, and platforms for LLM and AI-agent observability: OpenTelemetry GenAI conventions,…
- awesome-jev — A curated list of awesome Jev / TypeSafe System One applications, libraries, and resources.
- t3code — pingdotgg/t3code
awesome-evals in numbers
- Rank on the Stars Radar: #10 of 260
- Shared in 1 post by 1 account: @benchflow-ai
- 934 views on those posts
- First seen 1 h ago, last shared 5 days ago
- Pricing seen by Jev: open source
- Market: Learning & awesome lists
FAQ
What is awesome-evals?
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow. It was first shared on X 1 h ago and is ranked #10 on the Stars Radar.
Is awesome-evals free?
It is open source.
Who shared awesome-evals?
1 account on X, including @benchflow-ai, in 1 post totalling 934 views.
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 260 posts from 1 X accounts over the last 365 days, 260 tools, 10 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-30 15:59 UTC. Full method.