
Security News
White House Authorizes Private Companies to Conduct Offensive Cyber Operations
A new federal program will let vetted U.S. cybersecurity firms help investigate and disrupt foreign cybercrime groups under government direction.
Binary evals, trace-centric, error-analysis-first CLI for LLM evaluation. Built on Hamel Husain's principles.
Binary evals. Trace-centric. Error-analysis-first.
A command-line tool for evaluating LLM outputs using binary pass/fail judgments, built on Hamel Husain's evaluation principles.
📖 Full LLM Guide | 🚀 Quick Start | 📊 GitHub Pages
Most teams struggle with LLM evaluation because they:
EmbedEval fixes this with Hamel Husain's proven approach:
Option 1: Quick Install (Recommended)
curl -fsSL https://raw.githubusercontent.com/Algiras/embedeval/main/install.sh | bash
Option 2: npm (Global)
npm install -g embedeval
Option 3: npx (No Install)
npx embedeval <command>
# 1. COLLECT - Import your LLM traces
embedeval collect ./production-logs.jsonl --output traces.jsonl
# 2. ANNOTATE - Manual error analysis (30 min for 50-100 traces)
embedeval annotate traces.jsonl --user "expert@company.com"
# 3. TAXONOMY - Build failure taxonomy
embedeval taxonomy build --annotations annotations.jsonl
That's it. You'll now see:
Pass Rate: 73%
Top Failure Categories:
1. Hallucination: 12 traces (44%)
2. Incomplete: 8 traces (30%)
3. Wrong Format: 5 traces (19%)
If you're developing or contributing to EmbedEval, use the embedeval-dev script:
# Clone the repository
git clone https://github.com/Algiras/embedeval.git
cd embedeval
# Check your dev environment
./embedeval-dev --doctor
# Install dependencies
./embedeval-dev --install-deps
# Build TypeScript
./embedeval-dev --build
# Run CLI commands (no global install needed)
./embedeval-dev collect examples/v2/sample-traces.jsonl
./embedeval-dev view test-traces.jsonl
./embedeval-dev annotate test-traces.jsonl --user "dev@local"
# Development utilities
./embedeval-dev --watch # Watch mode for auto-rebuild
./embedeval-dev --test # Run test suite
./embedeval-dev --lint # Run ESLint
./embedeval-dev --types # TypeScript type check
./embedeval-dev --clean # Clean build artifacts
# Import traces from JSONL
embedeval collect ./logs.jsonl --output traces.jsonl
# Interactive annotation (p=pass, f=fail, s=save)
embedeval annotate traces.jsonl --user "pm@company.com"
# Read-only viewer
embedeval view traces.jsonl
# Build failure taxonomy
embedeval taxonomy build --user "pm@company.com"
# Display taxonomy
embedeval taxonomy show
# Add evaluator (interactive wizard)
embedeval eval add
# List evaluators
embedeval eval list
# Run evaluations
embedeval eval run traces.jsonl --config evals.yaml
# Generate report
embedeval eval report --results results.jsonl
# Create dimensions template
embedeval generate init
# Generate synthetic traces
embedeval generate create --dimensions dims.yaml --count 50
# Export to Jupyter notebook
embedeval export traces.jsonl --format notebook
# Generate HTML dashboard
embedeval report --traces traces.jsonl --annotations annotations.jsonl
# GOOD: Clear, fast decisions
evals:
- name: is_accurate
type: llm-judge
binary: true # Only PASS or FAIL
# BAD: Never do this
evals:
- name: quality_score
type: 1_to_5 # Creates disagreement
# Spend 60-80% of time here:
embedeval annotate traces.jsonl --user "expert@company.com"
# NOT here (automate only after understanding):
# embedeval eval run traces.jsonl # (do this AFTER annotation)
evals:
# Run cheap evals first
- name: has_content
type: assertion
check: "response.length > 100"
priority: cheap
# Expensive evals only for complex cases
- name: factual_accuracy
type: llm-judge
priority: expensive
# One "benevolent dictator" owns quality:
embedeval annotate traces.jsonl --user "product-manager@company.com"
# Not multiple people voting (causes conflict)
npm install -g embedeval
npx embedeval collect ./logs.jsonl
git clone https://github.com/Algiras/embedeval.git
cd embedeval
npm install
npm run build
npm link # Makes 'embedeval' command available globally
One JSON object per line:
{"id": "trace-001", "timestamp": "2026-01-30T10:00:00Z", "query": "What's your refund policy?", "response": "We offer full refunds within 30 days...", "metadata": {"provider": "google", "model": "gemini-1.5-flash", "latency": 180}}
{"id": "ann-001", "traceId": "trace-001", "annotator": "pm@company.com", "timestamp": "2026-01-30T10:05:00Z", "label": "fail", "failureCategory": "hallucination", "notes": "Made up refund time limit"}
evals:
- id: has_content
type: assertion
priority: cheap
config:
check: "response.length > 50"
- id: accurate
type: llm-judge
priority: expensive
config:
model: gemini-1.5-flash
prompt: "PASS or FAIL: Is this accurate?"
binary: true
# Collect week's traces
embedeval collect ./logs/week-$(date +%Y-%m-%d).jsonl
# Sample 100 for annotation
head -n 100 traces.jsonl > sample.jsonl
# Annotate
embedeval annotate sample.jsonl --user "pm@company.com"
# Build/update taxonomy
embedeval taxonomy update
# Run all evals
embedeval eval run traces.jsonl --config evals.yaml
# Generate report
embedeval report --traces traces.jsonl --annotations annotations.jsonl
# 1. Build taxonomy to see top failures
embedeval taxonomy build
# 2. Add eval for top category (e.g., hallucination)
embedeval eval add
# Interactive wizard asks for type, model, prompt
# 3. Run the new eval
embedeval eval run traces.jsonl --config evals.yaml
# 1. Create dimensions file
embedeval generate init
# Edit dimensions.yaml to define test scenarios
# 2. Generate synthetic traces
embedeval generate create -d dimensions.yaml -n 50
# 3. Run your system on synthetic queries
# (Implementation depends on your system)
# 4. Evaluate
embedeval annotate synthetic-traces.jsonl --user "tester@company.com"
For Claude, Cursor, or other MCP clients:
{
"mcpServers": {
"embedeval": {
"command": "npx",
"args": ["embedeval", "mcp-server"],
"env": {
"GEMINI_API_KEY": "your-api-key"
}
}
}
}
See LLM.md for detailed agent usage guide.
Already configured. Site updates automatically on push to main.
GitHub Actions workflow included. Runs on every PR:
v1 had 88 files with complex A/B testing, genetic algorithms, and BullMQ queues.
v2 has ~20 files with a simple philosophy: look at your traces first.
Before: Infrastructure-heavy, hard to understand
After: Simple CLI, clear workflow, Hamel Husain principles
MIT
Built with ❤️ following Hamel Husain's principles. The goal is understanding failures, not perfect metrics. Spend time looking at traces! 👀
FAQs
Binary evals, trace-centric, error-analysis-first CLI for LLM evaluation. Built on Hamel Husain's principles.
The npm package embedeval receives a total of 6 weekly downloads. As such, embedeval popularity was classified as not popular.
We found that embedeval demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Security News
A new federal program will let vetted U.S. cybersecurity firms help investigate and disrupt foreign cybercrime groups under government direction.

Research
/Security News
The campaign amassed more than 75,000 installs by targeting Russian-speaking users seeking access to blocked services.

Company News
Open source maintainers are under more pressure than ever. We're raising our open source program from the Team plan to the Business plan, free.