
Security News
/Company News
Securing the Financial Frontier: How Capital One Uses Socket for Open Source Security
Capital One is partnering with Socket to proactively secure its open source supply chain.
falsify-skill
Advanced tools
The scientific thinking protocol for AI agents — falsify before you believe. 5-stage falsification skill (axioms → hypothesis → adversarial → verify → converge) for Codex, Claude Code, Cursor, Gemini CLI & 20+ agents. Dual-model evals 26/28. 五段式科学思维协议:公理
The scientific thinking protocol for AI agents. Falsify before you believe.
像一流科学家一样思考:先证伪,再相信;先标不确定,再下结论。
falsify is a single-Markdown skill that installs a 5-stage scientific thinking protocol on any AI agent (Codex, Claude Code, DeepSeek Harness, Cursor, Gemini CLI, …). It stops the agent from giving confident answers it cannot falsify.
The Iron Law:
NO VERDICT WITHOUT A FALSIFIABLE HYPOTHESIS.
没有可证伪的假设,就没有结论。
| Before (typical agent) | After (falsify) | |
|---|---|---|
| Architecture question | Confident pro/con list → "Redis is a great fit" | Axioms → assumptions flagged → "I am 40% sure, because we have no volume data; cheapest first step is measuring, not adding Redis" |
| Bug diagnosis | "Probably a memory leak" | Hypothesis → adversarial check (deploy window? coincidence?) → evidence → calibrated verdict + residual risk |
| Data claim | "Yes, X is 5x faster" | Demands benchmark definition → labels claim hearsay if unverifiable → refuses to state it as fact |
| "Is this the best approach?" | Answers "yes, it's best" | Rewrites "best" as unfalsifiable → answers "best for [criteria] under [constraints]" |
Copy/paste into your CLI prompt (works for any agent that supports skills):
Install the falsify skill from https://github.com/263311487-ux/falsify, refer to the repo's AGENTS.md for instructions.
Or with the skills CLI:
npx skills add 263311487-ux/falsify
Or from npm (installs the SKILL.md into Codex and Claude Code skill dirs automatically):
npx falsify-skill
npm i -g falsify-skill && falsify-skill
Or manually: clone the repo and copy SKILL.md into your agent's skills directory
(~/.codex/skills/falsify/, ~/.claude/skills/falsify/, .cursor/skills/falsify/, …).
falsify is distilled from 70+ community sources and backed by academic work on how agents should reason:
The five stages (SKILL.md is the full protocol):
公理化 Axiomatize → separate axioms / assumptions / hearsay
假设化 Hypothesize → if [H] then we observe [O]; if [¬O], H is dead
对抗 Adversarialize → steelman the opponent, attack yourself first
验证 Verify → hunt disconfirming evidence, grade it, run the cheapest test
收束 Converge → calibrated verdict, remaining unknowns, lesson to the ledger
references/mental-models.md.templates/thinking-ledger.md) so reasoning is auditable.evals/ ships 28 cases + rubric so you can verify the skill changes behavior.See evals/cases.md and evals/rubric.md. Threshold: pass = 12/18 with no violation of the Iron Law.
Real-community cross-validation (external dogfood) is documented in evals/dogfood-external-20260827.md: 4 real questions from GitHub issues and Stack Overflow, 4/4 passed, and 3/3 cases with a known ground truth matched reality.
Cross-model proof (2026-08-27, v0.8.3): the full 28-case suite is run on two external DeepSeek models — not our own agent — in both directions:
deepseek-reasoner generator × deepseek-chat judge → 26/28 pass, avg 15.3/18deepseek-chat generator × deepseek-reasoner judge → 26/28 pass, avg 16.4/18The two rounds fail on disjoint cases (reasoner: 3/9; chat: 6/22); each failing case was regenerated manually and produced protocol-compliant output, so the failures are single-run variance, not stable protocol gaps. This round also fixed a real routing gap the reasoner exposed (live incidents must act at ~70% confidence, not run the full protocol) via a mandatory MODE SELECTION gate. Reproduce in one command: DEEPSEEK_API_KEY=... node evals/run_evals.mjs --model deepseek-reasoner. Reports: evals/results/deepseek-reasoner-2026-08-27.md · evals/results/deepseek-chat-2026-08-27-final.md.
Deployment note (reasoner-class models):
reasoning_contentandcontentshare themax_tokensbudget; on very deep debugging questions the reasoner can spend the entire budget on reasoning and return empty content (observed at 6k–16k tokens). Set a generousmax_tokens, add a retry-on-empty policy, or preferdeepseek-chatfor latency-constrained deployments.
The best coding agents are already excellent at producing answers. They are less good at not believing their own answers. falsify borrows the only epistemology that has a 400-year track record of not lying to itself — the scientific method — and turns it into five stages an agent can actually run.
Built on a simple inheritance: 公理 → 假设 → 对抗 → 验证 → 收束. Axiom → Hypothesis → Adversarialize → Verify → Converge.
falsify is one leg of a three-part workflow: think → verify → present.
Install any of them in one command:
npx skills add 263311487-ux/falsify
npx dsh-verify
npx imprint
MIT. See LICENSE.
FAQs
The scientific thinking protocol for AI agents — falsify before you believe. 5-stage falsification skill (axioms → hypothesis → adversarial → verify → converge) for Codex, Claude Code, Cursor, Gemini CLI & 20+ agents. Dual-model evals 26/28. 五段式科学思维协议:公理
The npm package falsify-skill receives a total of 21 weekly downloads. As such, falsify-skill popularity was classified as not popular.
We found that falsify-skill demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.

Security News
Socket CTO Ahmad Nassri discusses how to keep AI agents from bypassing package blocks, limit credential access, and monitor their actions.

Security News
GPT-6 Astra tried to plant malicious code in simulated open source projects using fake GitHub accounts and deceptive PRs during an assigned CTF challenge.