
Security News
arXiv Is Rate Limiting Authors Following a Flood of AI Slop Submissions
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.
falsify-skill
Advanced tools
The scientific thinking protocol for AI agents. Falsify before you believe. 像一流科学家一样思考:先证伪,再相信。
The scientific thinking protocol for AI agents. Falsify before you believe.
像一流科学家一样思考:先证伪,再相信;先标不确定,再下结论。
falsify is a single-Markdown skill that installs a 5-stage scientific thinking protocol on any AI agent (Codex, Claude Code, DeepSeek Harness, Cursor, Gemini CLI, …). It stops the agent from giving confident answers it cannot falsify.
The Iron Law:
NO VERDICT WITHOUT A FALSIFIABLE HYPOTHESIS.
没有可证伪的假设,就没有结论。
falsify is distilled from 70+ community sources and backed by academic work on how agents should reason:
Copy/paste into your CLI prompt (works for any agent that supports skills):
Install the falsify skill from https://github.com/263311487-ux/falsify, refer to the repo's AGENTS.md for instructions.
Or with the skills CLI:
npx skills add 263311487-ux/falsify
Or manually: clone the repo and copy SKILL.md into your agent's skills directory
(~/.codex/skills/falsify/, ~/.claude/skills/falsify/, .cursor/skills/falsify/, …).
| Before (typical agent) | After (falsify) | |
|---|---|---|
| Architecture question | Confident pro/con list → "Redis is a great fit" | Axioms → assumptions flagged → "I am 40% sure, because we have no volume data; cheapest first step is measuring, not adding Redis" |
| Bug diagnosis | "Probably a memory leak" | Hypothesis → adversarial check (deploy window? coincidence?) → evidence → calibrated verdict + residual risk |
| Data claim | "Yes, X is 5x faster" | Demands benchmark definition → labels claim hearsay if unverifiable → refuses to state it as fact |
| "Is this the best approach?" | Answers "yes, it's best" | Rewrites "best" as unfalsifiable → answers "best for [criteria] under [constraints]" |
The five stages (SKILL.md is the full protocol):
公理化 Axiomatize → separate axioms / assumptions / hearsay
假设化 Hypothesize → if [H] then we observe [O]; if [¬O], H is dead
对抗 Adversarialize → steelman the opponent, attack yourself first
验证 Verify → hunt disconfirming evidence, grade it, run the cheapest test
收束 Converge → calibrated verdict, remaining unknowns, lesson to the ledger
references/mental-models.md.templates/thinking-ledger.md) so reasoning is auditable.evals/ ships 28 cases + rubric so you can verify the skill changes behavior.See evals/cases.md and evals/rubric.md. Threshold: pass = 12/18 with no violation of the Iron Law.
Real-community cross-validation (external dogfood) is documented in evals/dogfood-external-20260827.md: 4 real questions from GitHub issues and Stack Overflow, 4/4 passed, and 3/3 cases with a known ground truth matched reality.
The best coding agents are already excellent at producing answers. They are less good at not believing their own answers. falsify borrows the only epistemology that has a 400-year track record of not lying to itself — the scientific method — and turns it into five stages an agent can actually run.
Built on a simple inheritance: 公理 → 假设 → 对抗 → 验证 → 收束. Axiom → Hypothesis → Adversarialize → Verify → Converge.
MIT. See LICENSE.
FAQs
The scientific thinking protocol for AI agents — falsify before you believe. A heuristic coach and five-stage skill (axioms → hypothesis → adversarial → verify → converge), with dated, source-linked historical eval reports; it is guidance, not a guarantee
The npm package falsify-skill receives a total of 4 weekly downloads. As such, falsify-skill popularity was classified as not popular.
We found that falsify-skill demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.

Research
/Security News
A new GhostAction wave hits hundreds of GitHub repos, expanding CI/CD secret theft to cloud and AI credentials in source code and git history.

Research
/Security News
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.