
Security News
Ruby's Bundler 4.0.18 Extends Cooldown to bundle lock and bundle cache
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.
@pseolint/core
Advanced tools
Programmatic SEO audit engine — 32 rules across 4 categories (integrity, discoverability, citation, data) for SpamBrain risk + AI Overview citability. v0.4 verdict ladder + site classifier.
Programmatic SEO audit engine — 48+ rules, surfaced per-template, on every monitored release.
The core engine behind pseolint v0.7.5. Use this package to embed pSEO auditing into your own tools, CI pipelines, or SaaS products.
npm install @pseolint/core
import { auditSource } from "@pseolint/core";
const result = await auditSource("https://example.com");
console.log(`Verdict: ${result.verdict}`);
console.log(`Findings: ${result.findings.length}`);
// v0.6: per-template breakdown
for (const t of result.templates) {
console.log(`${t.signature} ${t.verdict} risk=${t.risk}`);
if (t.variance.topDriver) {
const { ruleId, fireRate } = t.variance.topDriver;
const n = Math.round(fireRate * t.auditedUrls.length);
console.log(` ${n}/${t.auditedUrls.length} samples fail ${ruleId}`);
}
}
auditSource accepts a local directory, a single HTML file, a page URL, or a sitemap URL.
48+ rules grouped into 4 scoring super-categories (v0.4): Integrity (spam + content + cannibal, weight 0.50), Discoverability (links + tech, 0.20), Citation (aeo + schema, 0.25), Data (0.05). Source-tree namespaces remain spam/*, aeo/*, etc. for stable rule IDs.
tech/csr-bailout — render-diff rule (requires the --render / render option). Diffs raw server HTML against the post-hydration DOM and flags pages whose substantive content or interactivity exists only after client-side JS — invisible to crawlers and to Google's first indexing pass. High confidence when interactive elements are entirely absent from the server HTML; demoted to info on small marketing sites; a no-op without render.tech/soft-404 — synthetic-URL probe for programmatic directories. Probes one invented, nonexistent URL per template cluster; an HTTP 200 (instead of a 404) means the directory will silently index unbounded junk pages. One probe per cluster, capped, robots-respecting, and fail-open. (The static tech/soft-404 content-pattern check on real pages still runs alongside the probe.)llms.txt presence, AI-crawler access in robots.txt, freshness signals, FAQ coverage, answer-first opener, citable-fact density, content modularity, summary-bait (pages optimized for summarization over retention)title-overlap and keyword-collision were dropped in v0.4 due to high false-positive rates)The v0.7.1/v0.7.2 batch retired the binary thresholds that let a one-page change in crawl size flip a site's verdict, and tightened several rules to validate quality, not mere presence.
BREAKING config rename.
rules.uniqueValueMinWordsis gone — thecontent/unique-valuerule moved from an absolute word count to a rarity-density score, configured viarules.uniqueValueDensity: { passBelow, errorBelow }. Anyone carrying the old key inpseolint.config.*must update it.
spam/boilerplate-ratio, spam/template-diversity, content/value-add, content/wikipedia-paraphrase.schema/required-fields — an empty / whitespace-only / nameless author now counts as missing, not present.schema/json-ld-valid — accepts an @type that is a string or an all-string array (multi-type nodes no longer false-positive).tech/og-completeness — whitespace-only tag values count as missing; severity is graded rather than all-or-nothing.content/eeat-signals — reads the visible contentText (not raw markup), and an about-link only counts when it's same-host.links/orphan-pages & links/cluster-connectivity are suppressed on sampled crawls (a sampled graph can't distinguish a real orphan from sample shape).tech/canonical-consistency & tech/sitemap-completeness normalize URLs and collapse out-of-crawl-scope noise instead of flagging it.aeo/crawler-access honors robots Allow directives (RFC 9309) so an explicit allow no longer reads as a block.schema/consistency compares each page's own @type signature rather than the cluster-wide union, so multi-template sites no longer false-positive.The unit of analysis is now the template, not the URL. When ≥2 template clusters are detected (each with ≥1% URL coverage and ≥5 pages), the engine runs a two-phase pipeline:
clusterUrlTemplates, canonicality-verifies one sample per cluster. Cost: ~T HTTP fetches.RuleResult with its template field, computes per-template verdict + variance.AuditResult.templates is additive — old code reading findings continues to work.
Template typeimport type { Template, TemplateVariance, AuditResult } from "@pseolint/core";
// Each entry in result.templates:
// {
// signature: string; e.g. "/listing/:slug"
// totalUrls: number; cluster size in the sitemap
// auditedUrls: string[]; pages actually fetched in phase 2
// verdict: Verdict;
// risk: number; 0-100, independent of site-level risk
// categories: CategoryGrades;
// variance: {
// ruleFireRates: Record<string, number>; per-rule fraction of samples that fired
// uniformityScore: number; 0-1; high = same problems on every page
// topDriver: { ruleId: string; fireRate: number } | null;
// };
// findingIds: string[]; references into result.findings
// }
siteVerdictFromTemplates helperimport { siteVerdictFromTemplates } from "@pseolint/core";
// Returns the worst verdict among templates with ≥5% URL coverage.
// Falls back to "ready" when no template clears the 5% floor.
const verdict = siteVerdictFromTemplates(result.templates);
import { auditSource } from "@pseolint/core";
const result = await auditSource("https://example.com", {
samplingStrategy: "stratified",
cache: { dir: ".pseolint/cache" },
});
console.log(`Site verdict: ${result.verdict} risk: ${result.risk}`);
for (const t of result.templates) {
const td = t.variance.topDriver;
const driverLine = td
? ` top driver: ${td.ruleId} (${Math.round(td.fireRate * 100)}% of samples)`
: "";
console.log(` ${t.signature} ${t.verdict} uniformity ${Math.round(t.variance.uniformityScore * 100)}%${driverLine}`);
}
Design rationale: docs/superpowers/specs/2026-05-04-pseolint-v0.6-audit-as-template-reframe.md
content/title-uniqueness (raw, not entity-masked — catalog templates with per-record entity values still pass), content/heading-structure (H1 presence, single-H1, hierarchy), content/image-alt-text (skips role="presentation" / aria-hidden="true" / explicit alt=""), tech/og-completeness (the README-promised rule that finally ships).AuditOptions.authorityScore (0-100) — bring-your-own-DA. ≥80 shifts the verdict ladder one tier lenient (established brand can absorb shapes a newer site can't). ≤30 shifts one tier stricter (newer/lower-authority operator). Raw risk number unchanged so CI gates stay stable. The engine itself remains authority-blind by design — no Moz/Ahrefs/Semrush dependency.AuditOptions.contentEffort ({ enabled: boolean; model?: string; cacheDir?: string }, v0.7.3) — opt-in AI content-effort signal. When enabled, an LLM (default claude-sonnet-4-6, override via model) judges a 0-100 content originality/effort score from sampled page text (≤10 pages, content-hash cached under cacheDir or an OS-temp default) that moderates the verdict ±1 tier. Needs ANTHROPIC_API_KEY in the environment; fail-safe no-op (no verdict change) when the key is absent or the call fails. The resolved score is written to summary.contentEffort.score. Page text is passed as untrusted DATA (no URL/domain in the prompt; structured-output, no-tool judge — injection-resistant). The read-only counterpart on the result is AuditSummary.contentEffort ({ score: number }), present only when the signal ran and produced a score.AuditOptions.sampleSeed — deterministic mulberry32 PRNG plumbed through the stratified sampler. Repeated audits with the same seed pick the same pages and produce reproducible verdicts.spam/doorway-pattern cluster collapse — emits in the same pageUrl + relatedUrls[0] shape as spam/near-duplicate and is registered in CLUSTERABLE_RULES. C(N,2) per-pair findings on entity-swap-heavy catalogs collapse into one cluster finding per template-tied group.summary.appliedSeverityDemotions: string[] — engine emits the list of rule IDs whose severity was overridden by the active scoring profile so consumers (formatters, CI) can show which rules got demoted and why. Pass --strict to disable demotions entirely.links/unreachable-from-root skips on partial-sample audits (it can't distinguish real graph isolation from sample-shape).<details> so PR comments don't drown actionable items in 100+ info bullets.The full per-round iteration story (9 calibration rounds against a curated reputable-pSEO corpus) and the trade-offs we accepted are at docs/superpowers/specs/2026-05-03-calibration-against-reputable-pseo.md. The honest blind-spot audit (what we still don't detect, including the domain-authority gap that motivated --authority-score) is at docs/superpowers/specs/2026-05-03-pseolint-blind-spots.md. The dated user-facing methodology summary is at pseolint.dev/methodology.
auditSource(source, options?)Returns an AuditResult with verdict, risk score, category grades, enriched findings, and — since v0.6 — a templates: Template[] array with per-template verdicts and variance metrics. Also carries optional cache / state / AI-triage metadata.
Selected options (see AuditOptions in types.ts for the full surface):
await auditSource("https://example.com/sitemap.xml", {
concurrency: 5,
timeout: 30_000,
sampleSize: 200,
samplingStrategy: "stratified", // or "random"
ignore: ["**/api/**"],
maxFetchBytes: 52_428_800, // 50 MB hard cap per run
cache: { dir: ".pseolint/cache", ttlMs: 7 * 24 * 60 * 60 * 1000 },
state: {
path: ".pseolint/state.json",
mode: "monitoring", // v0.5+: pre-fetch decision matrix; "fresh" forces full re-audit.
// Omit to auto-monitor when prior state exists.
ageFloorDays: 7, // v0.5+: forces refetch on URLs older than N days
exitOnRegression: true,
since: true, // v0.5+ alias for mode: "monitoring" (back-compat)
},
pageGroups: {
blog: { match: "**/blog/**", rules: ["content/*", "spam/*"] },
products: { match: "**/p/**", overrides: { "spam/thin-content": { thinContentMinWords: 200 } } },
},
dataSource: { records: [{ url: "/p/*", data: { price: "$19", stock: 12 } }] },
entityPatterns: [{ placeholder: "[CITY]", pattern: "\\b(NYC|LA|SF)\\b", flags: "gi" }],
ai: { enabled: true, provider: "anthropic", model: "claude-haiku-4-5-20251001", maxCostUsd: 0.1 },
telemetry: { enabled: true, path: ".pseolint/telemetry.jsonl" },
// Safety (v0.3.2–v0.3.3)
safeMode: "saas", // "saas" | "cli" — flips guardSsrf + caps
guardSsrf: true, // DNS-validated SSRF check on every URL
respectRobotsTxt: true, // skip sitemap URLs Disallow'd by target robots.txt
followRedirects: true,
maxCrawlDiscovered: 2000, // hard ceiling on link-discovery fan-out
signal: controller.signal, // AbortSignal — ctrl-C / quota-exhausted cancels cleanly
rules: {
nearDuplicateThreshold: 0.85,
thinContentMinWords: 300,
titleOverlapThreshold: 0.8,
// ...
},
});
@pseolint/core ships a few primitives for hosts that run audits against
user-submitted URLs. All are opt-in; local CLI use doesn't change.
import {
safeFetch, // SSRF-safe fetch for non-audit use cases
validateTargetHost, // throws SSRFError on private-range / DNS-rebinding targets
isPrivateOrReservedHost,
SSRFError,
DnsResolutionError,
} from "@pseolint/core";
// Validate a user-submitted URL before enqueuing:
await validateTargetHost(new URL(userUrl).hostname);
// Fetch with SSRF guard baked in:
const res = await safeFetch(userUrl, { timeoutMs: 10_000, followRedirects: false });
The full audit picks up the same guard via auditSource(url, { safeMode: "saas" })
or via the individual guardSsrf / respectRobotsTxt / followRedirects flags.
Rendered audits (options.render = {...}) block known analytics endpoints
by default so the audit doesn't inject fake sessions into the site owner's
GA / Plausible / PostHog / Mixpanel / Hotjar / Sentry dashboards.
await auditSource(url, {
render: {
analyticsMode: "block", // default — blocks ~40 analytics hosts
// "allow-first-party" — block third-party only
// "allow" — don't intercept anything
extraBlockedHosts: ["my-internal-metrics.corp"],
},
});
import { formatConsole, formatJson, formatMarkdown, formatHtml } from "@pseolint/core";
const out = formatConsole(summary);
const json = formatJson(summary);
const md = formatMarkdown(summary);
const html = formatHtml(summary);
formatJson(summary) (and pseolint --format json) serializes the AuditSummary shape verbatim. This is the surface CI gates and other programmatic consumers (pseolint-gate-style scripts) should read. A machine-readable JSON Schema (draft 2020-12) is published with the package at schemas/audit-summary.schema.json — validate against it instead of hand-rolling shape assumptions.
schemaVersion — read it, branch on it{
"schemaVersion": "2026-06-v0.6", // bumps on EVERY breaking OR additive-public output change
"verdict": "caution", // "ready" | "caution" | "concerning" | "critical"
"risk": 42, // 0-100, lower = better. Never shown to humans; use for CI thresholds.
"headline": "3 ship-blockers, 16 should-fix",
"categories": { /* 5 fixed keys → { grade, issues } */ },
"issues": { /* severity buckets — see below */ },
"templates": [ /* v0.6 per-template breakdown; may be [] or absent */ ],
"pageCount": 150
// ... diagnostics, auditedUrls, truncated, etc.
}
The schemaVersion string carries a YYYY-MM-vX.Y tag (matching the $id of the published schema). It bumps whenever the output shape changes — breaking or additive. Gate scripts should assert the version they were written against (or accept a known range) so a shape change surfaces as a clear failure rather than silent misread.
issues is SEVERITY-bucketed (not a flat array, not category-keyed)This is the most common thing consumers get wrong:
"issues": {
"blockers": [ /* RuleResult — severity error | critical */ ],
"shouldFix": [ /* RuleResult — severity warning */ ],
"informational": [ /* RuleResult — severity info */ ]
}
It is not issues: RuleResult[] and not issues: { aeo: [...], spam: [...] }. To get a single flat list for a gate:
const all = [...summary.issues.blockers, ...summary.issues.shouldFix, ...summary.issues.informational];
const shouldFail = summary.issues.blockers.length > 0; // typical gate
categories (keyed by integrity | discoverability | citation | data | audit) carries grades + counts only — the per-category issue arrays live under issues, bucketed by severity, not under categories.
truncated: true means partial coverageWhen the crawl did not complete (e.g. the backpressure watchdog aborted on a degraded origin), the report sets truncated: true and a human-readable truncatedReason. The report still carries whatever findings were collected, but CI gates MUST treat pageCount, risk, verdict, and all counts as LOWER bounds — a "ready" verdict on a truncated run does not mean the site is clean, only that nothing was found before the abort. Gate on truncated explicitly if a partial pass should not count as a pass.
import Ajv2020 from "ajv/dist/2020.js";
import addFormats from "ajv-formats";
import schema from "@pseolint/core/schemas/audit-summary.schema.json" with { type: "json" };
const ajv = new Ajv2020({ strict: false });
addFormats(ajv);
const validate = ajv.compile(schema);
const report = JSON.parse(stdout); // pseolint --format json
if (!validate(report)) throw new Error(ajv.errorsText(validate.errors));
The schema's additionalProperties is permissive, so it pins the contract consumers gate on without rejecting forward-compatible field additions.
When ai.enabled is set, findings are clustered into root-causes by an LLM. Providers are loaded lazily from optional peer deps — install only the one you need:
npm install @ai-sdk/anthropic # or @ai-sdk/openai, @ai-sdk/google, @ai-sdk/mistral,
# @ai-sdk/groq, @ai-sdk/xai, @ai-sdk/cohere,
# ollama-ai-provider-v2
import { triageFindings, createLanguageModel, estimateCostUsd } from "@pseolint/core";
Cost and daily-budget caps are enforced pre-flight; results are cached on disk by default.
Net-new in v0.5. orchestrate() drives an LLM through 25 deterministic tools (sitemap fetch, template clustering, per-page rule checks, AEO probes against live answer engines, SerpAPI) and produces a fix manifest of concrete patches — not just a list of findings.
import { orchestrate } from "@pseolint/core";
const { session, manifest, validation, diff } = await orchestrate({
domain: "https://example.com",
userId: "demo",
budget: { maxSessionUsd: 3 }, // optional; default $5
onEvent: (e) => console.log(e), // optional; SSE-friendly callback
});
if (session.reason === "completed") {
console.log(`Verdict: ${manifest!.verdict}`);
console.log(`${validation!.validPatches}/${validation!.totalPatches} patches valid`);
}
What you get:
manifest — FixManifest with verdict, category grades, page/template/domain patches (replace_h1, rewrite_meta, add_jsonld, add_faq_block, rewrite_intro, add_internal_link, remove_thin_block, robots_txt, sitemap_xml, canonical_strategy)validation — patch-by-patch ManifestValidationReport. Failed patches are dropped from the manifest before it returns; failures carries the location + reason.diff — ManifestDiff of structured PatchDiff objects (5 kinds — text_replace, html_insert, html_remove, file_replace, guidance) suitable for direct UI rendering.Architecture: rules become tools the LLM calls. The LLM picks order. Budget caps (LLM tokens + external probe USD, pre-flight + reactive) bound spend. Watchdog injects a convergence reminder every N tool calls. AsyncLocalStorage-backed page cache means HTML never travels in conversation history — token cost stays bounded as audits scale.
Lower-level exports for callers who want individual pieces:
import {
runOrchestrator, // direct runner — bring your own LanguageModel
orchestratorTools, // the 25-tool registry
defineTool, // helper to add custom tools
validateManifest, // walk a manifest, return per-failure report
diffManifest, // produce a structured-diff projection
manifestSchema, // Zod schema for FixManifest
buildSystemPrompt, // canonical orchestrator system prompt
DEFAULT_BUDGET, // BudgetCaps defaults
} from "@pseolint/core";
External probes (query_serp, ask_ai_engine) read API keys from the call's apiKey arg or SERPAPI_API_KEY / ANTHROPIC_API_KEY / PERPLEXITY_API_KEY / GOOGLE_GENERATIVE_AI_API_KEY env vars.
When prior state exists, auditSource defaults to monitoring mode: the decision matrix decides which URLs to fetch BEFORE the network round-trip. URLs without change signals are skipped entirely; their findings are carried forward from prior state with carriedForward: true and lastVerifiedAt markers.
import { planScrapeStrategy, CORE_RULESET_VERSION, DEFAULT_AGE_FLOOR_DAYS } from "@pseolint/core";
// The decision matrix is also exposed as a pure function for callers that
// want to plan their own fetches:
const plan = planScrapeStrategy({
candidateUrls,
priorState,
sitemapLastmodByUrl, // Map<url, ISO-string>
currentRulesetVersion: CORE_RULESET_VERSION,
ageFloorDays: DEFAULT_AGE_FLOOR_DAYS,
now: new Date(),
// Optional Pro-only inputs:
// gscDeltasByUrl, gscThresholds
});
// plan.refetch: Map<url, RefetchReason>
// plan.skip: Map<url, "unchanged">
Reasons (first match wins): new → age → ruleset → recheck (warning/error/critical only — info findings carry forward) → lastmod → gsc → no-signal → else unchanged.
AuditSummary.scrapePlan reports { fetched, intended, carriedForward, reasonCounts, rulesetVersion, lastFullAuditAt } — populated only on monitoring runs.
Bump CORE_RULESET_VERSION when shipping a new rule or materially changing rule logic so monitoring runs re-evaluate previously-skipped URLs against the new ruleset.
Regression gating. state.exitOnRegression: true flags a run where a new rule ID fires on any previously clean URL (summary.hasRegression). Carried-forward findings are excluded from the regression baseline so a regression on a skipped URL isn't masked by stale findings.
UrlStateEntry v2 stores full finding records (not just IDs) so future runs can carry them forward. Persists lastModified, etag, sitemapLastmodAtAudit, rulesetVersion per URL. RunState adds lastFullAuditAt and rulesetVersion. Existing v1 state files (v0.4) are discarded on read with a warning, triggering one baseline re-audit.
Setting cache enables an ETag/Last-Modified-aware disk cache for HTTP fetches. summary.cacheStats reports { hits, total, bytesSavedEstimate }.
Classify pages by glob and apply different rule subsets or threshold overrides per group. Results are surfaced in summary.groupScores / summary.groupPageCounts.
For client-rendered pages, install playwright-core and pass render: { browserWsEndpoint } to connect to an existing browser endpoint.
Under the render option the auditor runs renderPages() (Playwright) to execute each sampled page in a headless browser and populate ParsedPage.renderedHtml with the post-hydration DOM, which the render-aware rules (tech/csr-bailout) diff against the raw server HTML. The render pass is Node-only (it does not run under Bun) and needs either a matching Chromium / chromium-headless-shell install (npx playwright install chromium) or a CDP endpoint (render.browserWsEndpoint or the PSEOLINT_BROWSER_WS env var). If no browser is available it degrades gracefully to a static audit — the run continues and render-aware rules are simply skipped.
All AI providers and playwright-core are optional peers — you only install the ones you actually use.
MIT
FAQs
Programmatic SEO audit engine — 32 rules across 4 categories (integrity, discoverability, citation, data) for SpamBrain risk + AI Overview citability. v0.4 verdict ladder + site classifier.
We found that @pseolint/core demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Security News
The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects.

Security News
During a UK cyber test, a Mythos 5 agent used sockpuppets, social engineering, and prompt injection to try to get a maintainer to merge malware.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.