
Research
/Security News
TensorLake npm SDK Compromised in ChainDrop Shai-Hulud Credential-Stealing Attack
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.
@serkanalgur/opencodev2-slim
Advanced tools
Smart context management plugin for OpenCode v2 - semantic compression, cost-aware pruning, adaptive thresholds
Smart context management plugin for OpenCode v2. Optimizes token usage through semantic compression, cost-aware pruning, and adaptive thresholds.
opencode plugin @serkanalgur/opencodev2-slim@latest --global
This installs the plugin globally. The TUI features (panel, slash commands) are automatically loaded when OpenCode starts.
If the CLI command doesn't work, add to your ~/.config/opencode/opencode.json:
{
"plugins": ["@serkanalgur/opencodev2-slim"]
}
After installation, these slash commands are available in the TUI:
/panel — Context window panel in a dialog: message/token breakdown for the
active context, the resolved compression trigger, and live measurements/compress — Sends the assistant a compression instruction (see below)/status — One-line context health report (tokens / limit / %, model, cost)/slim-debug — Toggles debug in ~/.config/opencode/slim.jsonc/panel, /status and /slim-debug are display-only: they render through
context.ui.dialog.alert and never append anything to the session transcript.
/compress is the exception — it is a deliberate instruction to the model.
/panel prints:
session.context, i.e. messages since the last compaction, not session
totals. The panel states this on a Scope: line.Trigger: … tokens (…% of … window) · floor …) — the same
line the panel tool printspanel toolThe panel tool prints a richer, boxed report. Two lines make the numbers
self-explanatory:
Source: measured (server-reported) / Source: estimated (approximate) —
where the headline Context token figure came from (see
Measurement trust). An estimate is a magnitude,
not an exact count.Prune: N outputs · ~X chars (~Y tokens) saved on last request — tool-output
pruning activity from the most recent request. It is a per-request figure
(the plan is re-applied every request), not a cumulative saving, and the line
appears only when pruning is enabled or the last request actually pruned
something.The enhanced compress tool supports multiple modes:
// Auto mode (default) - intelligently selects what to compress
compress({ focus: "old exploration" })
// Range mode - compress specific message range
compress({ focus: "completed tasks", mode: "range", start: 0, end: 50 })
// Topic mode - compress messages matching a topic
compress({ focus: "database work", mode: "topic", topic: "database" })
/compress delivery/compress does not compress anything itself — it writes a short instruction
(the focus/mode/keep-recent you passed) into the transcript so the assistant
calls the compress tool. Because the point is to make the model act now,
the synthetic message is sent with an explicit delivery: "steer":
steer (chosen) — delivered immediately: it interrupts an in-flight turn,
and session.synthetic wakes an idle session (resume stays at its default
true). This matches both the server default (delivery ?? "steer") and the
TUI's own default prompt delivery, so /compress behaves like typing a
message.queue (rejected) — a queued item is only taken when the runner is not
already consuming a turn, so the instruction would wait for the next user
turn instead of compressing now./panel, /status and /slim-debug write nothing to the session.
Create ~/.config/opencode/slim.jsonc:
{
"enabled": true,
"compress": {
"enabled": true,
"permission": "allow",
// Absolute token count (e.g. 200000) or percent of the model context
// window ("80%"). Broken values fall back to the default, never 0.
"maxContextLimit": "80%",
"minContextLimit": "40%",
"nudgeFrequency": 5,
"protectUserMessages": false,
"protectedTools": ["task", "skill", "todowrite", "todoread"]
},
"strategies": {
"deduplication": {
"enabled": true,
"protectedTools": []
},
"purgeErrors": {
"enabled": true,
"turns": 4,
"protectedTools": []
},
// OFF by default. When enabled, replaces the payload of old, large,
// non-protected tool results on the outgoing request only.
"pruneOutputs": {
"enabled": false,
"minChars": 2000,
"maxPerRequest": 50,
"protectedTools": []
},
// Never prune a tool result produced within the last `turns` turns.
"turnProtection": {
"enabled": true,
"turns": 4
}
},
// Measured-vs-estimated token accounting (see "Measurement trust" below).
"usage": {
"trustRatio": 0.5,
"capRatio": 3
},
"adaptive": {
"enabled": true,
"learningRate": 0.1,
"minCompressionRatio": 0.3
},
"costAware": {
"enabled": true,
"cacheBoostFactor": 0.5
},
"persistence": {
"enabled": true,
"directory": "~/.config/opencode/slim"
}
}
compress.maxContextLimit and compress.minContextLimit (and their per-model
overrides compress.modelMaxLimits / compress.modelMinLimits, keyed
"providerId/modelId") accept two forms:
| Form | Example | Meaning |
|---|---|---|
| Number | 200000 | Absolute token count — triggers once the session reaches 200k tokens. |
| Percent string | "80%" | Percentage of the model's context window (80% of 200k = 160k tokens). |
Percent strings accept a decimal comma for locale-typed configs: if the plain
parse fails, every comma is retried as the decimal separator, so "80,5%"
means 80.5% of the window (write absolute counts as JSON numbers — thousands
separators are not supported). A value that still cannot be parsed falls back
to the built-in default with a console warning, deduplicated per
(config key, issue, offending value) — repairing a value and breaking the key
again with a different bad value warns again.
Values that cannot be used (unparsable, negative, or a percent while the model
window is unknown) fall back to the built-in default (100000 / 50000) with a
one-time console warning — never to 0, which would disable triggering. An
absolute threshold above the context window is clamped to that window so it can
still fire. Both panel surfaces — the panel tool and the TUI /panel dialog —
show the resolved trigger as a token count and as its percentage of the window
(Trigger: … tokens (…% of … window) · floor …), with the percentage omitted
when the window is unknown.
What the threshold is actually compared against. maxContextLimit /
minContextLimit are compared against the size of the outgoing prompt we are
about to send — the real measured usage of the last completed request when the
provider reported one, otherwise a full-prompt estimate (system prompt, tool
schemas, message text, tool-call inputs and tool results). They are never
compared against the session's lifetime cumulative token counter, which grows
without bound because cache.read re-reads the whole context every turn. The
panel reflects this split explicitly:
Context: … tokens (the bar) is the current prompt size — the figure the
threshold applies to.Lifetime: … tokens · cumulative spend, NOT context size is the separate
lifetime total (Session.Info.tokens), shown only when it differs from the
prompt size. It is a cost statistic, not occupancy.The Context figure also carries a Source: line naming where it came from —
see Measurement trust below.
strategies.pruneOutputs replaces the payload of old, large, non-protected
tool results (e.g. a huge read, grep or bash output) with a short
placeholder on the outgoing request only — session history is never touched.
It is OFF by default (enabled: false): pruning rewrites an earlier part of
the prompt and therefore invalidates the provider's prefix cache, so it costs a
one-time full-price request per newly pruned turn.
| Key | Default | Meaning |
|---|---|---|
pruneOutputs.enabled | false | Master switch. Opt-in; an absent block leaves the prompt untouched. |
pruneOutputs.minChars | 2000 | Minimum serialized size (characters) for an output to be eligible. |
pruneOutputs.maxPerRequest | 50 | At most this many outputs pruned per request. Never splits a turn. |
pruneOutputs.protectedTools | [] | Extra tool names kept, on top of the always-protected set (task, skill, todowrite, todoread, write, edit, …). purgeErrors.protectedTools is honoured here too. |
turnProtection.enabled | true | Keep the most recent turns intact. |
turnProtection.turns | 4 | Number of recent turns never pruned — the working set the model is actively using. |
Because the prune plan is rebuilt and re-applied on every request, any
saving it produces is a per-request figure, not a cumulative or permanent one.
The panel tool states this on its Prune: line (… saved on last request) and
shows it only when pruning is enabled or the last request actually pruned
something; with the default (off) the line is absent.
usage)The trigger merges two independent numbers:
session.step.ended: input + output + reasoning + cache.read + cache.write). Exact for that request, but it describes the previous prompt,
so it goes stale the moment pruning/compression shrinks the next one.Neither is safe alone, so the two are clamped in both directions:
| Key | Default | Meaning |
|---|---|---|
usage.trustRatio | 0.5 | A measurement below this fraction of the estimate is treated as stale and the estimate wins. |
usage.capRatio | 3 | A measurement above this multiple of the estimate is treated as a broken reading and capped at capRatio × estimated. |
Both fields are optional and fall back to the defaults, so an absent block is
safe. The panel's Source: line reports where the panel's own headline
figure came from — measured (server-reported) when the server reported usage
for that request, otherwise estimated (approximate) — so an approximation is
never mistaken for an exact count. It does not replay the merge above: the
trigger applies its own trustRatio/capRatio rules to the measured total
(input + output + reasoning + cache.read + cache.write) against the
outgoing-prompt estimate, a different quantity from the panel's Context figure.
The panel's Context: line is the last request's prompt size
(input + cache.read + cache.write) and its Lifetime: line is the session's
cumulative spend; the two are never mixed, and Source: labels only the
Context figure.
Global limits are only resolved when they are actually needed: a model with a
valid compress.modelMaxLimits / compress.modelMinLimits override never
reads — and never warns about — the global maxContextLimit /
minContextLimit. The fallback chain itself is unchanged and covered by the
test suite: a broken per-model override degrades to the configured global
(not straight to the built-in default), a broken global degrades to the
built-in default, and warnings fire once per (key, issue, value) across
repeated calls.
Unlike simple text truncation, slim analyzes the semantic content of messages and groups related tool calls together. This preserves context while removing redundancy.
Slim considers the cost of tokens when deciding what to compress. It prioritizes compressing expensive operations (like large file reads) while preserving cheap but important context.
The plugin learns from your compression patterns and adjusts thresholds over time. If you tend to need more context, it will compress less aggressively. If you're efficient, it will compress more.
State is saved to disk, so compression history and learning persist across restarts.
| Command | Description |
|---|---|
/panel | Open the Slim TUI panel in a dialog with context usage, stats, and trigger thresholds |
/compress | Send the assistant a compression instruction (delivery: "steer") |
/status | Show a one-line context health report in a dialog (never written to the session) |
/slim-debug | Toggle debug in ~/.config/opencode/slim.jsonc and show the result in a dialog |
Note: Compression is performed by the AI assistant using the compress tool. The slash command provides guidance on usage; it is the only slash command that writes to the session, and it does so with an explicit delivery: "steer" (see Compress Tool).
BREAKING CHANGES
Session.Info.tokens lifetime
cumulative cost counter is no longer used to decide the trigger, so existing
config files with tuned maxContextLimit/minContextLimit can behave
differently and may need re-tuning. In the panel, Context: is the last
request's prompt size (input + cache.read + cache.write) and Lifetime: is
the session's cumulative spend; the two are shown separately. When the
transcript carries no per-turn usage the old figure is still shown, explicitly
labelled as lifetime cumulative./panel, /status and /slim-debug no longer write to the session (no
transcript message); their output is shown in a modal dialog
(context.ui.dialog.alert). Previous versions wrote the output into the
message stream. Only /compress still writes (it is an instruction to the
model) and now passes an explicit delivery: "steer" instead of relying on
the server default.FIXES
client.session.synthetic,
data.session.message.*) and migrated to the OpenCode v2 API.id field, so the old code always found an empty id and
deleted every block. Now uses deterministic content-based key generation plus
an ambiguity lock.session.step.ended) plus
a full-prompt-scoped estimate.String()
produced [object Object].RangeError when above 100% (negative repeat on a full bar).await, so it
always fell back to 200000.80,5%).NEW
maxContextLimit/minContextLimit accept absolute token counts (a bare
number); percentage forms (80%, 80,5%) keep working.strategies.pruneOutputs prunes tool output (default OFF, opt-in) with turn
protection. Auto-compress keys are captured before pruning replaces pruned
messages with clones, so pruning no longer desynchronises the persisted
compression block's anchors from the messages it covered (which left the block
inert and re-triggered every throttle window).usage.trustRatio / usage.capRatio measurement-trust settings.compressionBlocks, nudge anchors, token
measurement).Source: measured/estimated line and prune statistics. The
Source: line describes where the headline figure came from —
measured (server-reported) when the server reported usage, otherwise
estimated (approximate) — so an approximation is never read as an exact
count. The Prune: line reports outputs pruned and characters/tokens saved on
the last request (a per-request figure, since pruning is re-applied every
request; the line is absent while the default-off feature has not pruned
anything)./panel states its scope: the stats come from session.context ("all
messages after the last compaction"), so Messages:/Tokens (est) are window
counts, not session totals, and it now shows the resolved compression trigger
(Trigger: … tokens (…% of … window) · floor …), matching the panel tool.deriveStats counts unknown message types (agent-switched, model-switched,
location-switched, idle, …) as system instead of assistant, keeping
user + assistant + system === total messages.strategies.pruneOutputs (off by default) and
strategies.turnProtection in the config reference, plus the usage
measurement-trust ratios (trustRatio / capRatio), and clarified that
maxContextLimit is compared against the outgoing prompt estimate/measurement
— not the lifetime cumulative counter./panel in the TUI not reflecting real context usage:
Session.Info.tokens + cost + model + context window) via context.client.session.get() and prints them (measured tokens, %, cost, model) at the bottom of the panel, matching what the panel tool reports.context.ui.router.current() instead of a non-existent context.router, so the panel targets the focused session rather than always the first one.context.data.session.message.sync() before reading the transcript so stats aren't computed from an empty/stale cache.ctx.model.default().data.limit.context instead of the hard-coded 200k./panel now reads live server measurements (Session.Info.tokens + cost) via ctx.session.get() and feeds them to buildPanelData, so the headline tokens/percent/cost match what OpenCode's UI reports.solid-js, @opentui/core, @opentui/solid to devDependencies.
The workflow runs npm ci --legacy-peer-deps, which skips peer deps, so loading
@opencode/plugin/tui failed with Cannot find package 'solid-js'./panel output showing User tokens: 0: user/system messages carry their text
on a top-level text field (not inside content), which deriveStats now captures.feat(panel-as-message): /panel and slim-panel now print the context stats as plain
text into the message stream via client.session.synthetic, instead of taking over
OpenCode's own panel UI (session.panel slot + ui.panel.open removed).keymap.provider is missing in the CLI plugin: register the keymap layer inside
an app slot render (where the keymap provider is available) instead of in setup().Cannot find package 'react' when the plugin is loaded from the global npm cache:
add a per-file /** @jsxImportSource @opentui/solid */ pragma to src/tui.tsx so JSX
always compiles against @opentui/solid/jsx-runtimetsconfig.json in the published package so loaders that read jsxImportSource from config pick it upcompaction hook so history actually shrinks (the context hook only affects the outgoing request)compress/panel tools with options.codemode so they appear in agent/codemode environmentstools always equalled zerosession.panel slot (slim-panel / /panel)Breaking Changes: Migrated to OpenCode v2 plugin API.
@opencode-ai/plugin to @opencode/pluginPlugin.define() pattern instead of server functionctx.tool.transform() with JSON Schema inputctx.session.hook()ctx.event.subscribe()MIT
FAQs
Smart context management plugin for OpenCode v2 - semantic compression, cost-aware pruning, adaptive thresholds
The npm package @serkanalgur/opencodev2-slim receives a total of 127 weekly downloads. As such, @serkanalgur/opencodev2-slim popularity was classified as not popular.
We found that @serkanalgur/opencodev2-slim demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Research
/Security News
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.

Research
/Security News
Socket found 16 malicious Firefox extensions designed to steal crypto wallet recovery phrases and private keys using cloned Rabby and OKX interfaces.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.