
Research
/Security News
TensorLake npm SDK Compromised in ChainDrop Shai-Hulud Credential-Stealing Attack
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.
@serkanalgur/opencodev2-slim
Advanced tools
Smart context management plugin for OpenCode v2 - semantic compression, cost-aware pruning, adaptive thresholds
Smart context management plugin for OpenCode v2 — semantic compression, cost-aware pruning, adaptive thresholds
Installation • Usage • Features • Configuration • How It Works • Commands • Changelog
opencode plugin @serkanalgur/opencodev2-slim@latest --global
This installs the plugin globally. The TUI features (panel, slash commands) are automatically loaded when OpenCode starts.
If the CLI command doesn't work, add to your ~/.config/opencode/opencode.json:
{
"plugins": ["@serkanalgur/opencodev2-slim"]
}
After installation, these slash commands are available in the TUI:
/panel — Context window panel in a dialog: message/token breakdown for the
active context, the resolved compression trigger, and live measurements/compress — Sends the assistant a compression instruction (see below)/status — One-line context health report (tokens / limit / %, model, cost)/slim-debug — Toggles debug in ~/.config/opencode/slim.jsonc/panel, /status and /slim-debug are display-only: they render through
context.ui.dialog.alert and never append anything to the session transcript.
/compress is the exception — it is a deliberate instruction to the model.
/panel prints:
session.context, i.e. messages since the last compaction, not session
totals. The panel states this on a Scope: line.Trigger: … tokens (…% of … window) · floor …) — the same
line the panel tool printspanel toolThe panel tool prints a richer, boxed report. Two lines make the numbers
self-explanatory:
Source: measured (server-reported) / Source: estimated (approximate) —
where the headline Context token figure came from (see
Measurement trust). An estimate is a magnitude,
not an exact count.Prune: N outputs · ~X chars (~Y tokens) saved on last request — tool-output
pruning activity from the most recent request. It is a per-request figure
(the plan is re-applied every request), not a cumulative saving, and the line
appears only when pruning is enabled or the last request actually pruned
something.The enhanced compress tool supports multiple modes:
// Auto mode (default) - intelligently selects what to compress
compress({ focus: "old exploration" })
// Range mode - compress specific message range
compress({ focus: "completed tasks", mode: "range", start: 0, end: 50 })
// Topic mode - compress messages matching a topic
compress({ focus: "database work", mode: "topic", topic: "database" })
/compress delivery/compress does not compress anything itself — it writes a short instruction
(the focus/mode/keep-recent you passed) into the transcript so the assistant
calls the compress tool. Because the point is to make the model act now,
the synthetic message is sent with an explicit delivery: "steer":
steer (chosen) — delivered immediately: it interrupts an in-flight turn,
and session.synthetic wakes an idle session (resume stays at its default
true). This matches both the server default (delivery ?? "steer") and the
TUI's own default prompt delivery, so /compress behaves like typing a
message.queue (rejected) — a queued item is only taken when the runner is not
already consuming a turn, so the instruction would wait for the next user
turn instead of compressing now./panel, /status and /slim-debug write nothing to the session.
Create ~/.config/opencode/slim.jsonc:
{
"enabled": true,
"compress": {
"enabled": true,
"permission": "allow",
// Absolute token count (e.g. 200000) or percent of the model context
// window ("80%"). Broken values fall back to the default, never 0.
"maxContextLimit": "80%",
"minContextLimit": "40%",
"nudgeFrequency": 5,
"protectUserMessages": false,
"protectedTools": ["task", "skill", "todowrite", "todoread"]
},
"strategies": {
"deduplication": {
"enabled": true,
"protectedTools": []
},
"purgeErrors": {
// OFF by default. In 3.0.2 this was configured as `true` but
// matched nothing, because it looked for the pairing id on
// `toolCallID`/`callID` while the id lives on `part.id` — see
// "Purge-errors migration" below before turning it on.
"enabled": false,
"turns": 4,
"protectedTools": []
},
// OFF by default. When enabled, replaces the payload of old, large,
// non-protected tool results on the outgoing request only.
"pruneOutputs": {
"enabled": false,
"minChars": 2000,
"maxPerRequest": 50,
"protectedTools": []
},
// Never prune a tool result produced within the last `turns` turns.
"turnProtection": {
"enabled": true,
"turns": 4
},
// ON by default, escape hatch only. Set to false only if you have a
// reason to: it lets a range split a tool pair again, and the provider
// will reject the next request with `invalid_request_error`.
"guardToolPairs": true
},
// Measured-vs-estimated token accounting (see "Measurement trust" below).
"usage": {
"trustRatio": 0.5,
"capRatio": 3
},
"adaptive": {
"enabled": true,
"learningRate": 0.1,
"minCompressionRatio": 0.3
},
"costAware": {
"enabled": true,
"cacheBoostFactor": 0.5
},
"persistence": {
"enabled": true,
"directory": "~/.config/opencode/slim"
}
}
compress.maxContextLimit and compress.minContextLimit (and their per-model
overrides compress.modelMaxLimits / compress.modelMinLimits, keyed
"providerId/modelId") accept two forms:
| Form | Example | Meaning |
|---|---|---|
| Number | 200000 | Absolute token count — triggers once the session reaches 200k tokens. |
| Percent string | "80%" | Percentage of the model's context window (80% of 200k = 160k tokens). |
Percent strings accept a decimal comma for locale-typed configs: if the plain
parse fails, every comma is retried as the decimal separator, so "80,5%"
means 80.5% of the window (write absolute counts as JSON numbers — thousands
separators are not supported). A value that still cannot be parsed falls back
to the built-in default with a console warning, deduplicated per
(config key, issue, offending value) — repairing a value and breaking the key
again with a different bad value warns again.
Values that cannot be used (unparsable, negative, or a percent while the model
window is unknown) fall back to the built-in default (100000 / 50000) with a
one-time console warning — never to 0, which would disable triggering. An
absolute threshold above the context window is clamped to that window so it can
still fire. Both panel surfaces — the panel tool and the TUI /panel dialog —
show the resolved trigger as a token count and as its percentage of the window
(Trigger: … tokens (…% of … window) · floor …), with the percentage omitted
when the window is unknown.
What the threshold is actually compared against. maxContextLimit /
minContextLimit are compared against the size of the outgoing prompt we are
about to send — the real measured usage of the last completed request when the
provider reported one, otherwise a full-prompt estimate (system prompt, tool
schemas, message text, tool-call inputs and tool results). They are never
compared against the session's lifetime cumulative token counter, which grows
without bound because cache.read re-reads the whole context every turn. The
panel reflects this split explicitly:
Context: … tokens (the bar) is the current prompt size — the figure the
threshold applies to.Lifetime: … tokens · cumulative spend, NOT context size is the separate
lifetime total (Session.Info.tokens), shown only when it differs from the
prompt size. It is a cost statistic, not occupancy.The Context figure also carries a Source: line naming where it came from —
see Measurement trust below.
strategies.pruneOutputs replaces the payload of old, large, non-protected
tool results (e.g. a huge read, grep or bash output) with a short
placeholder on the outgoing request only — session history is never touched.
It is OFF by default (enabled: false): pruning rewrites an earlier part of
the prompt and therefore invalidates the provider's prefix cache, so it costs a
one-time full-price request per newly pruned turn.
| Key | Default | Meaning |
|---|---|---|
pruneOutputs.enabled | false | Master switch. Opt-in; an absent block leaves the prompt untouched. |
pruneOutputs.minChars | 2000 | Minimum serialized size (characters) for an output to be eligible. |
pruneOutputs.maxPerRequest | 50 | At most this many outputs pruned per request. Never splits a turn. |
pruneOutputs.protectedTools | [] | Extra tool names kept, on top of the always-protected set (task, skill, todowrite, todoread, write, edit, …). purgeErrors.protectedTools is honoured here too. |
turnProtection.enabled | true | Keep the most recent turns intact. |
turnProtection.turns | 4 | Number of recent turns never pruned — the working set the model is actively using. |
Because the prune plan is rebuilt and re-applied on every request, any
saving it produces is a per-request figure, not a cumulative or permanent one.
The panel tool states this on its Prune: line (… saved on last request) and
shows it only when pruning is enabled or the last request actually pruned
something; with the default (off) the line is absent.
strategies.purgeErrors used to default to enabled: true, and if you never
wrote the key into slim.jsonc you were running it. It did nothing. The
strategy matches an errored tool result to its call, and it was reading the
pairing id from toolCallID / callID — fields that do not exist on the v2
message shape, where the id lives on part.id. Every lookup came back empty, so
the purge never fired, on any request, for any session.
The lookup is fixed, which means the strategy now does what its name says: for
a tool call whose result errored, and which is at least turns positions behind
the end of the conversation, every string value longer than 80 characters in its
input is replaced with [input removed due to failed tool call]. The error
itself and the rest of the input are kept.
Because that is a visible change to the outgoing prompt from a path that had
never executed, the default is now enabled: false.
purgeErrors.enabled: true explicitly. It now works. Expect the
rewritten inputs described above on the first request after upgrading, and
check that losing those long input values does not break anything you rely on
for error recovery — a failed call is exactly the one you may want to re-read.| Key | Default | Meaning |
|---|---|---|
purgeErrors.enabled | false | Master switch. Opt-in after 3.0.2. |
purgeErrors.turns | 4 | A call is only purged once it is at least this many messages from the end. |
purgeErrors.protectedTools | [] | Tool names exempt from the purge. Also honoured by pruneOutputs.protectedTools. |
Compression blocks and deduplication both drop whole messages, which can
split a tool call from its result. The host repairs one direction of that split
(it synthesises Tool result missing for a surviving call) but not the other: a
surviving role:"tool" result whose call is gone is emitted with an orphan
tool_call_id and the next request fails with [invalid_request_error] invalid request.
strategies.guardToolPairs therefore forbids removing a tool-call whose
tool-result is not removed by the same pass. It is ON by default; set it
to false only as an escape hatch, because turning it off on a range that
splits a pair restores the 400. A block whose covered range ends up entirely
locked by the guard is skipped altogether rather than injecting a summary with
no removal behind it.
The summary itself is injected as a role:"user" message wrapped in a
<conversation-checkpoint> envelope, so the model can see where the replaced
range began and ended. The tags are part of what your model reads on every
compressed turn — mention them if you need to.
| Key | Default | Meaning |
|---|---|---|
guardToolPairs | true | ON by default, escape hatch only. Never remove a message carrying a tool-call whose tool-result survives. Applies to compression blocks and deduplication; anything other than an explicit false counts as on. |
usage)The trigger merges two independent numbers:
session.step.ended: input + output + reasoning + cache.read + cache.write). Exact for that request, but it describes the previous prompt,
so it goes stale the moment pruning/compression shrinks the next one.Neither is safe alone, so the two are clamped in both directions:
| Key | Default | Meaning |
|---|---|---|
usage.trustRatio | 0.5 | A measurement below this fraction of the estimate is treated as stale and the estimate wins. |
usage.capRatio | 3 | A measurement above this multiple of the estimate is treated as a broken reading and capped at capRatio × estimated. |
Both fields are optional and fall back to the defaults, so an absent block is
safe. The panel's Source: line reports where the panel's own headline
figure came from — measured (server-reported) when the server reported usage
for that request, otherwise estimated (approximate) — so an approximation is
never mistaken for an exact count. It does not replay the merge above: the
trigger applies its own trustRatio/capRatio rules to the measured total
(input + output + reasoning + cache.read + cache.write) against the
outgoing-prompt estimate, a different quantity from the panel's Context figure.
The panel's Context: line is the last request's prompt size
(input + cache.read + cache.write) and its Lifetime: line is the session's
cumulative spend; the two are never mixed, and Source: labels only the
Context figure.
Global limits are only resolved when they are actually needed: a model with a
valid compress.modelMaxLimits / compress.modelMinLimits override never
reads — and never warns about — the global maxContextLimit /
minContextLimit. The fallback chain itself is unchanged and covered by the
test suite: a broken per-model override degrades to the configured global
(not straight to the built-in default), a broken global degrades to the
built-in default, and warnings fire once per (key, issue, value) across
repeated calls.
Unlike simple text truncation, slim analyzes the semantic content of messages and groups related tool calls together. This preserves context while removing redundancy.
Slim considers the cost of tokens when deciding what to compress. It prioritizes compressing expensive operations (like large file reads) while preserving cheap but important context.
The plugin learns from your compression patterns and adjusts thresholds over time. If you tend to need more context, it will compress less aggressively. If you're efficient, it will compress more.
State is saved to disk, so compression history and learning persist across restarts.
| Command | Description |
|---|---|
/panel | Open the Slim TUI panel in a dialog with context usage, stats, and trigger thresholds |
/compress | Send the assistant a compression instruction (delivery: "steer") |
/status | Show a one-line context health report in a dialog (never written to the session) |
/slim-debug | Toggle debug in ~/.config/opencode/slim.jsonc and show the result in a dialog |
Note: Compression is performed by the AI assistant using the compress tool. The slash command provides guidance on usage; it is the only slash command that writes to the session, and it does so with an explicit delivery: "steer" (see Compress Tool).
BREAKING CHANGES
strategies.purgeErrors now defaults to enabled: false. The strategy was
configured to run but could not match a call to its errored result: it read the
pairing id from toolCallID / callID, fields the v2 message shape does not
carry, so every lookup came back empty and the purge never fired on any
request. The lookup is fixed, and because it rewrites the outgoing prompt from
a path that had never executed, it is now opt-in. If you set the key
explicitly you get the working behaviour; if you never set it, nothing changes
for you until you ask for it — see
Purge-errors migration.FIXES
part.output was silently missing from every
protected tool's section, so restoring it exposed a section bounded only by the
tool's own output size. The section is now capped and says so when it has been
truncated, so a compression is always smaller than the range it replaced.0%
rather than a negative figure, and the panel no longer prints negative tokens
or a negative dollar amount saved. A compression that does not shrink is not a
saving.NEW
DOCS
purgeErrors opt-in default in
Purge-errors migration, alongside the
Tool-pair guard it interacts with on the same messages.FIXES
[invalid_request_error] invalid request. The covered range was selected on
token size, and a message carrying only tool calls counts as zero text tokens —
so such a message fell outside the range while its (large) tool result fell
inside it, leaving an orphaned tool_call_id on the wire. This affected the
compress tool and automatic compression alike.<conversation-checkpoint> envelope,
so the model can see where the replaced range began and ended.NEW
strategies.guardToolPairs (default on) is the escape hatch for the pair
protection described above. Turn it off only if you are debugging the guard
itself — see Tool-pair guard for the full semantics.FIXES
default() limit and, failing that, the documented
DEFAULT_MODEL_LIMIT with a warning. A usage percentage computed against
another model's context window is meaningless. The TUI's resolver was
tightened the same way, from a modelID-only match to an exact
providerID+modelID match.cache.read re-reads the whole context, and
which therefore produced bogus "100% critical" readings. The cumulative figure
is still reported where it is legitimate (as lifetime spend/cost) and is still
explicitly labelled as not being a context size. The status/CRITICAL
derivation is now clamped to 0..100.Slim Plugin vX.Y.Z
title was a hardcoded literal that had drifted to v2.1.0; it is now derived
from a single PLUGIN_VERSION constant, guarded by a test that keeps it in
sync with package.json.DOCS
## Credits section.BREAKING CHANGES
Session.Info.tokens lifetime
cumulative cost counter is no longer used to decide the trigger, so existing
config files with tuned maxContextLimit/minContextLimit can behave
differently and may need re-tuning. In the panel, Context: is the last
request's prompt size (input + cache.read + cache.write) and Lifetime: is
the session's cumulative spend; the two are shown separately. When the
transcript carries no per-turn usage the old figure is still shown, explicitly
labelled as lifetime cumulative./panel, /status and /slim-debug no longer write to the session (no
transcript message); their output is shown in a modal dialog
(context.ui.dialog.alert). Previous versions wrote the output into the
message stream. Only /compress still writes (it is an instruction to the
model) and now passes an explicit delivery: "steer" instead of relying on
the server default.FIXES
client.session.synthetic,
data.session.message.*) and migrated to the OpenCode v2 API.id field, so the old code always found an empty id and
deleted every block. Now uses deterministic content-based key generation plus
an ambiguity lock.session.step.ended) plus
a full-prompt-scoped estimate.String()
produced [object Object].RangeError when above 100% (negative repeat on a full bar).await, so it
always fell back to 200000.80,5%).NEW
maxContextLimit/minContextLimit accept absolute token counts (a bare
number); percentage forms (80%, 80,5%) keep working.strategies.pruneOutputs prunes tool output (default OFF, opt-in) with turn
protection. Auto-compress keys are captured before pruning replaces pruned
messages with clones, so pruning no longer desynchronises the persisted
compression block's anchors from the messages it covered (which left the block
inert and re-triggered every throttle window).usage.trustRatio / usage.capRatio measurement-trust settings.compressionBlocks, nudge anchors, token
measurement).Source: measured/estimated line and prune statistics. The
Source: line describes where the headline figure came from —
measured (server-reported) when the server reported usage, otherwise
estimated (approximate) — so an approximation is never read as an exact
count. The Prune: line reports outputs pruned and characters/tokens saved on
the last request (a per-request figure, since pruning is re-applied every
request; the line is absent while the default-off feature has not pruned
anything)./panel states its scope: the stats come from session.context ("all
messages after the last compaction"), so Messages:/Tokens (est) are window
counts, not session totals, and it now shows the resolved compression trigger
(Trigger: … tokens (…% of … window) · floor …), matching the panel tool.deriveStats counts unknown message types (agent-switched, model-switched,
location-switched, idle, …) as system instead of assistant, keeping
user + assistant + system === total messages.strategies.pruneOutputs (off by default) and
strategies.turnProtection in the config reference, plus the usage
measurement-trust ratios (trustRatio / capRatio), and clarified that
maxContextLimit is compared against the outgoing prompt estimate/measurement
— not the lifetime cumulative counter./panel in the TUI not reflecting real context usage:
Session.Info.tokens + cost + model + context window) via context.client.session.get() and prints them (measured tokens, %, cost, model) at the bottom of the panel, matching what the panel tool reports.context.ui.router.current() instead of a non-existent context.router, so the panel targets the focused session rather than always the first one.context.data.session.message.sync() before reading the transcript so stats aren't computed from an empty/stale cache.ctx.model.default().data.limit.context instead of the hard-coded 200k./panel now reads live server measurements (Session.Info.tokens + cost) via ctx.session.get() and feeds them to buildPanelData, so the headline tokens/percent/cost match what OpenCode's UI reports.solid-js, @opentui/core, @opentui/solid to devDependencies.
The workflow runs npm ci --legacy-peer-deps, which skips peer deps, so loading
@opencode/plugin/tui failed with Cannot find package 'solid-js'./panel output showing User tokens: 0: user/system messages carry their text
on a top-level text field (not inside content), which deriveStats now captures.feat(panel-as-message): /panel and slim-panel now print the context stats as plain
text into the message stream via client.session.synthetic, instead of taking over
OpenCode's own panel UI (session.panel slot + ui.panel.open removed).keymap.provider is missing in the CLI plugin: register the keymap layer inside
an app slot render (where the keymap provider is available) instead of in setup().Cannot find package 'react' when the plugin is loaded from the global npm cache:
add a per-file /** @jsxImportSource @opentui/solid */ pragma to src/tui.tsx so JSX
always compiles against @opentui/solid/jsx-runtimetsconfig.json in the published package so loaders that read jsxImportSource from config pick it upcompaction hook so history actually shrinks (the context hook only affects the outgoing request)compress/panel tools with options.codemode so they appear in agent/codemode environmentstools always equalled zerosession.panel slot (slim-panel / /panel)Breaking Changes: Migrated to OpenCode v2 plugin API.
@opencode-ai/plugin to @opencode/pluginPlugin.define() pattern instead of server functionctx.tool.transform() with JSON Schema inputctx.session.hook()ctx.event.subscribe()MIT
FAQs
Smart context management plugin for OpenCode v2 - semantic compression, cost-aware pruning, adaptive thresholds
The npm package @serkanalgur/opencodev2-slim receives a total of 388 weekly downloads. As such, @serkanalgur/opencodev2-slim popularity was classified as not popular.
We found that @serkanalgur/opencodev2-slim demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Research
/Security News
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.

Research
/Security News
Socket found 16 malicious Firefox extensions designed to steal crypto wallet recovery phrases and private keys using cloned Rabby and OKX interfaces.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.