@clawplays/ospec-cli
Advanced tools
@@ -96,3 +96,3 @@ --- | ||
| - شغّل deterministic preflight للتصميم والخطة، ثم اشتق task graph، ثم نفّذ combined planning review مستقلة واحدة. يسمح بإصلاح تخطيط مجمّع واحد وfresh re-review واحدة فقط، وتبقى task review وfinal combined review وverification مطلوبة. | ||
| - شغّل deterministic preflight للتصميم والخطة، ثم اشتق task graph، ثم نفّذ combined planning review مستقلة واحدة. يسمح بإصلاح تخطيط مجمّع واحد وإعادة مراجعة تفاضلية واحدة كحد أقصى؛ يعيد فشل المنفّذ دون تعديل التخطيط تسليح الحصة، وتُقبل النتائج التي لا تتجاوز medium حتمياً كـ `APPROVED_WITH_CONCERNS` بعد الإصلاح، وتبقى task review وfinal combined review وverification مطلوبة. | ||
| - يحل worker/reviewer logical model profile حسب dispatch target الفعلي، بما في ذلك launch override. افصل requested/configured model عن provider-observed model؛ بدون provider/usage evidence يبقى observed model غير معروف. | ||
@@ -99,0 +99,0 @@ - يتلقى command runner المسار `OSPEC_USAGE_FILE` ويجمع sidecar تلقائيا؛ ويبقى `ospec execute complete ... --usage-file usage.json` للإدخال اليدوي. تسجل metrics المصدر والحقول المرصودة وتغطية complete/partial/missing، ولا تعرض القيمة غير المبلّغ عنها كصفر مقاس. |
@@ -40,3 +40,3 @@ --- | ||
| - تستخدم allowlist الاختيارية الاستبدال الآمن. عند طلب حد إضافي استخدم `ospec loop allowlist derive/check/apply --from-task-graph`، وراجع فرق CAS، ووافق صراحة على توسيع الصلاحيات المقصود. | ||
| - تستخدم مرحلة design/plan deterministic inline preflight ثم combined planning review مستقل واحد. يسمح بمحاولة grouped planning repair واحدة وfresh re-review واحدة، وبعد فشل متكرر يتوقف المسار بثبات. | ||
| - تستخدم مرحلة design/plan deterministic inline preflight ثم combined planning review مستقل واحد. يسمح بمحاولة grouped planning repair واحدة وإعادة مراجعة تفاضلية واحدة كحد أقصى؛ فشل المنفّذ دون تعديل محتوى التخطيط يعيد تسليح الإصلاح دون استهلاك الحصة، وبعد اكتمال الإصلاح تُقبل النتائج التي لا تتجاوز medium حتمياً كـ `APPROVED_WITH_CONCERNS`، وبعد فشل دلالي متكرر يتوقف المسار بثبات. | ||
| - يجب أن تتقارب عمليات review repair. ترث task downstream التي تعدل ملفات مشتركة التزامات regression من upstream المتعدية. تمثل الجولتان الافتراضيتان عتبة تقارب. يستمر التنفيذ عندما تتغير structured finding IDs، ولا يستمر المعرف الثابت إلا عند تغير structured finding fingerprint وcode snapshot المصرح بهما معا. في continuous mode تحصل مجموعة findings المتوقفة للـ task أو final review على strategy escalation دائمة واحدة مرتبطة بالـ scope والمعرفات نفسها؛ نفذ packet إعادة تحليل root cause وfocused regression مرة واحدة، ثم توقف إذا بقيت المجموعة نفسها متوقفة. يحتفظ strict mode بالحد المضبوط. عندما تكون final review بحالة `BLOCKED` يجب التوقف لحل blocker وعدم بدء grouped repair. لا ترفع الحدود لتكرار عمل لم يتغير. | ||
@@ -101,3 +101,3 @@ - يجب أن يكون انتظار native child محدوداً في كل harness. يعيد `wait_agent` في Codex/GPT وpolling لـ Claude Task وأي native wait آخر التحكم خلال 60 ثانية؛ هذا حد poll واحدة للـ controller وليس حد تشغيل child. حدّث كل child حي قبل `heartbeatDueAt`، وعند الاكتمال استخدم `loop finalize` الصادر مع action لحفظ evidence وresult ذرياً، ثم أعد tick بعد كل poll. يمكن للـ child الاستمرار عبر عدة polls حتى action deadline، وتوجد result grace محدودة بعد اكتمال evidence. عند غياب capacity يستخدم implementation التوازي الافتراضي وهو ثلاث مهام من دون تقليل توازي review الآمن. إذا كان harness يعرف بثقة capacity حالية أكبر للـ child فيمكنه ربطها بجلسة controller النشطة ورفع `maxParallel` وفقاً لذلك؛ لا تخمّن capacity ولا تعِد استخدام قيمة قديمة. | ||
| - شغّل deterministic preflight للتصميم والخطة، ثم اشتق task graph، ثم نفّذ combined planning review مستقلة واحدة؛ يسمح بإصلاح تخطيط مجمّع واحد وfresh re-review واحدة فقط. | ||
| - شغّل deterministic preflight للتصميم والخطة، ثم اشتق task graph، ثم نفّذ combined planning review مستقلة واحدة؛ يسمح بإصلاح تخطيط مجمّع واحد وإعادة مراجعة تفاضلية واحدة كحد أقصى. يعيد فشل المنفّذ دون تعديل التخطيط تسليح الحصة، وتُقبل النتائج التي لا تتجاوز medium حتمياً كـ `APPROVED_WITH_CONCERNS` بعد الإصلاح؛ لا يبطل اعتماد التخطيط إلا تغيير دلالي في محتواه، ولا يبطله تقدم التنفيذ. | ||
| - تربط `.skillrc.workflow.model_profiles` ملفات `mechanical` و`standard` و`strong_reasoning` و`review` و`final_review` المنطقية بنماذج target؛ وعند غياب الربط يستخدم harness default مع warning في packet. | ||
@@ -104,0 +104,0 @@ - يستخدم command runner المتغير `OSPEC_USAGE_FILE` لجمع normalized usage تلقائيا، ويبقى `--usage-file` كتجاوز يدوي. يجمع `execution-metrics.json` حسب capability tier وmodel profile وworkflow stage ويعرض تغطية complete/partial/missing. |
@@ -117,3 +117,3 @@ --- | ||
| - Run deterministic design and plan preflights, derive the task graph, then run one independent combined planning review. One grouped planning repair and one fresh re-review are the maximum; task review, final combined review, and verification remain required. | ||
| - Run deterministic design and plan preflights, derive the task graph, then run one independent combined planning review. One grouped planning repair and at most one delta-scoped re-review are the maximum; no-edit executor failures re-arm the repair, all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS`, and task review, final combined review, and verification remain required. | ||
| - Workers and reviewers use logical model profiles resolved against the actual dispatch target, including launch overrides. Keep requested/configured model separate from provider-observed model; absent provider/usage evidence means observed model is unknown, not selected by assertion. | ||
@@ -120,0 +120,0 @@ - Command runners receive `OSPEC_USAGE_FILE` and automatically ingest that sidecar; `ospec execute complete ... --usage-file usage.json` remains available for manual ingestion. Metrics record their source, observed fields, and complete/partial/missing coverage, so an unreported counter is not presented as a measured zero. |
@@ -42,3 +42,3 @@ --- | ||
| - Optional configured allowlists use secure replacement semantics. When an extra boundary is requested, use `ospec loop allowlist derive/check/apply --from-task-graph`, review the CAS-bound diff, and explicitly approve intended expansions. | ||
| - Design and plan use deterministic inline preflights, followed by one independent combined planning review. One grouped planning repair and one fresh re-review are allowed; a repeated failure blocks stably. `--force` never bypasses Loop lifetime, token, STOP, no-progress, context, or executor-provenance guards. | ||
| - Design and plan use deterministic inline preflights, followed by one independent combined planning review. One grouped planning repair and at most one delta-scoped re-review are allowed; an executor failure with no planning edits re-arms the repair, all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS` after the repair, and a repeated semantic failure blocks stably. `--force` never bypasses Loop lifetime, token, STOP, no-progress, context, or executor-provenance guards. | ||
| - Review repair must converge. A downstream task that shares files inherits transitive upstream regression obligations. The default two-round values are convergence thresholds. Continue when structured finding IDs change. A stable ID may also continue only when both its structured finding fingerprint and its prior authorized repair-scope code snapshot changed. In continuous mode, stalled task or final findings receive one durable strategy escalation for that exact scope and finding-ID set; execute its root-cause and focused-regression packet once, then stop if the same set remains stalled. Strict mode retains the configured limit. A blocked final review stops for blocker resolution; it never enters grouped repair. Do not raise a limit to repeat unchanged evidence. | ||
@@ -121,3 +121,3 @@ - Native child waiting is bounded on every harness. Codex/GPT `wait_agent`, Claude Task polling, and other native waits return within 60 seconds; this limits one controller poll, not the child runtime. Refresh every live child before `heartbeatDueAt`, persist each completed result with its emitted `loop finalize` command, and re-tick after every poll. A live child may continue across polls up to its action deadline, and evidence-complete work receives a bounded result grace period. Unknown capacity uses the default implementation concurrency of three without reducing safe review parallelism. A harness that authoritatively knows a larger current child capacity may report it for the active controller session and raise `maxParallel` accordingly; never guess or reuse stale capacity. | ||
| - Run deterministic design and plan preflights, derive the task graph, then run one independent combined planning review. Allow at most one grouped planning repair and one fresh re-review. | ||
| - Run deterministic design and plan preflights, derive the task graph, then run one independent combined planning review. Allow one grouped planning repair and at most one delta-scoped re-review; a repair executor failure with no planning edits re-arms the allowance, and a completed repair whose findings were all medium or lower settles deterministically as `APPROVED_WITH_CONCERNS`. Planning approvals invalidate only on semantic planning changes, never on execution progress. | ||
| - `.skillrc.workflow.model_profiles` maps the `mechanical`, `standard`, `strong_reasoning`, `review`, and `final_review` logical profiles to target-specific models; missing mappings use the harness default and produce a packet warning. | ||
@@ -124,0 +124,0 @@ - Command runners use `OSPEC_USAGE_FILE` for automatic normalized usage ingestion; `--usage-file` remains a manual override. `execution-metrics.json` aggregates by capability tier, model profile, and workflow stage and reports complete/partial/missing coverage. |
@@ -96,3 +96,3 @@ --- | ||
| - design/plan deterministic preflight、task graph 導出、独立 combined planning review の順に実行する。grouped planning repair と fresh re-review は各 1 回までとし、task review、final combined review、verification は引き続き必須。 | ||
| - design/plan deterministic preflight、task graph 導出、独立 combined planning review の順に実行する。grouped planning repair 1 回と delta re-review 最大 1 回までとし、planning 内容を変更しない executor 失敗は許容量を再アームし、findings がすべて medium 以下なら repair 後に `APPROVED_WITH_CONCERNS` として決定論的に確定する。task review、final combined review、verification は引き続き必須。 | ||
| - worker/reviewer の logical model profile は launch override を含む実際の dispatch target に対して解決する。requested/configured model と provider-observed model を分離し、provider/usage evidence がなければ observed model は unknown とする。 | ||
@@ -99,0 +99,0 @@ - command runner は `OSPEC_USAGE_FILE` を受け取り sidecar を自動集計する。`ospec execute complete ... --usage-file usage.json` は手動入力として残す。metrics は source、observed fields、complete/partial/missing coverage を記録し、未報告値を測定済み 0 として扱わない。 |
@@ -40,3 +40,3 @@ --- | ||
| - optional allowlist は安全な置換であり、暗黙の追加ではない。追加境界が必要な場合だけ `ospec loop allowlist derive/check/apply --from-task-graph` を使い、CAS 差分と権限拡張を明示的に確認する。 | ||
| - design/plan stage は deterministic inline preflight を使い、その後に独立した combined planning review を 1 回実行する。grouped planning repair と fresh re-review は各 1 回だけ許可し、再失敗は安定して block する。 | ||
| - design/plan stage は deterministic inline preflight を使い、その後に独立した combined planning review を 1 回実行する。grouped planning repair 1 回と delta re-review 最大 1 回を許可する。planning 内容を変更しない executor 失敗は許容量を消費せず repair を再アームし、repair 完了後に findings がすべて medium 以下なら `APPROVED_WITH_CONCERNS` として決定論的に確定する。意味的な再失敗は安定して block する。 | ||
| - review repair は収束させる。共有ファイルを変更する downstream task は推移的 upstream の regression obligation を継承する。既定の 2 round は収束しきい値であり、structured finding ID が変化すれば自動続行する。同じ ID でも structured finding fingerprint と直前に許可された repair scope 内の code snapshot が両方変化した場合だけ続行できる。continuous mode では、停滞した task または final finding 集合に、その正確な scope と finding ID に対する durable strategy escalation を 1 回だけ発行する。root cause の再評価と focused regression を要求する packet を 1 回実行し、同じ集合が停滞したままなら停止する。strict mode は設定済み上限を維持する。final review が `BLOCKED` の場合は blocker の解決まで停止し、grouped repair に進めない。変化しない作業を繰り返すために上限を引き上げてはならない。 | ||
@@ -101,3 +101,3 @@ - すべての harness で native child の待機を bounded にする。Codex/GPT の `wait_agent`、Claude Task polling、その他の native wait は 60 秒以内に戻る。この 60 秒は controller poll 1 回の上限であり、child runtime の上限ではない。各 live child を `heartbeatDueAt` 前に更新し、完了時は action の `loop finalize` で evidence と result を atomic に保存して poll ごとに再 tick する。child は action deadline まで複数 poll にまたがって実行でき、evidence 完了後には bounded result grace がある。capacity 不明時の implementation は既定の並列数 3 を使用し、安全な review 並列性は維持する。harness が現在のより大きな child capacity を確実に把握できる場合は active controller session に結び付けて報告し、必要に応じて `maxParallel` を引き上げられるが、capacity を推測したり古い値を再利用してはならない。 | ||
| - design/plan deterministic preflight、task graph 導出、独立 combined planning review の順に実行し、grouped planning repair と fresh re-review は各 1 回までとする。 | ||
| - design/plan deterministic preflight、task graph 導出、独立 combined planning review の順に実行し、grouped planning repair 1 回と delta re-review 最大 1 回までとする。planning 内容を変更しない executor 失敗は許容量を再アームし、repair 完了後に findings がすべて medium 以下なら `APPROVED_WITH_CONCERNS` として決定論的に確定する。planning 承認は planning の意味的変更のみで無効化され、実行進捗では無効化されない。 | ||
| - `.skillrc.workflow.model_profiles` は `mechanical`、`standard`、`strong_reasoning`、`review`、`final_review` logical profile を target-specific model に対応付ける。未設定時は harness default と packet warning を使う。 | ||
@@ -104,0 +104,0 @@ - command runner は `OSPEC_USAGE_FILE` で normalized usage を自動集計し、`--usage-file` は手動 override として残る。`execution-metrics.json` は capability tier、model profile、workflow stage 別に集計し、complete/partial/missing coverage を報告する。 |
@@ -148,3 +148,3 @@ --- | ||
| - 依次执行 design/plan 确定性预检,派生 task graph,再执行一次独立 combined planning review;最多允许一次整体规划修复和一次 fresh re-review。task review、最终 combined review 和验证仍然保留。 | ||
| - 依次执行 design/plan 确定性预检,派生 task graph,再执行一次独立 combined planning review;最多允许一次整体规划修复和一次差量复审。未改动规划内容的执行器失败会重新武装修复额度,修复后 findings 全部不高于 medium 时确定性通过为 `APPROVED_WITH_CONCERNS`。task review、最终 combined review 和验证仍然保留。 | ||
| - worker/reviewer 使用逻辑 model profile,并按实际 dispatch target(包括 launch override)解析。requested/configured model 与 provider observed model 必须分开;没有 provider/usage 证据时 observed model 是未知,不能宣称已选择。 | ||
@@ -151,0 +151,0 @@ - 命令执行器会收到 `OSPEC_USAGE_FILE` 并自动归集该 sidecar;`ospec execute complete ... --usage-file usage.json` 继续作为手工入口。指标必须记录来源、实际观测字段和 complete/partial/missing 覆盖率,未上报的计数不能显示成已测得的零。 |
@@ -42,3 +42,3 @@ --- | ||
| - 可选白名单采用安全替换语义。需要额外边界时使用 `ospec loop allowlist derive/check/apply --from-task-graph`,检查 CAS 绑定的差异,并仅对预期的权限扩大显式确认。 | ||
| - design/plan 阶段使用确定性 inline preflight,随后执行一次独立的合并规划复审。只允许一次 grouped planning repair 和一次 fresh re-review;重复失败后稳定阻断。 | ||
| - design/plan 阶段使用确定性 inline preflight,随后执行一次独立的合并规划复审。只允许一次 grouped planning repair 和最多一次差量复审;执行器失败且未改动规划内容时重新武装修复而不消耗额度,修复完成后 findings 全部不高于 medium 时确定性通过为 `APPROVED_WITH_CONCERNS`;语义层面重复失败后稳定阻断。 | ||
| - review repair 必须收敛。共享文件的下游任务要继承传递上游的回归义务;默认两轮是收敛阈值。结构化 finding ID 变化时自动继续;同一 ID 只有在结构化 finding 指纹与上一轮授权 repair scope 内的代码快照同时变化时也可继续。连续模式下,停滞的 task 或 final finding 集合会按精确 scope 与 finding ID 获得一次持久化策略升级;只执行一次要求重新定位根因并加强聚焦回归的 packet,同一集合仍停滞时必须停止。严格模式继续遵守配置上限。final review 为 `BLOCKED` 时必须停下解决 blocker,不得进入 grouped repair。不得提高上限重复未变化的工作。 | ||
@@ -127,3 +127,3 @@ - 所有 harness 的 native child 等待都必须有界。Codex/GPT 的 `wait_agent`、Claude Task 轮询以及其它原生等待必须在 60 秒内返回;60 秒只限制一次 controller poll,不是 child 的执行上限。每个 live child 都要在 `heartbeatDueAt` 前续租,完成后用 action 给出的 `loop finalize` 原子提交证据和结果,并在每轮 poll 后重新 tick。child 可跨多个 poll 运行到 action 的绝对期限,证据完成后还有有界的结果宽限期。capacity 未知时 implementation 使用默认并发 3,且不降低安全 review 的并行度。harness 能可靠获知当前更大的 child 容量时,可以把它绑定到当前 controller session 并相应提高 `maxParallel`;绝不能猜测或复用过期容量。 | ||
| - 依次执行 design/plan 确定性预检,派生 task graph,再执行一次独立 combined planning review;最多允许一次整体规划修复和一次 fresh re-review。 | ||
| - 依次执行 design/plan 确定性预检,派生 task graph,再执行一次独立 combined planning review;最多允许一次整体规划修复和一次差量复审。未改动规划内容的执行器失败会重新武装修复额度,修复完成后 findings 全部不高于 medium 时确定性通过为 `APPROVED_WITH_CONCERNS`;规划批准只因规划语义变化失效,执行进度不会使其失效。 | ||
| - `.skillrc.workflow.model_profiles` 将 `mechanical`、`standard`、`strong_reasoning`、`review`、`final_review` 逻辑 profile 映射到各 target 模型;未配置时使用 harness 默认并在 packet 中警告。 | ||
@@ -130,0 +130,0 @@ - 命令执行器通过 `OSPEC_USAGE_FILE` 自动归集标准化 usage;`--usage-file` 保留为手工覆盖入口。`execution-metrics.json` 按 capability tier、model profile 和 workflow stage 汇总,并报告 complete/partial/missing 覆盖率。 |
@@ -45,3 +45,3 @@ --- | ||
| - `Announce-Before-Act`: never run the change flow silently. Announce in one line which skill you are using (`ospec-change`) and the current stage, which `ospec` command you are about to run and the artifact it writes, and which gate is blocking when progress stops. | ||
| - `Brainstorm-First`: before implementing, confirm scope and acceptance with the user when anything is ambiguous, and ask one question at a time instead of silently assuming direction, API, UI, risk, or scope. Record the agreed scope in `proposal.md` rather than guessing. **NEVER auto-select the recommended option or resolve a decision gate yourself — `recommended` is only a hint to show the user; present every gate and wait for the user's actual choice instead of running the change in one shot.** **Present each open decision using the best interactive mechanism your harness has — a native question UI (Claude Code `AskUserQuestion`, Gemini `ask_user`) if available, otherwise your plan/approval UI (Codex Plan mode) if available, otherwise plain chat text — you always ask the user, only the presentation differs.** When you run `ospec brainstorm`, do not leave it as an unanswered template: ask the user the decision gates and record each answer with `ospec brainstorm resolve [path] --brainstorm <id> --gate <gate-id> --select <option-id>` so the brainstorm has a result. | ||
| - `Brainstorm-First` (forked decisions only): raise a decision gate only when the requirement has a **genuine fork** — mutually exclusive API shapes, competing UI approaches, data-model or storage choices, destructive or hard-to-reverse operations, or a scope conflict with what the user asked for. For routine unambiguous changes — a bug fix with an evident cause, a mechanical refactor, a docs update, a small addition with one reasonable implementation — do **not** open decision gates or run `ospec brainstorm`: proceed with the reasonable default and record the assumptions you made in `proposal.md` so the user can correct them. When a genuine fork exists: ask one question at a time; **NEVER auto-select the recommended option or resolve a gate yourself — `recommended` is only a hint to show the user; present each gate and wait for the user's actual choice.** **Use the best interactive mechanism your harness has — a native question UI (Claude Code `AskUserQuestion`, Gemini `ask_user`) if available, otherwise your plan/approval UI (Codex Plan mode) if available, otherwise plain chat text — you always ask the user, only the presentation differs.** If you did run `ospec brainstorm`, do not leave it as an unanswered template: record each answer with `ospec brainstorm resolve [path] --brainstorm <id> --gate <gate-id> --select <option-id>` so the brainstorm has a result. | ||
| - `Zero-Setup`: the user only describes the change; you run every `ospec` command yourself and never ask them to type setup or execution commands. In a Claude Code harness, if `.claude/settings.json` does not yet reference `.ospec/hooks/claude/ospec-claude-hook.cjs`, run `ospec session hook --target claude --apply` once (idempotent). | ||
@@ -48,0 +48,0 @@ |
@@ -18,3 +18,3 @@ --- | ||
| - **Executor lifecycle is durable and bounded.** After native subagent dispatch, record `ospec loop heartbeat <goal> --action-item <id> --executor <child-id>` and refresh every live child before its `heartbeatDueAt`. Never make one indefinite native wait: Codex/GPT use `wait_agent` for at most 60 seconds per poll, Claude uses bounded background Task polling when available, and every other native adapter follows its published `maxWaitMs`. Sixty seconds limits one controller poll, not the child runtime; a live child continues across polls up to its action deadline. Commit each finished child with its emitted `ospec loop finalize ...` command, persist completed siblings immediately, and re-run `loop run --once --json` after every poll. Each successful bounded controller poll renews the short lease for already-claimed live children without extending the absolute deadline. A poll may recover that same claim only within one bounded 60-second wait after the short-lease boundary; direct late results remain rejected, a controller that stops polling still lets orphan leases expire, and no renewal moves the absolute deadline. Successful finalize requires authoritative durable evidence; evidence-complete work receives a bounded result grace period. Legacy `loop result` remains supported. Use `ospec loop recover --force` only when the prior session/child is known to be gone. Expired items requeue; completed siblings do not. | ||
| - **Planning quality is fast and bounded.** Run design preflight, then implementation-plan preflight, derive the task graph, and let Loop issue one independent combined planning review. The two preflights use no reviewer child. A `NEEDS_CHANGES` planning review permits one grouped repair and one fresh re-review; another failure is a stable blocker, never an open-ended loop. | ||
| - **Planning quality is fast and bounded.** Run design preflight, then implementation-plan preflight, derive the task graph, and let Loop issue one independent combined planning review. The two preflights use no reviewer child. A `NEEDS_CHANGES` planning review permits one grouped repair and at most one delta-scoped re-review; a repair executor failure with no planning edits re-arms instead of consuming the allowance, all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS` after the repair, and another semantic failure is a stable blocker, never an open-ended loop. | ||
| - **Required decisions always block.** Present each required decision to the user, never auto-select the recommendation, and record the answer with `ospec execute decision ... --select ... --answered-by user` before the loop proceeds. New brainstorm resolutions require the same `--answered-by user` provenance. | ||
@@ -21,0 +21,0 @@ - **`/goal` is capability-probed, not inferred from a target name.** `ospec execute launch --primitive goal` produces a native-`/goal` instruction only when the current harness explicitly reports support; otherwise the same controller runs the verify-driven loop through native subagents. |
@@ -58,3 +58,3 @@ --- | ||
| - قبل اشتقاق task graph، شغّل `ospec execute preflight [changes/active/<change>] --stage design` ثم `--stage plan` لإنشاء deterministic inline preflight packets وapproval artifacts. اشتق أو حدّث task graph بعد نجاح المرحلتين فقط، ولا تشغّل أي مرحلة reviewer child. اجمع red test العادي وproduction implementation ودليل green/refactor في atomic task واحدة | ||
| - بعد اشتقاق task graph يجب أن يصدر Loop combined planning review مستقلة واحدة قبل workspace أو worker dispatch. يسمح بإصلاح تخطيط مجمّع واحد وfresh re-review واحدة فقط، ثم يتوقف بثبات عند تكرار الفشل | ||
| - بعد اشتقاق task graph يجب أن يصدر Loop combined planning review مستقلة واحدة قبل workspace أو worker dispatch. يسمح بإصلاح تخطيط مجمّع واحد وإعادة مراجعة تفاضلية واحدة كحد أقصى؛ يعيد فشل المنفّذ دون تعديل التخطيط تسليح الحصة، وتُقبل النتائج التي لا تتجاوز medium حتمياً كـ `APPROVED_WITH_CONCERNS` بعد الإصلاح، ثم يتوقف بثبات عند تكرار الفشل الدلالي | ||
| - قبل handoff إلى worker استخدم `ospec execute workspace [changes/active/<change>]` لتسجيل سلامة git workspace في `artifacts/agents/workspace-status.json` (`workspace-status.json`)؛ يسمح Goal قائم فقط بالمسارات التابعة لأهداف task غير `PENDING`، أو ملف `tsconfig.tsbuildinfo` الدقيق داخل حزمة task بدأ فعلا وصرح بأمر build/typecheck، أو لإثبات `ospec update` حالي متحقق من الهاش، وتظهر أي مسارات أخرى بالحالة `needs_isolation` | ||
@@ -61,0 +61,0 @@ - Use `ospec execute route [changes/active/<change>]` to write `workflow-route.json` and `workflow-route.md` with the next recommended OSpec command; this records workflow routing artifacts only and does not edit source files |
@@ -58,3 +58,3 @@ --- | ||
| - Before deriving the task graph, run `ospec execute preflight [changes/active/<change>] --stage design`, then `--stage plan`, to create deterministic inline preflight packets and approval artifacts. Derive or refresh the task graph only after both pass; neither stage launches a reviewer child. Keep ordinary red tests with their production implementation and green/refactor evidence in one atomic task | ||
| - After task graph derivation, Loop must issue one independent combined planning review before workspace or worker dispatch. One grouped planning repair and one fresh re-review are the maximum; repeated failure stops stably | ||
| - After task graph derivation, Loop must issue one independent combined planning review before workspace or worker dispatch. One grouped planning repair and at most one delta-scoped re-review are the maximum; a no-edit executor failure re-arms the repair, all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS` after the repair, and repeated semantic failure stops stably | ||
| - Before worker handoff, use `ospec execute workspace [changes/active/<change>]` to record git workspace safety in `artifacts/agents/workspace-status.json` (`workspace-status.json`); existing Goals may retain only dirty paths owned by non-`PENDING` task targets, exact package-local `tsconfig.tsbuildinfo` derived from a started task's declared build/typecheck verification, or current hash-verified `ospec update` provenance, and every other dirty path reports `needs_isolation` | ||
@@ -61,0 +61,0 @@ - Use `ospec execute route [changes/active/<change>]` to write `workflow-route.json` and `workflow-route.md` with the next recommended OSpec command; this records workflow routing artifacts only and does not edit source files |
@@ -58,3 +58,3 @@ --- | ||
| - task graph 導出前に `ospec execute preflight [changes/active/<change>] --stage design`、続いて `--stage plan` を実行し、deterministic inline preflight packet と approval artifact を作成する。両方の通過後に task graph を導出または更新し、どちらの stage も reviewer child を起動しない。通常の red test、production implementation、green/refactor evidence は 1 つの atomic task にまとめる | ||
| - task graph 導出後、workspace または worker dispatch より前に Loop が独立 combined planning review を 1 回発行する。grouped planning repair と fresh re-review は各 1 回までで、再失敗は安定して停止する | ||
| - task graph 導出後、workspace または worker dispatch より前に Loop が独立 combined planning review を 1 回発行する。grouped planning repair 1 回と delta re-review 最大 1 回までで、planning 内容を変更しない executor 失敗は許容量を再アームし、findings がすべて medium 以下なら repair 後に `APPROVED_WITH_CONCERNS` として決定論的に確定し、意味的な再失敗は安定して停止する | ||
| - worker handoff の前に `ospec execute workspace [changes/active/<change>]` で git workspace safety を `artifacts/agents/workspace-status.json`(`workspace-status.json`)に記録する。既存 Goal では、非 `PENDING` task の target file、開始済み task の宣言済み build/typecheck 検証から導出される package-local の exact `tsconfig.tsbuildinfo`、または現在のハッシュ検証済み `ospec update` provenance に属する dirty path だけを許可し、それ以外は `needs_isolation` を示す | ||
@@ -61,0 +61,0 @@ - Use `ospec execute route [changes/active/<change>]` to write `workflow-route.json` and `workflow-route.md` with the next recommended OSpec command; this records workflow routing artifacts only and does not edit source files |
@@ -58,3 +58,3 @@ --- | ||
| - 派生 task graph 前,依次运行 `ospec execute preflight [changes/active/<change>] --stage design` 和 `--stage plan`,生成确定性 inline preflight packet 与 approval artifacts;两步通过后再派生或刷新 task graph,任何阶段都不启动 reviewer child。普通 red test、对应生产实现和 green/refactor 证据应放在同一个原子 task | ||
| - task graph 派生后,Loop 必须在 workspace 或 worker 派发前执行一次独立 combined planning review。最多允许一次整体规划修复和一次 fresh re-review;重复失败必须稳定停止 | ||
| - task graph 派生后,Loop 必须在 workspace 或 worker 派发前执行一次独立 combined planning review。最多允许一次整体规划修复和一次差量复审;未改动规划内容的执行器失败会重新武装修复额度,修复后 findings 全部不高于 medium 时确定性通过为 `APPROVED_WITH_CONCERNS`;语义层面重复失败必须稳定停止 | ||
| - 派发 worker 前,用 `ospec execute workspace [changes/active/<change>]` 在 `artifacts/agents/workspace-status.json`(`workspace-status.json`)中记录 git 工作区安全状态;已有 Goal 只允许非 `PENDING` 任务目标文件、由已启动任务声明的 build/typecheck 验证精确派生且位于其包内的 `tsconfig.tsbuildinfo`,或当前哈希校验通过的 `ospec update` 证明所归属的脏路径,其余脏路径显示 `needs_isolation` | ||
@@ -61,0 +61,0 @@ - 需要把下一条 OSpec 命令持久化给人或 AI 接手时,用 `ospec execute route [changes/active/<change>]` 写入 `workflow-route.json` 和 `workflow-route.md`;该命令只记录 workflow routing artifacts,不会编辑源码 |
+1
-1
@@ -65,3 +65,3 @@ #!/usr/bin/env node | ||
| const services_1 = require("./services"); | ||
| const CLI_VERSION = '1.9.1'; | ||
| const CLI_VERSION = '1.9.2'; | ||
| function showInitUsage() { | ||
@@ -68,0 +68,0 @@ console.log('Usage: ospec init [root-dir] [--summary "..."] [--tech-stack node,react] [--architecture "..."] [--document-language en-US|zh-CN|ja-JP|ar]'); |
+1
-1
| { | ||
| "name": "@clawplays/ospec-cli", | ||
| "version": "1.9.1", | ||
| "version": "1.9.2", | ||
| "description": "Official OSpec CLI package for spec-driven development (SDD) and document-driven development in AI coding agent and CLI workflows.", | ||
@@ -5,0 +5,0 @@ "main": "dist/index.js", |
+1
-1
@@ -331,3 +331,3 @@ <h1><a href="https://ospec.ai/" target="_blank" rel="noopener noreferrer">OSpec.ai</a></h1> | ||
| - **Task graph controller**: `ospec execute bootstrap` writes a one-change startup/resume snapshot; `preflight` records deterministic design and implementation-plan evidence under `artifacts/agents/planning-preflights/`; `workspace` records git safety; `dispatch`, `launch`, `complete`, and `review` settle native-subagent packets; `debug`, `tdd`, and `verify` record durable evidence; `sync` rebuilds derived status. | ||
| - **Fast planning quality**: deterministic preflights cost no model round trip, while one independent combined planning reviewer checks requirement, architecture, task-graph, dependency, and verification semantics. At most one grouped planning repair and one re-review are allowed before a stable blocker. | ||
| - **Fast planning quality**: deterministic preflights cost no model round trip, while one independent combined planning reviewer checks requirement, architecture, task-graph, dependency, and verification semantics. A `NEEDS_CHANGES` decision permits one grouped planning repair and at most one delta-scoped re-review; a repair executor failure with no planning edits re-arms instead of consuming the allowance, and all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS` after the repair. Planning approvals are invalidated only by semantic planning changes, never by execution progress; repeated semantic failure is a stable blocker. | ||
| - **Measured execution and grouped repair**: command runners can write authoritative usage to `OSPEC_USAGE_FILE` for automatic ingestion, while `--usage-file` remains a manual input. Metrics distinguish complete, partial, and missing coverage. `ospec execute repair` turns all structured `NEEDS_CHANGES` findings into one repair task. | ||
@@ -334,0 +334,0 @@ - **Verified durable documentation**: declared documentation targets capture before/after normalized content hashes, so an unchanged file cannot satisfy a new run. Feature indexes link completed work directly to the durable project documents it updated. |
+1
-1
@@ -50,3 +50,3 @@ --- | ||
| - Run `ospec execute preflight ... --stage design`, then `--stage plan`, before deriving the task graph. These zero-token checks validate document readiness, required decisions, ordering, and provenance inline and never launch reviewer children. | ||
| - After task graph derivation, let Loop issue one independent combined planning review across proposal, design, plan, tasks, graph, and acceptance-to-verification coverage. `NEEDS_CHANGES` permits one grouped planning repair and one fresh re-review; another failure is a stable blocker. | ||
| - After task graph derivation, let Loop issue one independent combined planning review across proposal, design, plan, tasks, graph, and acceptance-to-verification coverage. `NEEDS_CHANGES` permits one grouped planning repair and at most one delta-scoped re-review; a repair executor failure with no planning edits re-arms instead of consuming the allowance, all-medium-or-lower findings settle deterministically as `APPROVED_WITH_CONCERNS` after the repair, and another semantic failure is a stable blocker. | ||
| - Resolve required decisions and workspace isolation before dispatch. | ||
@@ -53,0 +53,0 @@ - Dispatch scoped worker packets, use the launch plan with the current harness native agent mechanism, record completion, then perform one combined task review. |
Sorry, the diff of this file is too big to display
Sorry, the diff of this file is too big to display
Sorry, the diff of this file is too big to display
Sorry, the diff of this file is too big to display
Long strings
Supply chain riskContains long string literals, which may be a sign of obfuscated or packed code.
URL strings
Supply chain riskPackage contains fragments of external URLs or IP addresses, which the package may be accessing at runtime.
Long strings
Supply chain riskContains long string literals, which may be a sign of obfuscated or packed code.
URL strings
Supply chain riskPackage contains fragments of external URLs or IP addresses, which the package may be accessing at runtime.
3332042
0.81%58352
0.62%