Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 6 additions & 18 deletions plugins/aidd-refine/skills/05-improve/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,29 +8,17 @@ argument-hint: conversation | export
```mermaid
flowchart LR
start([conversation ID or export]) --> read-conversation
read-conversation -->|complete| scopes{two or more scopes?}
read-conversation -->|complete| recommend --> target-edits
read-conversation -->|unavailable| unavailable([stop])
scopes -->|no| recommend-local[recommend locally] --> target-edits
scopes -->|yes| isolation{isolated artifact context?}
isolation -->|yes| recommend-parallel[recommend in parallel] --> target-edits
isolation -->|no| recommend-local
target-edits --> question([ask next intent]) --> stop([stop])
```

## Actions

Run the flow above. Read only the next action file.
Run all three actions in order without confirmation. Read only the next file in `actions/`. Keep the project read-only; write only the unique temporary report.

| Action | Does |
| Order | Action |
| --- | --- |
| read-conversation | freeze complete evidence and measure visible cost |
| recommend | analyze relevant scopes and merge grounded findings |
| target-edits | render minimal edits and an executable prompt |

## Transversal rules

- After resolving the source, run all three actions without pausing for confirmation.
- Analyze only a complete conversation and the exact skills or documents it names.
- Exclude every current or previous `improve` invocation and report from the analysis evidence.
- Never invent time, tokens, source coverage, or document status.
- Do not modify, stage, or persist any project file.
| 1 | `01-read-conversation.md` |
| 2 | `02-recommend.md` |
| 3 | `03-target-edits.md` |
Original file line number Diff line number Diff line change
@@ -1,35 +1,23 @@
# 01 - Read conversation

Load complete conversation evidence once and profile its visible cost.
Freeze complete evidence and measure visible cost.

## Input

A current conversation, an exact conversation ID, or a complete transcript export.

## Output

An evidence boundary, a `## Timing` table with `Activity | Observed time | Share | Evidence`, a `## Usage` table with `Metric | Value | Evidence`, and a scope index.
An evidence boundary, a `## Timing` table with `Activity | Observed time | Share | Evidence`, a `## Usage` table with `Metric | Value | Evidence`, and a scope index of encountered resources.

## Process

1. **Resolve.** Choose the host-specific transcript source from [conversation sources](../assets/conversation-sources.md).
- Stop when no complete transcript can be resolved for the exact conversation.
2. **Bound.** Freeze the evidence at the message before the current invocation, or at the supplied export boundary.
3. **Read.** Load every in-scope message, tool call, tool result, and timestamp from that transcript.
4. **Index.** Record relevant turns and invoked or named skills and knowledge files for downstream analysis.
5. **Measure.** Calculate visible time and usage only from timestamps, elapsed records, or host usage data.
- Group explicit tool activity as `research and diagnosis`, `implementation`, or `validation`; keep mixed or unknown time `unattributed`.
- Include tokens, requests, and monetary cost only when the source exposes them.
- Never label unattributed wait as reasoning time.
6. **Render.** Mark a missing metric `unavailable` and order known activity times from longest to shortest.

## Test

| Case | Pass |
| --- | --- |
| A conversation is analyzed | its source resolves to one exact complete transcript |
| The same conversation is analyzed again | its evidence excludes every `improve` invocation and report |
| A timing metric is shown | its evidence identifies visible timestamp or elapsed records |
| A metric is unavailable | the report does not estimate or call it reasoning time |
| A usage metric is shown | its evidence identifies the host record that exposes it |
| A scope is indexed | it identifies exact turns or artifact paths, not a generated summary |
1. **Resolve.** Use the host route in [conversation sources](../assets/conversation-sources.md). Prefer its complete export or host reader. Match the exact session ID before direct storage; search only for that ID, never unrelated sessions, configuration, or authentication. Stop unless the exact complete transcript is available.
2. **Freeze.** End before this invocation, or at the export boundary. Exclude every `improve` invocation and report.
3. **Read once.** Load every in-scope message, tool call, result, and timestamp. Index relevant turns; invoked skills and their used actions or resources; applicable `AGENTS.md`; and task-relevant memory references. Do not read maintained sources yet.
4. **Measure.** Use only timestamps, elapsed records, and exposed host usage.
- Group tool time as `research and diagnosis`, `implementation`, `validation`, or `unattributed`.
- Include tokens, requests, and cost only when exposed; otherwise mark them `unavailable`.
- Separate background-process lifetime from blocking time.
- Treat private reasoning duration as unavailable unless exposed directly; never infer it.
5. **Deliver.** Order known activity times descending and cite every value.
63 changes: 25 additions & 38 deletions plugins/aidd-refine/skills/05-improve/actions/02-recommend.md
Original file line number Diff line number Diff line change
@@ -1,50 +1,37 @@
# 02 - Recommend

Analyze each relevant scope with the smallest grounded change.
Reduce avoidable work without reducing reliability.

## Input

The evidence boundary, timing and usage tables, complete conversation, and scope index.
The evidence boundary, timing and usage tables, complete conversation, and resource index.

## Output

A `## Recommendations` table with `ID | Question | Type | Diagnostic | Evidence | Smallest change | Target | Saving`.
An updated `Scope | Status | Evidence` index, a `## Recommendations` table with `ID | Focus | Type | Diagnostic | Evidence | Smallest change | Source | Saving`, and non-actionable limitations.

## Process

1. **Scope.** Select `behavior`, plus `skill` and `knowledge` only when the scope index names relevant artifacts.
2. **Dispatch.** Analyze locally when one scope exists; otherwise dispatch one read-only analyst per scope in parallel.
- Prefer a lightweight available model and low reasoning effort when the host supports per-agent overrides; otherwise inherit the run defaults.
- Give the behavior analyst the complete frozen transcript; give artifact analysts the same boundary, indexed turns, and exact artifact paths.
- Request isolated or minimal context for artifact analysts when supported; otherwise analyze every scope locally instead of duplicating the transcript.
- Do not dispatch an analyst with no relevant evidence.
- Require table rows only and forbid file writes.
3. **Question.** Make each analyst answer every prompt for its scope.
1. **Verify.** Resolve and read indexed maintained skill resources, applicable `AGENTS.md`, and task-relevant memory references once; reuse a still-valid read. Add and read only relevant paths from the project memory index; never scan unrelated memory. Neither judge from nor target caches or installs. Account for host transforms. Mark only read sources `checked`; mark unresolved or skipped ones `missing` or `not reviewed`.
2. **Scope.** Always assess `behavior`; add `skill` or `knowledge` only for checked indexed sources.
3. **Dispatch.** With multiple isolated scopes, run one read-only analyst per scope in parallel. Otherwise analyze locally.
- Prefer a cheap model and low reasoning effort when per-agent overrides are supported; otherwise inherit host defaults.
- Give the behavior analyst the frozen transcript. Give artifact analysts the boundary, indexed turns, and verified excerpts with evidence; allow a targeted read-only refresh only when validity is uncertain.
- Give every analyst steps 4–6 and require table rows.
4. **Reflect.** Ask yourself, for each scope; keep this assessment internal:
- How could the next run be faster or better?
- What information should be removed or clarified?
- Where should the change live?
- How could it save time or tokens?
- What work was counterproductive?
4. **Verify.** Read a named skill or knowledge file before assessing its information.
5. **Assess.** Label relevant information `obsolete`, `over-specific-or-time-bound`, `duplicate`, `inconsistent`, `counterproductive`, or `correct`.
- Use `correct` when no evidence supports another label, and never render it as a recommendation.
6. **Merge.** Deduplicate findings across scopes and verify only their cited evidence against the frozen source.
7. **Render.** Order by question then `behavior`, `skill`, `knowledge`, and describe each change with the fewest unambiguous words.
- Use `skill`, `behavior`, `knowledge`, or `tooling` as the target type.
- State `time`, `tokens`, `both`, or `unknown` as its saving.
- Render `no change` when a scope has no evidence-backed recommendation.

## Test

| Case | Pass |
| --- | --- |
| A question is shown | it is answered from conversation evidence |
| One relevant scope exists | no parallel analyst is dispatched |
| Several relevant scopes exist | their analysts run in parallel and return the same columns |
| An artifact analyst cannot receive isolated context | every scope is analyzed locally instead of duplicating the transcript |
| A type is shown | it is `skill`, `behavior`, `knowledge`, or `tooling` |
| Information is assessed | it has one allowed label backed by evidence |
| Information is correct | its scope says `no change` and no recommendation is rendered |
| A named target is assessed | the target file was read before the verdict |
| A saving is shown | it is categorical and never an invented amount |
| The same evidence is analyzed again | recommendations keep the same order and do not cite an earlier `improve` report |
- What should be removed or clarified? What was counterproductive?
- Where should change live? How could it save time or tokens?
- Which back-and-forth, bottlenecks, or tool calls could be removed, batched, parallelized, or replaced?
5. **Assess.** Use `obsolete`, `over-specific-or-time-bound`, `duplicate`, `inconsistent`, `counterproductive`, or `correct`.
- Compare observed work with the shortest reliable path to the requested result. Keep checks for mutable state or unresolved uncertainty; reuse still-valid results instead of repeating research.
- Raw call count never proves waste. Ground the smallest general correction in evidence.
- Recommend memory only for concise, durable knowledge that prevents recurring rediscovery, never transient state or one-off implementation details.
- For skills, especially knowledge: delete evidence-backed waste or inconsistency first, without quota; then consolidate or clarify.
- Reuse or strengthen existing content; add only for a demonstrated gap it cannot cover. Preserve useful context and requirements; classify sufficient content as `correct` => `no change`.
- Claim time savings from blocking-time or resource-impact evidence, never background-process uptime alone.
6. **Deliver.** Deduplicate, verify cited evidence, and order by focus, then `behavior`, `skill`, `knowledge`. Report findings, not the questionnaire, in the fewest actionable words.
- Every recommendation needs an exact maintained source path, current excerpt, minimal correction or explicit `+`/`−` edit, and a short target/action prompt.
- If no source resolves or no useful persistent edit exists, report a non-actionable limitation.
- Type: `skill`, `behavior`, `knowledge`, or `tooling`.
- Saving: `time`, `tokens`, `both`, or `unknown`.
Original file line number Diff line number Diff line change
@@ -1,35 +1,20 @@
# 03 - Target edits

Map recommendations to minimal edits and render the report.
Map minimal edits and render the report.

## Input

The timing, usage, and recommendations tables.
The timing, usage, recommendations, and scope index tables.

## Output

An HTML report in a temporary directory with local `report.css` and `report.js`, plus its path and the next-intent question.
An HTML report in a temporary directory with local `report.css` and `report.js`, plus its path and a next-intent request.

## Process

1. **Target.** Map each file-targeted recommendation to a real project path.
- Omit behavior-only recommendations from this table.
2. **Summarize.** Build a `Fichier | + Ajout | − Retrait ou clarification` table.
- Consolidate repeated targets and render each diagnostic's smallest change.
- Render one `Aucun fichier recommandé | — | —` row when no recommendation needs a file edit.
3. **Render.** Fill [the report template](../assets/report-template.html) with only measured values and grounded findings, then copy its local CSS and JavaScript beside it.
- Remove every sample value and sample finding from the produced report.
- HTML-escape every injected value; allow only template-owned markup and local asset references.
- Write only to a unique temporary directory, never the project.
4. **Return.** Provide the report path and end with this exact question: `Quel changement d’intention général, même minime, appliquons-nous au prochain run pour rendre notre amélioration cumulative et mesurable ?`

## Test

| Case | Pass |
| --- | --- |
| A file is shown | it is the target of an earlier recommendation |
| No file needs editing | the empty-table row is shown |
| The report is rendered | its HTML, CSS, and JavaScript resolve locally with no network dependency |
| Conversation evidence is rendered | it is escaped as text and cannot create markup, scripts, or URLs |
| A value is unavailable | the report says `unavailable` instead of showing sample data |
| The run closes | the returned chat response ends with the exact next-intent question |
1. **Target.** Consolidate the provided recommendations by source into `File | + Added | − Removed or clarified`; use `No recommended file | — | —` when empty. Never resolve, judge, or apply sources here.
2. **Render.** Fill [the report template](../assets/report-template.html) from the provided data; remove its sample and copy its local CSS and JavaScript into the allowed unique temporary directory.
- Escape injected values; allow only template markup and local assets. Localize the HTML `lang`, visible text, and dynamic `data-label-*` text to the user's language.
- Put findings first with diagnostic/type tags, filters, accept controls, the target table, and an editable global prompt. Copy each recommendation's short target/action prompt into `data-prompt` without redrafting it.
- Show source paths and minimal edits. Collapse raw evidence, secondary metrics, timing, coverage, source records, and limitations; keep the run ID in source records and avoid repeated metadata.
3. **Return.** Return its path. Ask the user, in their language, which small general change in intent to apply next run for cumulative, measurable improvement.
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,3 @@ Use this static routing map to load one complete conversation. Do not fill or co
| Codex | host thread reader with the exact thread ID, then TUI `/export` Markdown | `CODEX_HOME/history.jsonl`, normally `~/.codex/history.jsonl` | TUI `/status` estimated thread credits or cost; tool elapsed records when exposed |
| Claude Code | `/export` text, or the documented script interface for an exact session ID | `~/.claude/projects/{project}/{session-id}.jsonl` | transcript timestamps and tool records when exposed |
| OpenCode | `opencode export {session-id} --sanitize` JSON | `~/.local/share/opencode/project/{project-slug}/storage/` | `opencode stats`; `~/.local/share/opencode/log/` |

## Rules

- Prefer the complete export or host reader over direct storage parsing.
- Use direct storage only after matching one exact session ID.
- Search shared history or logs only for the exact session ID and ignore unrelated records.
- Never enumerate unrelated sessions or inspect configuration or authentication files.
- Treat private reasoning duration as unavailable unless the host exposes it directly.
Loading
Loading