Capability Evolution
Runtime signals and production changes are separate layers
Task traces, user feedback, and deliverables describe what happened; they do not identify what should change. The same failure may come from missing context, stale knowledge, a tool error, a permission boundary, or model judgment. Appending the trace to a global Prompt would turn an incident into a permanent rule.
aibuddy therefore does not expose one unrestricted self-rewriting learner. It first routes a signal to the update surface with the right scope, then lets that surface’s lifecycle decide when content becomes effective in later tasks. The changed object, impact scope, and release action remain explicit instead of accumulating in a growing system prompt.
Information ownership and execution semantics select the update surface
| Surface | Appropriate content | Activation boundary |
|---|---|---|
| User memory | One user’s durable preferences, identity, and corrections | Written in user scope; manually authored content takes precedence over automatic content |
| Knowledge base | Sourced facts, documents, and document structure | Reingestion returns a document to pending; only ready documents are searchable by an Agent |
| Skill | Procedures, policies, and references loaded on demand | An explicit upload advances the revision and invalidates metadata caches |
| Prompt | One Agent’s role, boundaries, and default behavior | A draft must be published before activation; activation updates the Agent’s runtime instructions |
| Program and tool | Cross-task rules that require deterministic execution or verification | An engineering release changes every applicable execution path |
The preferred surface has the smallest scope and the clearest verification method. Missing internal material belongs in knowledge; a deterministic check belongs in program logic or a tool; only stable behavioral constraints belong in a Prompt.
User memory provides a guarded online update path
Memory is the only surface that can currently produce content automatically after an ordinary task, and it is constrained before and after generation:
- Pre-call gate: deterministic cues first decide whether the recent conversation may contain a preference, identity fact, long-running project, or explicit correction. Without a cue, no extraction model is called.
- Bounded judgment window: the extractor reads only the latest three user–assistant pairs. A memory must be about the user, remain useful next week, and not be recoverable from files or connected systems.
- Mutual exclusion with explicit writes: extraction is skipped when the Agent already used
memory_save,memory_update, ormemory_deletein the turn, preventing competing write paths. - Pre-write reconciliation: low-salience candidates are dropped; same-type content is reconciled with character-bigram similarity; equivalent content is a no-op; automatic content never overwrites a manual memory.
- Non-blocking execution: extraction runs in the background. A failure is recorded without changing the task result.
The objective is not maximum retention. It is a small set of explainable, reusable user facts maintained without extending the critical task path.
Knowledge updates are separated from visibility
A knowledge update does not immediately replace material visible to an Agent. Import or reingestion first normalizes the Markdown, rebuilds its heading outline, and moves the document to pending. The content is available for management review, but it enters the Agent’s knowledge map and search scope only after becoming ready.
Reingestion also clears the summary derived from the prior body. The replacement summary is a non-blocking enhancement: slow or failed generation cannot undo the content write, and the service rereads the document before persisting a result so a late response cannot overwrite a reviewer-authored summary. Body, derived outline, summary, and publication status therefore have an explicit update order.
Skills and Prompts use different release semantics
Both carry reusable behavior, but their lifecycles are deliberately different:
| Skill | Prompt | |
|---|---|---|
| Change operation | Upload a complete file tree and validate root SKILL.md, frontmatter, and paths | Create a new draft version for one Agent |
| Runtime visibility | Upload advances .version and invalidates metadata caches; an Agent can observe the revision and reload | Only a published version can activate; activation synchronizes content to agent.instructions |
| Immutability boundary | The current implementation replaces same-name content; the revision is not a historical snapshot | Published content cannot be edited in place, and the active version cannot be deleted |
| Current gap | No built-in history comparison or rollback | No automated evaluation, canary release, or automatic rollback |
The distinction reflects their roles: a Skill is a deployable capability package, while a Prompt is Agent configuration. A revision tag is not presented as a version repository, and activation is not presented as proof of quality.
Message feedback is evidence, not an optimizer
Positive, negative, and cleared feedback is stored against a specific message and task. It can be analyzed with the trace, step states, usage, and deliverables. One negative rating does not automatically modify memory, a Prompt, a Skill, or a tool.
A system-wide evolution loop would additionally need to turn signals into fixed cases and link each candidate change to baseline comparison, release observation, and rollback. aibuddy already provides layered evidence and several controlled update surfaces; automatic attribution, cross-version evaluation, canary release, and unified rollback remain explicit engineering boundaries.
Implementation anchors
- Memory extraction: deterministic gate → bounded extraction → salience and near-duplicate reconciliation → background write.
- Knowledge publication: persist body and outline → generate summary asynchronously → control Agent visibility with review status.
- Skill revision: validate upload → issue a server-side monotonic revision → invalidate cache → reload on demand at runtime.
- Prompt release: draft → publish → activate; published content is immutable.
- Runtime feedback: persist at message level while remaining decoupled from capability release.
Related reading
- Memory & Knowledge for scope, publication, and progressive disclosure across persistent context.
- Evaluation & Observability for the runtime evidence consumed by capability changes.
- Learning Signals for the distinction between retaining experience and building durable capability.
- Evolution Loop for the complete baseline, candidate, release, and rollback method.