Capability Evolution

Runtime signals and production changes are separate layers

Task traces, user feedback, and deliverables describe what happened; they do not identify what should change. The same failure may come from missing context, stale knowledge, a tool error, a permission boundary, or model judgment. Appending the trace to a global Prompt would turn an incident into a permanent rule.

aibuddy therefore does not expose one unrestricted self-rewriting learner. It first routes a signal to the update surface with the right scope, then lets that surface’s lifecycle decide when content becomes effective in later tasks. The changed object, impact scope, and release action remain explicit instead of accumulating in a growing system prompt.

aibuddy capability evolution change control Runtime signals are stored as evidence and routed by scope to user memory, knowledge, Skills, or Prompts. Each surface has its own activation boundary; unified attribution, cross-version evaluation, canary release, and rollback are not yet closed. Runs produce signals; releases change capability There is no automatic write path from one negative rating to a global Prompt. Runtime evidence Task trace and state Message feedback Deliverables and failures Step metrics and usage Route by ownership and execution semantics Evidence remains a signal; the update surface defines scope and activation. User memory One user · durable · not retrievable Knowledge base Sourced facts and document structure Skill Capability package loaded on demand Prompt Runtime configuration for one Agent Guarded automatic write Gate · exclusion · deduplication Manual memory wins Reviewed visibility pending → ready Outline and summary re-derived Explicit upload revision .version ↑ · cache invalidation Not a historical snapshot Immutable published version draft → published → active Activation enters runtime instructions Only effective state enters the next run Local update paths are controlled and do not block the current task. No unified system loop yet Attribution · eval · canary · rollback

Information ownership and execution semantics select the update surface

SurfaceAppropriate contentActivation boundary
User memoryOne user’s durable preferences, identity, and correctionsWritten in user scope; manually authored content takes precedence over automatic content
Knowledge baseSourced facts, documents, and document structureReingestion returns a document to pending; only ready documents are searchable by an Agent
SkillProcedures, policies, and references loaded on demandAn explicit upload advances the revision and invalidates metadata caches
PromptOne Agent’s role, boundaries, and default behaviorA draft must be published before activation; activation updates the Agent’s runtime instructions
Program and toolCross-task rules that require deterministic execution or verificationAn engineering release changes every applicable execution path

The preferred surface has the smallest scope and the clearest verification method. Missing internal material belongs in knowledge; a deterministic check belongs in program logic or a tool; only stable behavioral constraints belong in a Prompt.

User memory provides a guarded online update path

Memory is the only surface that can currently produce content automatically after an ordinary task, and it is constrained before and after generation:

  1. Pre-call gate: deterministic cues first decide whether the recent conversation may contain a preference, identity fact, long-running project, or explicit correction. Without a cue, no extraction model is called.
  2. Bounded judgment window: the extractor reads only the latest three user–assistant pairs. A memory must be about the user, remain useful next week, and not be recoverable from files or connected systems.
  3. Mutual exclusion with explicit writes: extraction is skipped when the Agent already used memory_save, memory_update, or memory_delete in the turn, preventing competing write paths.
  4. Pre-write reconciliation: low-salience candidates are dropped; same-type content is reconciled with character-bigram similarity; equivalent content is a no-op; automatic content never overwrites a manual memory.
  5. Non-blocking execution: extraction runs in the background. A failure is recorded without changing the task result.

The objective is not maximum retention. It is a small set of explainable, reusable user facts maintained without extending the critical task path.

Knowledge updates are separated from visibility

A knowledge update does not immediately replace material visible to an Agent. Import or reingestion first normalizes the Markdown, rebuilds its heading outline, and moves the document to pending. The content is available for management review, but it enters the Agent’s knowledge map and search scope only after becoming ready.

Reingestion also clears the summary derived from the prior body. The replacement summary is a non-blocking enhancement: slow or failed generation cannot undo the content write, and the service rereads the document before persisting a result so a late response cannot overwrite a reviewer-authored summary. Body, derived outline, summary, and publication status therefore have an explicit update order.

Skills and Prompts use different release semantics

Both carry reusable behavior, but their lifecycles are deliberately different:

SkillPrompt
Change operationUpload a complete file tree and validate root SKILL.md, frontmatter, and pathsCreate a new draft version for one Agent
Runtime visibilityUpload advances .version and invalidates metadata caches; an Agent can observe the revision and reloadOnly a published version can activate; activation synchronizes content to agent.instructions
Immutability boundaryThe current implementation replaces same-name content; the revision is not a historical snapshotPublished content cannot be edited in place, and the active version cannot be deleted
Current gapNo built-in history comparison or rollbackNo automated evaluation, canary release, or automatic rollback

The distinction reflects their roles: a Skill is a deployable capability package, while a Prompt is Agent configuration. A revision tag is not presented as a version repository, and activation is not presented as proof of quality.

Message feedback is evidence, not an optimizer

Positive, negative, and cleared feedback is stored against a specific message and task. It can be analyzed with the trace, step states, usage, and deliverables. One negative rating does not automatically modify memory, a Prompt, a Skill, or a tool.

A system-wide evolution loop would additionally need to turn signals into fixed cases and link each candidate change to baseline comparison, release observation, and rollback. aibuddy already provides layered evidence and several controlled update surfaces; automatic attribution, cross-version evaluation, canary release, and unified rollback remain explicit engineering boundaries.

Implementation anchors

  • Memory extraction: deterministic gate → bounded extraction → salience and near-duplicate reconciliation → background write.
  • Knowledge publication: persist body and outline → generate summary asynchronously → control Agent visibility with review status.
  • Skill revision: validate upload → issue a server-side monotonic revision → invalidate cache → reload on demand at runtime.
  • Prompt release: draft → publish → activate; published content is immutable.
  • Runtime feedback: persist at message level while remaining decoupled from capability release.
Was this page helpful?