
Google WikiSkill gives AI agents long-term memory by storing detailed experience in a wiki and converting relevant lessons into compact executable skills. This lets agents reuse successful strategies and avoid known failures without continuously expanding their production prompts.
Key takeaways
- WikiSkill separates rich long-term knowledge from the short instructions used during task execution.
- The Wiki Layer stores successful strategies, recurring failures, conditions, and recovery procedures extracted from execution traces.
- The Skill Layer converts relevant wiki knowledge into focused procedures that agents can use at runtime.
- Evolution logs and impact tracking measure whether skill updates improve, preserve, or reduce agent performance.
- Traceable links between skills, wiki patterns, and execution evidence make self-improvement more auditable.
AI agents can complete complex tasks, but they often struggle to learn from previous attempts. An agent may encounter a failed tool call, discover a reliable workaround, and then forget both experiences when its current session ends. The usual solution is to place more information in the prompt. That can make the system slower, more expensive, and harder to manage.
Google WikiSkill takes a different approach. Described as a framework from Google Research and Virginia Tech, it separates an agent's detailed long-term knowledge from the compact instructions it needs during execution. The result is a two-layer memory system: a wiki preserves what the agent has learned, while procedural skills provide focused guidance at runtime.
Why AI Agents Need Memory Beyond the Current Prompt
An execution trace contains more than an answer. It can show which tools the agent used, where it failed, what assumptions proved incorrect, and which sequence of actions eventually worked. Across many tasks, these traces reveal recurring failure modes and successful strategies.
Most agents do not automatically turn that history into reusable knowledge. If each new task begins with only the latest prompt and a limited set of instructions, the agent may repeatedly make the same mistake. It can also spend inference time rediscovering a solution it has already found.
Expanding the production prompt is not an ideal fix. A large prompt can increase context-processing costs and introduce irrelevant information. It may also force the agent to search through detailed historical records when it needs only one narrow procedure, such as how to recover from a particular API error.
WikiSkill addresses this problem by preserving detailed experience outside the execution context and exposing only the most useful procedures when needed.
How the WikiSkill Framework Separates Knowledge From Execution
WikiSkill has two connected components: the Wiki Layer and the Skill Layer.
The Wiki Layer stores successful strategies and recurring failure modes
The Wiki Layer acts as structured procedural memory for AI agents. It organizes information extracted from execution traces into pages that describe patterns such as:
- A recurring tool-use failure and its likely cause
- Conditions under which a strategy succeeds
- A sequence of actions that reliably completes a task
- A failed approach that should be avoided
- A recovery method for a known error
This layer can retain rich detail. For example, a page might record that an agent failed when it called a tool before validating an identifier, then succeeded after checking the identifier format and retrying with a corrected value. The page preserves the lesson without forcing every future task to carry the entire trace.
The wiki therefore becomes a growing record of the agent's experience. It can contain both positive and negative knowledge, which is important because avoiding a known failure may be as valuable as repeating a successful action.
The Skill Layer delivers compact procedures at runtime
The Skill Layer translates relevant wiki knowledge into executable instructions. These skills are narrower than the underlying knowledge base. Instead of presenting every related trace, the system might provide a short procedure such as: validate the identifier, call the lookup tool, inspect the response, and retry only when the error matches a defined condition.
The production agent receives these compact skills during task execution. It does not need to read the full wiki or reason through every historical example. This separation lets the system retain detailed memory while keeping the active context focused.
How WikiSkill Turns Execution Traces Into Better Skills
WikiSkill uses accumulated execution evidence to identify knowledge that can be reused. Full traces are converted into wiki pages covering successful strategies and recurring failure modes. A separate agent then uses those pages to propose narrow updates to executable skills.
This division matters. The agent performing the task is not required to rewrite its own entire prompt after every attempt. Instead, another process examines the accumulated knowledge and produces a targeted procedural change. One update might add a precondition, change the order of tool calls, or introduce a recovery step for a known failure.
Evolution logs and impact tracking show whether changes work
Skill updates are recorded through an evolution log and a skill-impact tracker. The evolution log captures proposed changes and the knowledge behind them. The impact tracker records what happened after an update was applied, including whether performance improved, stayed the same, or declined.
This creates a feedback loop:
- The agent attempts a task and produces an execution trace.
- The system extracts useful patterns and failures into wiki pages.
- A separate agent proposes a focused skill modification.
- The modified skill is evaluated on subsequent tasks.
- The result is recorded and used to guide future changes.
Tracking impact helps prevent uncontrolled skill evolution. A change that sounds sensible may not improve outcomes in practice. By recording its effect, the framework can distinguish useful adaptations from changes that should be rejected or revised.
Links between skills and wiki patterns make updates auditable
Each executable skill can link back to the wiki patterns that motivated its creation or modification. That connection provides a traceable path from experience to instruction.
If a skill tells an agent to verify a value before using a tool, an evaluator can inspect the underlying failure pattern and the execution traces that support the recommendation. If the skill later causes new problems, the team can identify which knowledge page and update led to the behavior.
This is an important auditability benefit. The system preserves execution traces, knowledge pages, proposed improvements, and measured outcomes instead of treating the agent's behavior as an opaque result of a continually changing prompt.
Why Externalized Memory Can Reduce Inference Costs
Keeping the full wiki outside the production agent's context can reduce runtime overhead. The agent does not repeatedly process all historical attempts, and its prompt does not grow every time the system learns something new.
The cost advantage comes from selectivity. The external memory can be extensive, but the execution agent receives only compact skills relevant to its task. This can reduce context length and limit the amount of information the model must interpret before acting.
The approach also avoids a common weakness of deriving skills from isolated trajectories. If each update is based only on the latest attempt, the system may rediscover failures it has already encountered. WikiSkill gives the improvement process access to a cumulative record, allowing new updates to build on prior knowledge rather than starting from a single example.
External memory does not eliminate inference costs. The system still needs processes to retrieve, organize, evaluate, and convert wiki knowledge into skills. However, those costs can be separated from the main execution path and managed as an improvement workflow instead of being paid through an ever-larger prompt on every task.
What Google and Virginia Tech's Research Means for Self-Improving Agents
WikiSkill presents a practical model for agentic AI skills and long-term learning. Its central idea is that knowledge and action instructions should not be stored in the same form.
The wiki can preserve detailed, searchable experience. Skills can remain short, operational, and task-focused. Together, they provide a form of procedural memory for AI agents: the system remembers not only what happened, but how that experience should change future behavior.
The framework's broader significance lies in targeted adaptation. Instead of asking an agent to absorb every past interaction, it converts selected lessons into constrained procedures and measures their effects. That can support more reliable tool use, clearer debugging, lower context overhead, and better oversight of self-improving systems.
WikiSkill is not a guarantee that an agent will always learn the right lesson. Poorly extracted patterns or incorrectly evaluated updates could still lead to bad skills. Its contribution is a structured way to manage that risk by preserving the evidence behind changes and separating rich memory from runtime execution.
As AI agents take on longer, more complex workflows, systems that can remember failures without carrying every failure in the prompt may become increasingly important. Follow the latest developments in agentic AI and AI agent memory to see how systems are becoming more reliable, efficient, and auditable.
By the numbers
WikiSkill uses two connected layers: a Wiki Layer for detailed procedural memory and a Skill Layer for compact runtime instructions.
Architecture described in the Google Research and Virginia Tech WikiSkill framework discussed in the article.
The framework's improvement loop contains five stages: trace capture, knowledge extraction, skill modification, task evaluation, and impact recording.
The five-stage workflow is summarized from the article's execution-to-feedback process.
Three core evidence artifacts support traceability: wiki pages, evolution logs, and skill-impact records.
These artifacts are identified in the article's discussion of memory organization, skill updates, and auditability.
Step by step
- 01
Capture execution traces
Record tool calls, failed assumptions, successful action sequences, error conditions, and recovery attempts from agent workflows.
- 02
Extract reusable knowledge
Convert recurring failure modes and successful strategies into structured wiki pages with conditions, causes, and recommended actions.
- 03
Generate focused skills
Use the accumulated wiki knowledge to propose narrow procedural updates, such as new preconditions, tool-call ordering, or recovery steps.
- 04
Evaluate skill changes
Test each updated skill on subsequent tasks and compare outcomes with the previous version to determine whether performance improves.
- 05
Track and audit impact
Store proposed changes, supporting wiki patterns, measured results, and reversions in evolution and impact records.
Frequently asked questions
What is Google WikiSkill?
Google WikiSkill is a framework for helping AI agents learn from prior execution experiences without expanding their runtime prompts. Developed in work associated with Google Research and Virginia Tech, it separates detailed knowledge in a wiki from compact procedures called skills.
How does WikiSkill help AI agents learn from mistakes?
WikiSkill records recurring failures and their successful recovery strategies in structured wiki pages. A separate improvement process uses that knowledge to update executable skills, allowing future agents to avoid known errors instead of rediscovering the same solution.
Does WikiSkill make AI agent prompts larger?
No, WikiSkill is designed to keep runtime prompts focused by retrieving only relevant compact skills. Detailed execution traces and historical knowledge remain outside the production context, although the system still incurs costs for organizing, retrieving, and evaluating that memory.
What is the difference between the Wiki Layer and Skill Layer?
The Wiki Layer stores detailed procedural knowledge, while the Skill Layer delivers concise instructions for execution. The wiki may contain traces, failure patterns, and alternative strategies; a skill distills the most relevant lesson into an operational procedure.
How does WikiSkill make self-improving agents auditable?
WikiSkill links skill changes to the wiki patterns and execution evidence that motivated them. Evolution logs record proposed updates, impact tracking records their results, and these connections help teams inspect or reverse changes that produce unexpected behavior.



