The issue with most AI agents is that they lose context. You set up an agent, use it for a week or two, and then you have to explain the context again, correct mistakes, and the loop continues. This is what Hermes, the open-source agent from Nous Research, is solving with a specific mechanism.
It compiles multi-step execution trajectories into standardized SKILL.md files. So instead of re-deriving a workflow from scratch and burning tokens to do so, the agent recognizes “I have done this before” and loads a compact, purpose-built instruction file and executes. Over time, those files get corrected and refined through use. This is what Hermes calls its procedural memory layer. But how do you build a self-improving AI agent with a Hermes SKILL.md file? This article explains this further.

Understand the two memory layers

Before configuring anything, let’s understand the two memory layers, which are:
- Declarative memory (MEMORY.md + USER.md): Compact facts that should always be in context, such as the preferred communication style, stable project paths, deployment conventions, and corrections you don’t want repeated.
- Procedural memory (skills): Step-by-step know-how for a specific type of task, such as the commands to run, the pitfalls to avoid, and how to verify success. Skills only load when relevant. Plus, they can be far more detailed than a memory entry without bloating every session.
In this guide, we’ll build the second layer.
What is procedural memory in Hermes?
Procedural memory is essential memory for how to do something. For instance, you can ask an agent to research several sources, compare the information, extract key findings, create a structured report, etc.
The first time, Hermes may require multiple tool calls, corrections, and reasoning steps to complete the task.
Instead of discarding that execution, Hermes can preserve the useful workflow as a reusable skill.
Hermes documentation describes agent-managed skills as its procedural memory. When the agent discovers a non-trivial workflow, encounters a useful workaround, or learns from a correction, it can create or update a SKILL.md file containing that procedure.
A skill can include:
- When the procedure should be used
- Step-by-step instructions
- Tools or commands required
- Common failure cases
- Workarounds
- Verification steps
- Constraints and safety notes
How to build self-improving AI agents with Hermes SKILL.md procedural memory?
To build a self-improving AI agent with Hermes, focus on the core idea: letting the agent turn useful experiences from completed tasks into reusable procedures.
Instead of storing an entire conversation or execution history, Hermes can distill the key parts of a successful multi-step workflow into a SKILL.md file.md file. When a similar task appears later, the agent can load that skill and follow the proven procedure rather than solving the workflow from scratch.
How is a skill created?

The agent itself writes the skill via a skill_manage tool (create, patch, edit, delete). It triggers automatically when you complete a complex task, hit errors, receive a user correction, etc.
As a user, you don’t need to do any setup. Just use the agent on real multi-step tasks. You can also compile a skill on demand with /learn <description, doc, or codebase> instead of waiting for organic discovery.
Why it saves tokens: progressive disclosure
Skills do not load all at once. Instead, they load in three progressively deeper stages. At the lightest level, Hermes keeps only an index in context — just the name and description of every installed skill. When a task actually calls for a particular skill, Hermes pulls in the next level, loading that skill’s full content. After that, if that skill points to deeper reference material — a longer guide, an example file, documentation for an edge case — that gets loaded as well.
Overall, Hermes’ SKILL.md system turns one-off problem solving into a long-lasting capability. The agent itself captures what worked, refines it with each use, and only loads what’s needed to keep costs low. The result is an agent that gets measurably better at recurring tasks over time, rather than starting from zero every session.