SkillRC: Is It the Skill or the Stack?
Auditing Experience Memory in Agent Harnesses
University of Notre Dame · ylu37@nd.edu
Manuscript · October 2026
Abstract
Agent harnesses increasingly learn from experience: past executions are distilled into insights, skills, or playbooks that are injected back into the context of a frozen model. Such memories are usually credited by comparing an agent with the memory to one without it. That comparison cannot tell whether the gain carries over to a different executor, or whether it comes from what the memory says rather than from simply adding context. SkillRC, a reality check (RC) for agent skills, treats the memory as a controlled intervention on a fixed agent and asks two questions in turn: does the memory's gain transfer across executors, and does it survive a placebo that keeps the memory's rules, order, and length but scrambles its words? On ALFWorld, a compact skill memory clearly improves GPT-4o-mini but not Qwen3-8B, and the placebo reproduces most of this difference: the gap reflects how each executor responds to added context more than how it uses the skills.
Key numbers
ALFWorld, 134 unseen tasks, 33 ExpeL-style insights (645 tokens), BM25 presentation, three demonstration pairs weighted equally; 95% task-bootstrap intervals.
Problem
Experience-driven agents convert past executions into insights, skills, or playbooks that re-enter the model's context. The evidence that such a memory helps is usually a with-versus-without comparison on one base model, and the memory is rarely separated from the other changes it brings to the prompt. Three questions stay open:
- Does the gain transfer? The executor is a stack (weights, tokenizer, template, runtime, decoding).
- Is it the content? When the pool fits in the budget, retrievers change order, not membership, and any added block can move a model.
- Where does it fail? A positive mean can combine rescues with harm.
Harness factors
Each comparison varies one factor on identical tasks. The agent loop (ReAct, admissible actions, step limit), decoding (temperature 0), and action parsing are shared by every configuration and are not factors.
| Factor | Role | Levels in this study |
|---|---|---|
| EExecutor stack | varied | GPT-4o-mini (hosted API), Qwen3-8B (vLLM, thinking off) |
| cMemory content | varied | none, true insights, order placebo |
| rRepresentation | varied | 33 compact insights, raw trajectories, Markdown procedures |
| pPresentation | varied | BM25, dense, reference-ranked placebo (order, not membership) |
| BFootprint | measured, matched | 2,048-token budget; 645 tokens realized by the insights |
| dDemonstration pair | averaged | pairs 01, 02, 12, weighted equally |
Results
RQ1Does experience memory improve the harness? Yes, for the stack it was built with. GPT-4o-mini rises from 0.587 to 0.714 on every demonstration pair, with the most compact memory.
RQ2Does the improvement transfer? No. Qwen3-8B moves from 0.704 to 0.683, and the transport gap of −0.148 excludes zero.
RQ3Is the improvement driven by experience content? Only partly. A placebo with the same items, order, and 645 tokens but no original word pair improves GPT by +0.089 and harms Qwen by −0.108. Coherent content adds +0.047 and +0.064, a gap of +0.017 that crosses zero.
RQ4Where does transfer succeed or fail? Only 9 of 134 tasks improve under both stacks; 45 improve under GPT but not Qwen, and 13 of the 14 GPT-help/Qwen-hurt tasks have a high Qwen baseline.
RQ5Is more context better, and does it hold elsewhere? No. Trajectory and procedural memories are about three times larger and never beat the compact insights, and in WebShop a ten-rule memory lowers Qwen3-8B reward from 0.194 to 0.105 (8 improvements, 20 declines, 72 ties).
What the tasks look like
ALFWorld gives the agent text only, but each game comes from an ALFRED trial in the AI2-THOR simulator, and the 134 unseen tasks take place in just four rooms. Each figure renders the trial's room (AI2-THOR 2.1.0, no model run) and labels the places named in the agent's observation: orange is where the object starts, purple the appliance, green the destination, and the map arrow is the expert's route. Cases 4–6 show an example task of the family.
Task 0: put a mug in desk
Task 1: heat some egg and put it in garbagecan
Task 2: put two pillow in sofa
The clean family (31 tasks)
The cool family (21 tasks)
The examine family (18 tasks)
WebShop: 100 train-disjoint sessions
Get started
The unit tests run the whole harness on a mock executor and environment, with no API calls and no GPU.
git clone https://github.com/EvenEureka/SkillRC.git && cd SkillRC
pip install -r requirements.txt
python tests/test_smoke_mock.py
# rebuild the order placebo and audit its footprint
PYTHONPATH=. python scripts/analysis/build_order_placebo.py
# ALFWorld with a local Qwen3-8B served by vLLM (thinking off)
alfworld-download && bash scripts/fetch_alfworld_prompts.sh
vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --port 8000 --max-model-len 16384
python -m skillrc.runner --config configs/papero_placebo_qwen.yaml --max-tasks 3
BibTeX
@misc{skillrc2026,
title = {SkillRC: Is It the Skill or the Stack? Auditing Experience Memory in Agent Harnesses},
author = {Lu, Yiwen},
year = {2026},
note = {Manuscript},
url = {https://eveneureka.github.io/SkillRC/}
}












