SkillRC: Is It the Skill or the Stack?
Auditing Experience Memory in Agent Harnesses

Yiwen Lu

University of Notre Dame · ylu37@nd.edu

Manuscript · October 2026

Abstract

Agent harnesses increasingly learn from experience: past executions are distilled into insights, skills, or playbooks that are injected back into the context of a frozen model. Such memories are usually credited by comparing an agent with the memory to one without it. That comparison cannot tell whether the gain carries over to a different executor, or whether it comes from what the memory says rather than from simply adding context. SkillRC, a reality check (RC) for agent skills, treats the memory as a controlled intervention on a fixed agent and asks two questions in turn: does the memory's gain transfer across executors, and does it survive a placebo that keeps the memory's rules, order, and length but scrambles its words? On ALFWorld, a compact skill memory clearly improves GPT-4o-mini but not Qwen3-8B, and the placebo reproduces most of this difference: the gap reflects how each executor responds to added context more than how it uses the skills.

Two-phase SkillRC overview. Phase A: distill experience into a skill memory, insert it into an otherwise unchanged agent, run two executors with and without the memory, and compare their gains (the transport gap). Phase B: build a placebo with the same rules, order and length but scrambled words, run true versus placebo, split the gain into a context effect and a content effect, and decide by a rule fixed in advance.
Overview of SkillRC. Phase A asks whether a memory's gain transfers across executors; Phase B asks whether the gain comes from what the memory says, by splitting it into a context effect (no memory to placebo) and a content effect (placebo to true memory).

Key numbers

ALFWorld, 134 unseen tasks, 33 ExpeL-style insights (645 tokens), BM25 presentation, three demonstration pairs weighted equally; 95% task-bootstrap intervals.

+0.127
GPT-4o-mini gain with memory
[+0.072, +0.184]
−0.021
Qwen3-8B gain with memory
[−0.062, +0.020]
−0.148
Transport gap (Qwen − GPT)
[−0.218, −0.080]
+0.017
Gap in content increment (vs. placebo)
[−0.075, +0.108]

Problem

Experience-driven agents convert past executions into insights, skills, or playbooks that re-enter the model's context. The evidence that such a memory helps is usually a with-versus-without comparison on one base model, and the memory is rarely separated from the other changes it brings to the prompt. Three questions stay open:

Harness factors

Each comparison varies one factor on identical tasks. The agent loop (ReAct, admissible actions, step limit), decoding (temperature 0), and action parsing are shared by every configuration and are not factors.

FactorRoleLevels in this study
EExecutor stackvariedGPT-4o-mini (hosted API), Qwen3-8B (vLLM, thinking off)
cMemory contentvariednone, true insights, order placebo
rRepresentationvaried33 compact insights, raw trajectories, Markdown procedures
pPresentationvariedBM25, dense, reference-ranked placebo (order, not membership)
BFootprintmeasured, matched2,048-token budget; 645 tokens realized by the insights
dDemonstration pairaveragedpairs 01, 02, 12, weighted equally

Results

RQ1Does experience memory improve the harness? Yes, for the stack it was built with. GPT-4o-mini rises from 0.587 to 0.714 on every demonstration pair, with the most compact memory.

RQ2Does the improvement transfer? No. Qwen3-8B moves from 0.704 to 0.683, and the transport gap of −0.148 excludes zero.

ALFWorld success without and with memory: GPT-4o-mini 0.587 to 0.714, Qwen3-8B 0.704 to 0.683. Paired gain per demonstration pair: GPT positive on all pairs, Qwen near zero or negative; equal-pair mean +0.127 versus −0.021.
RQ1–RQ2. Success without and with the memory, and the paired gain on each demonstration pair with 95% intervals on the mean.
Observation. The second stack starts from higher no-memory success (70.4% vs. 58.7%) and responds to the added block differently; the gap identifies stack-level moderation, not which stack component causes it.

RQ3Is the improvement driven by experience content? Only partly. A placebo with the same items, order, and 645 tokens but no original word pair improves GPT by +0.089 and harms Qwen by −0.108. Coherent content adds +0.047 and +0.064, a gap of +0.017 that crosses zero.

Order-placebo gate on 60 frozen tasks: success under no memory, placebo, and true memory per stack, and paired effects with intervals.
RQ3. Pre-registered order-placebo gate on 60 frozen tasks. The grey band is the frozen equivalence region [−0.05, +0.05].
Observation. Under the frozen rule the content mechanism is inconclusive. In point estimates the placebo alone opens a stack gap of −0.197, as large as the true memory's −0.181.

RQ4Where does transfer succeed or fail? Only 9 of 134 tasks improve under both stacks; 45 improve under GPT but not Qwen, and 13 of the 14 GPT-help/Qwen-hurt tasks have a high Qwen baseline.

Task-level transport map by sign of effect under each stack, and GPT gain by task type.
RQ4. Tasks by the sign of their paired effect under each stack, and GPT gain by task type.

RQ5Is more context better, and does it hold elsewhere? No. Trajectory and procedural memories are about three times larger and never beat the compact insights, and in WebShop a ten-rule memory lowers Qwen3-8B reward from 0.194 to 0.105 (8 improvements, 20 declines, 72 ties).

WebShop: reward and exact success without and with memory, and the sign of paired reward change over 100 sessions.
RQ5. Train-disjoint WebShop comparison with local Qwen3-8B.

What the tasks look like

ALFWorld gives the agent text only, but each game comes from an ALFRED trial in the AI2-THOR simulator, and the 134 unseen tasks take place in just four rooms. Each figure renders the trial's room (AI2-THOR 2.1.0, no model run) and labels the places named in the agent's observation: orange is where the object starts, purple the appliance, green the destination, and the map arrow is the expert's route. Cases 4–6 show an example task of the family.

Task 0: put a mug in desk

Task 0, put a mug in desk: rendered bedroom with bed, desks, shelves, drawers, safe, laundry hamper, and garbage can labelled; the mug is on a shelf of the desk hutch and goes to the other desk; GPT-4o-mini solves it in 19 steps without memory and 6 with memory.
The mug sits on a shelf of the desk hutch and must go to the other desk; the expert needs 4 steps. With memory, GPT-4o-mini still solves the task but in 6 instead of 19 steps, so success rate alone does not show the gain.

Task 1: heat some egg and put it in garbagecan

Task 1, heat some egg and put it in garbagecan: rendered kitchen with the egg on a countertop, the microwave, and the garbage can labelled; expert solution 6 steps; GPT-4o-mini fails at the 30-step limit with and without memory.
The egg lies on a countertop, must be heated in the microwave, and ends in the garbage can (6 expert steps). The matching rule is the first item of the memory, yet GPT-4o-mini still fails within 30 steps, at a higher token cost.

Task 2: put two pillow in sofa

Task 2, put two pillow in sofa: rendered living room with both pillows on the armchair and the sofa labelled; expert solution 8 steps; GPT-4o-mini fails at 30 steps in both conditions.
Both pillows sit on the armchair next to the sofa, so the expert repeats one 4-step routine. GPT-4o-mini does not finish in 30 steps either way, although the memory has rules about tracking the second item.

The clean family (31 tasks)

Clean family example: rendered bathroom with the soap bar on the toilet, the sink basin, and the countertop labelled; bars of the memory effect on tasks 30, 57 and 50 for both executors.
An example task: the soap bar starts on the toilet, is cleaned at the sink basin, and goes on the countertop. The memory lifts GPT-4o-mini from 0.452 to 0.667 on this family while Qwen3-8B moves from 0.656 to 0.613.

The cool family (21 tasks)

Cool family example: rendered kitchen with the mug on a countertop, the fridge, and an upper cabinet labelled; bars of the memory effect on tasks 5, 16, 19 and 39 for both executors.
An example task: the mug starts on a countertop, is cooled in the fridge, and goes into a cabinet. This is GPT-4o-mini's strongest family without memory, and the only one where the memory hurts it (0.857 to 0.746).

The examine family (18 tasks)

Examine family example: rendered bedroom with the book on the bed and the desk lamp labelled; bars of the memory effect on tasks 7, 45 and 89 for both executors.
An example task: the book lies on the bed and must be looked at under the desk lamp. Task 7 improves under both executors, while task 45 of the same family helps GPT-4o-mini and hurts Qwen3-8B.

WebShop: 100 train-disjoint sessions

WebShop screenshots: start page with the request, search results with the requested pillow cover 5th, product page with the 28 by 28 inch size and Buy Now, score page with reward 1.0; with memory Qwen3-8B mean reward drops from 0.194 to 0.105.
Screenshots of the official web app on session 0: the request, a search, the product page where the size is chosen, and the score page (reward 1.0 for a full match; the query and clicks are ours). A ten-rule memory distilled from six successful sessions lowers Qwen3-8B's mean reward from 0.194 to 0.105.

Get started

The unit tests run the whole harness on a mock executor and environment, with no API calls and no GPU.

git clone https://github.com/EvenEureka/SkillRC.git && cd SkillRC
pip install -r requirements.txt
python tests/test_smoke_mock.py

# rebuild the order placebo and audit its footprint
PYTHONPATH=. python scripts/analysis/build_order_placebo.py

# ALFWorld with a local Qwen3-8B served by vLLM (thinking off)
alfworld-download && bash scripts/fetch_alfworld_prompts.sh
vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --port 8000 --max-model-len 16384
python -m skillrc.runner --config configs/papero_placebo_qwen.yaml --max-tasks 3

BibTeX

@misc{skillrc2026,
  title  = {SkillRC: Is It the Skill or the Stack? Auditing Experience Memory in Agent Harnesses},
  author = {Lu, Yiwen},
  year   = {2026},
  note   = {Manuscript},
  url    = {https://eveneureka.github.io/SkillRC/}
}