Know2Act: What Knowledge Reaches
a Tool-Using Agent’s Actions?

Yiwen Lu

University of Notre Dame · ylu37@nd.edu

Manuscript · October 2026

Abstract

Recent work trains language agents to predict not only their own actions but also the observations that the environment returns, and reports that this makes later reinforcement learning (RL) more effective because the agent learns the consequences of its actions. We ask whether knowledge learned in this way reaches the agent’s decisions. In Know2Act, a pre-registered study of a multi-turn text-to-SQL agent trained with supervised fine-tuning and then GRPO, we measure what each model learns and how it acts. Observation supervision is learned: on unseen databases the agent predicts the format and headers of result tables far better. Yet it changes the agent’s actions less than reshuffling the training data does, and it brings no gain after RL. We then vary what the agent predicts and where, up to a contract that states the shape of the correct answer and is checked against every query result. Every prediction is learned, but none changes when the agent submits an answer or keeps querying. The knowledge is usable from outside: selecting among sampled answers with the agent’s own contract improves accuracy. RL, which rewards only the final answer, removes even this gain and makes the agent more willing to submit after its own check has failed. Learning to predict the environment is therefore not the same as acting on the prediction. We also find that 44% of the training tasks in a public SQL RL environment reward an empty answer.

One real trajectory for the question 'How many people live in Gelderland district?': a query returning 4 rows, a corrected SUM query returning 1 row, and the solution. Below it, five training variants and which segments carry loss: ActionSFT trains on actions only; ActObs also on observations; Outcome-Aux adds an 'Expected result' line as an auxiliary target; Outcome-in-Action puts that line inside the action; Contract starts with the expected answer shape and checks each result against it. A brace marks that the four non-baseline variants all leave decisions unchanged.
One real trajectory, five ways to train on it. Each row is a training variant: −log p marks tokens that carry loss and masked tokens are context only. Boxes are inserted lines; a dashed box is an auxiliary target that is never generated at test time, a solid one is part of the action. Every variant learns its target, but none changes when the agent submits an answer or keeps querying.

Key numbers

Qwen3-4B student, Qwen3-32B teacher, SkyRL text-to-SQL environment; evaluation on Spider-dev and BIRD-dev with 16 samples per question; 95% intervals from a bootstrap over questions.

+0.26
ActObs minus ActionSFT, pooled pass@16 after GRPO, two seeds each (pre-registered outcome)
[−0.38, +0.91] points
0.0049
Action KL from ActObs to ActionSFT on unseen databases
data-order noise: 0.0070
89 / 82%
Contract states the correct answer shape (Spider / BIRD)
from the question alone: 71 / 49%
≤ 0.016
Largest change in submit-or-continue sensitivity from any SFT variant (Spider)
data-order noise: 0.033

The question

A tool-using agent’s trajectory interleaves decisions with their consequences: it issues a query, the environment returns a result, and the agent decides what to do next. Standard supervised fine-tuning puts the loss on the agent’s actions only. ActObs (Zhang et al., 2026) removes this mask and also trains the agent to predict the observations. It reports that GRPO then reaches higher pass@k on Terminal-Bench, and attributes the gain to the agent learning what its actions do.

That explanation has a testable step in the middle. For a predictive objective to help, what the model learns must change what it does, and that link has not been measured. We measure it at three levels: whether the knowledge is learned, whether it moves the action distribution, and whether it changes the one decision a SQL agent makes after every result, namely to submit an answer or to keep querying.

Training variants

All variants train on the same 24,043 teacher trajectories, in the same order and with the same hyper-parameters, and then go through the same GRPO. They differ only in which tokens carry loss or in a short line inserted into the text. The main study was pre-registered, and the predictions for the later variants were written down before they were trained.

VariantWhat changesWhat it tests
ActionSFTLoss on actions only; observations are context.Baseline
ActObsLoss on actions and observations, pooled as in ActObs.Does predicting observations help RL?
ActionSFT-LRAction-only, learning rate lowered until it moves as far from the base model as ActObs.Is the effect just a smaller step?
ActionSFT+WikiThe observation loss replaced by the same number of Wikipedia tokens.Is it dilution of the action loss?
Outcome-AuxExpected result: 3 rows, 2 columns. after each query, as an auxiliary target.Predicting the outcome, outside the action
Outcome-in-ActionThe same line before each query, inside the action.Predicting the outcome, inside the action
ContractExpected answer: 1 row, 1 column. at the start, and a check against it after every result.Knowing what a correct answer looks like
Contract-PlaceboThe same lines, with contracts taken from other questions.Format versus content

Findings

RQ1Does observation supervision make RL stronger? No. Before RL all variants are within about one point of each other. After the same GRPO, ActObs and ActionSFT reach the same pooled pass@16 (+0.26 points, 95% CI [−0.38, +0.91], averaged over two RL seeds of each), and GRPO extracts no consistently larger gain from ActObs than from the action-only models.

Gain during GRPO for each SFT initialisation at pass@1 and pass@16 on Spider and BIRD, with 95% intervals; ActObs does not gain more than ActionSFT.
RQ1. Gain during RL from each initialisation: pass@k after GRPO minus pass@k before it, with 95% bootstrap intervals over questions.
Observation. Training itself is noisy. Two RL seeds of the same ActObs checkpoint end 0.97 points apart on Spider pass@1, and two checkpoints that act the same before RL (ActionSFT-LR and ActObs) end 1.48 points apart on BIRD pass@1. With variance of this size, an effect as small as the one ActObs reports for a 4B model (+1.1 pass@16) would go undetected here.

RQ2What does the model learn, and does it reach the actions? It learns the observations’ structure but not their content. On databases it never saw, ActObs predicts the output template almost perfectly and the result header far better than ActionSFT, yet predicts row values no better (1.92 vs. 1.88 nats). Its action distribution differs from ActionSFT’s by less than two action-only models trained on reshuffled data differ from each other, and its behaviour after errors and empty results is unchanged.

(a) Action KL of each SFT variant to ActionSFT against a grey data-order noise band; ActObs lies inside the band. (b) Observation cross-entropy by part on unseen and training databases: ActObs learns template and header everywhere but row values only on training databases.
RQ2. (a) Action KL of each SFT variant to ActionSFT; the grey band is the data-order noise floor. (b) Observation cross-entropy by part on unseen and on training databases.

RQ3Does it matter what the agent predicts and where? For the action tokens, yes. The same outcome line moves the action distribution only when it is written inside the action (2.3 times the noise floor); as an auxiliary target it does not. The contract is learned on unseen databases and states the correct answer shape in 89% (Spider) and 82% (BIRD) of rollouts, against 71% and 49% for a classifier that reads only the question.

Observation. The in-action predictions are blind to failure. Outcome-in-Action writes “error” before only 9 of the 8,311 calls that do fail. Its labels come from teacher trajectories in which only 2.8% of calls fail.

RQ4Does any of it change the agent’s decisions? No. After every result the agent either submits or keeps querying. We measure how much more often it submits when the result has the right shape than when it does not. ActionSFT already shows this asymmetry, inherited from the teacher, and no variant changes it by more than retraining ActionSFT on a reshuffled data order does. This holds even for Contract, whose own check reports every mismatch in words.

Forest plot of the change in decision sensitivity relative to ActionSFT for every variant after SFT and after GRPO, with 95% intervals; all lie within the grey noise references.
RQ4. Change in decision sensitivity, P(submit | right shape) minus P(submit | wrong shape), relative to ActionSFT, with 95% intervals. Grey rows are noise references: another data order after SFT, another RL seed after GRPO.
Observation. What RL does change is shared by every variant: it makes the agent submit more readily after a mismatching result. Measured with Contract’s own checks on BIRD, submitting after a reported mismatch rises from 31% to 46%.

RQ5Can the knowledge be used from outside? Before RL, yes. Picking among 16 sampled answers by the contract beats a filter that drops errors and empty results, and adds to execution majority voting. After GRPO, it adds nothing: RL has already removed most of the shape errors a contract would catch.

Choosing one of 16 ActionSFT samplesBefore RLAfter GRPO
SpiderBIRDSpiderBIRD
Random sample (= pass@1)71.449.780.558.1
No error, non-empty result72.553.581.160.5
Execution majority vote75.858.382.663.9
Matches the contract76.956.581.360.6
Matches the contract, then vote78.459.782.563.3
Oracle: gold result shape78.259.582.663.4

Accuracy (%) on the primary questions. The contracts come from the Contract model and are applied to ActionSFT’s samples, so the selector uses knowledge that ActionSFT itself never had.

Takeaway. Learning about the environment, or even about the answer, is not enough for an agent to act on it. Decisions follow imitation and the outcome reward; making knowledge reach them probably requires rewarding its use, or training on the agent’s own, failure-rich interaction.

What the agent writes

Three real rollouts on Spider-dev. Inserted lines are in purple, queries in blue and observations on grey; […] and ... mark text we shortened.

Question 728 · “How many people live in Gelderland district?”

Expected answer: 1 row, 1 column.
<sql>SELECT Population FROM city WHERE District = 'Gelderland';</sql>
152463 / 101574 / 138020 / 153491
Check: got several rows, 1 column; this does not match the expected answer. […] Since the question asks for the total number of people living in the district, we need to sum these values.
<sql>SELECT SUM(Population) FROM city WHERE District = 'Gelderland';</sql>
545548
Check: got 1 row, 1 column; this matches the expected answer.
<solution>SELECT SUM(Population) FROM city WHERE District = 'Gelderland';</solution>

The contract is correct, the check catches the four-row result, and the agent revises to a sum. This is the behaviour the contract was meant to produce; the overview figure uses the same question.

Question 361 · “Show paragraph details for paragraph with text ‘Korea’.”

Expected answer: 1 row, 1 column.
<sql>SELECT * FROM Paragraphs WHERE Paragraph_Text = 'Korea';</sql>
510442723 2394 Korea None / 7 2394 Korea None
Check: got several rows, 4 columns; this does not match the expected answer.
<solution>SELECT * FROM Paragraphs WHERE Paragraph_Text = 'Korea';</solution>

After GRPO the contract is still correct and the check still reports the mismatch, yet the agent submits the query it has just rejected, and the answer is wrong.

Question 660 · “Return the names of all the poker players.”

Expected result: 10 rows, 1 column.
<sql>SELECT p.Name FROM poker_player p JOIN people p2 ON ...;</sql>
Error executing SQL: no such column: p.Name
Expected result: 10 rows, 1 column.
<sql>SELECT p2.Name FROM poker_player p JOIN people p2 ON ...;</sql>
Aleksey Ostapenko / Yevgeni Sivozhelez / ... (5 rows)
<solution>SELECT p2.Name FROM poker_player p JOIN people p2 ON ...;</solution>

The prediction written before each call ignores what the call will do: the first call fails and the second returns 5 rows, yet both predictions say 10 rows.

Environment audit

The public SkyRL text-to-SQL environment rewards a solution when its result set equals the gold query’s. Running the unmodified environment with the constant answer SELECT 1 WHERE 0; earns full reward on 290 of its 653 training tasks (44%): 283 of the 540 SynSQL tasks and 7 of the 113 Spider tasks. Their gold queries return no rows, and for 18 of them the result depends on the system clock through filters such as year > strftime('%Y','now') - 5. We train RL only on the 363 tasks with non-empty gold results, and reported the problem to the maintainers with a reproduction script (SkyRL issue #2451).

BibTeX

@misc{lu2026know2act,
  title  = {Know2Act: What Knowledge Reaches a Tool-Using Agent's Actions?},
  author = {Lu, Yiwen},
  year   = {2026},
  note   = {Manuscript}
}