Know2Act: What Knowledge Reaches
a Tool-Using Agent’s Actions?
University of Notre Dame · ylu37@nd.edu
Manuscript · October 2026
Abstract
Recent work trains language agents to predict not only their own actions but also the observations that the environment returns, and reports that this makes later reinforcement learning (RL) more effective because the agent learns the consequences of its actions. We ask whether knowledge learned in this way reaches the agent’s decisions. In Know2Act, a pre-registered study of a multi-turn text-to-SQL agent trained with supervised fine-tuning and then GRPO, we measure what each model learns and how it acts. Observation supervision is learned: on unseen databases the agent predicts the format and headers of result tables far better. Yet it changes the agent’s actions less than reshuffling the training data does, and it brings no gain after RL. We then vary what the agent predicts and where, up to a contract that states the shape of the correct answer and is checked against every query result. Every prediction is learned, but none changes when the agent submits an answer or keeps querying. The knowledge is usable from outside: selecting among sampled answers with the agent’s own contract improves accuracy. RL, which rewards only the final answer, removes even this gain and makes the agent more willing to submit after its own check has failed. Learning to predict the environment is therefore not the same as acting on the prediction. We also find that 44% of the training tasks in a public SQL RL environment reward an empty answer.
−log p
marks tokens that carry loss and masked tokens are context only. Boxes are inserted lines; a dashed box is an
auxiliary target that is never generated at test time, a solid one is part of the action. Every variant learns its
target, but none changes when the agent submits an answer or keeps querying.Key numbers
Qwen3-4B student, Qwen3-32B teacher, SkyRL text-to-SQL environment; evaluation on Spider-dev and BIRD-dev with 16 samples per question; 95% intervals from a bootstrap over questions.
The question
A tool-using agent’s trajectory interleaves decisions with their consequences: it issues a query, the environment returns a result, and the agent decides what to do next. Standard supervised fine-tuning puts the loss on the agent’s actions only. ActObs (Zhang et al., 2026) removes this mask and also trains the agent to predict the observations. It reports that GRPO then reaches higher pass@k on Terminal-Bench, and attributes the gain to the agent learning what its actions do.
That explanation has a testable step in the middle. For a predictive objective to help, what the model learns must change what it does, and that link has not been measured. We measure it at three levels: whether the knowledge is learned, whether it moves the action distribution, and whether it changes the one decision a SQL agent makes after every result, namely to submit an answer or to keep querying.
Training variants
All variants train on the same 24,043 teacher trajectories, in the same order and with the same hyper-parameters, and then go through the same GRPO. They differ only in which tokens carry loss or in a short line inserted into the text. The main study was pre-registered, and the predictions for the later variants were written down before they were trained.
| Variant | What changes | What it tests |
|---|---|---|
| ActionSFT | Loss on actions only; observations are context. | Baseline |
| ActObs | Loss on actions and observations, pooled as in ActObs. | Does predicting observations help RL? |
| ActionSFT-LR | Action-only, learning rate lowered until it moves as far from the base model as ActObs. | Is the effect just a smaller step? |
| ActionSFT+Wiki | The observation loss replaced by the same number of Wikipedia tokens. | Is it dilution of the action loss? |
| Outcome-Aux | Expected result: 3 rows, 2 columns. after each query, as an auxiliary target. | Predicting the outcome, outside the action |
| Outcome-in-Action | The same line before each query, inside the action. | Predicting the outcome, inside the action |
| Contract | Expected answer: 1 row, 1 column. at the start, and a check against it after every result. | Knowing what a correct answer looks like |
| Contract-Placebo | The same lines, with contracts taken from other questions. | Format versus content |
Findings
RQ1Does observation supervision make RL stronger? No. Before RL all variants are within about one point of each other. After the same GRPO, ActObs and ActionSFT reach the same pooled pass@16 (+0.26 points, 95% CI [−0.38, +0.91], averaged over two RL seeds of each), and GRPO extracts no consistently larger gain from ActObs than from the action-only models.
RQ2What does the model learn, and does it reach the actions? It learns the observations’ structure but not their content. On databases it never saw, ActObs predicts the output template almost perfectly and the result header far better than ActionSFT, yet predicts row values no better (1.92 vs. 1.88 nats). Its action distribution differs from ActionSFT’s by less than two action-only models trained on reshuffled data differ from each other, and its behaviour after errors and empty results is unchanged.
RQ3Does it matter what the agent predicts and where? For the action tokens, yes. The same outcome line moves the action distribution only when it is written inside the action (2.3 times the noise floor); as an auxiliary target it does not. The contract is learned on unseen databases and states the correct answer shape in 89% (Spider) and 82% (BIRD) of rollouts, against 71% and 49% for a classifier that reads only the question.
RQ4Does any of it change the agent’s decisions? No. After every result the agent either submits or keeps querying. We measure how much more often it submits when the result has the right shape than when it does not. ActionSFT already shows this asymmetry, inherited from the teacher, and no variant changes it by more than retraining ActionSFT on a reshuffled data order does. This holds even for Contract, whose own check reports every mismatch in words.
RQ5Can the knowledge be used from outside? Before RL, yes. Picking among 16 sampled answers by the contract beats a filter that drops errors and empty results, and adds to execution majority voting. After GRPO, it adds nothing: RL has already removed most of the shape errors a contract would catch.
| Choosing one of 16 ActionSFT samples | Before RL | After GRPO | ||
|---|---|---|---|---|
| Spider | BIRD | Spider | BIRD | |
| Random sample (= pass@1) | 71.4 | 49.7 | 80.5 | 58.1 |
| No error, non-empty result | 72.5 | 53.5 | 81.1 | 60.5 |
| Execution majority vote | 75.8 | 58.3 | 82.6 | 63.9 |
| Matches the contract | 76.9 | 56.5 | 81.3 | 60.6 |
| Matches the contract, then vote | 78.4 | 59.7 | 82.5 | 63.3 |
| Oracle: gold result shape | 78.2 | 59.5 | 82.6 | 63.4 |
Accuracy (%) on the primary questions. The contracts come from the Contract model and are applied to ActionSFT’s samples, so the selector uses knowledge that ActionSFT itself never had.
What the agent writes
Three real rollouts on Spider-dev. Inserted lines are in purple, queries in blue and observations
on grey; […] and ... mark text we shortened.
Question 728 · “How many people live in Gelderland district?”
Expected answer: 1 row, 1 column. <sql>SELECT Population FROM city WHERE District = 'Gelderland';</sql> 152463 / 101574 / 138020 / 153491 Check: got several rows, 1 column; this does not match the expected answer. […] Since the question asks for the total number of people living in the district, we need to sum these values. <sql>SELECT SUM(Population) FROM city WHERE District = 'Gelderland';</sql> 545548 Check: got 1 row, 1 column; this matches the expected answer. <solution>SELECT SUM(Population) FROM city WHERE District = 'Gelderland';</solution>
The contract is correct, the check catches the four-row result, and the agent revises to a sum. This is the behaviour the contract was meant to produce; the overview figure uses the same question.
Question 361 · “Show paragraph details for paragraph with text ‘Korea’.”
Expected answer: 1 row, 1 column. <sql>SELECT * FROM Paragraphs WHERE Paragraph_Text = 'Korea';</sql> 510442723 2394 Korea None / 7 2394 Korea None Check: got several rows, 4 columns; this does not match the expected answer. <solution>SELECT * FROM Paragraphs WHERE Paragraph_Text = 'Korea';</solution>
After GRPO the contract is still correct and the check still reports the mismatch, yet the agent submits the query it has just rejected, and the answer is wrong.
Question 660 · “Return the names of all the poker players.”
Expected result: 10 rows, 1 column. <sql>SELECT p.Name FROM poker_player p JOIN people p2 ON ...;</sql> Error executing SQL: no such column: p.Name Expected result: 10 rows, 1 column. <sql>SELECT p2.Name FROM poker_player p JOIN people p2 ON ...;</sql> Aleksey Ostapenko / Yevgeni Sivozhelez / ... (5 rows) <solution>SELECT p2.Name FROM poker_player p JOIN people p2 ON ...;</solution>
The prediction written before each call ignores what the call will do: the first call fails and the second returns 5 rows, yet both predictions say 10 rows.
Environment audit
The public SkyRL text-to-SQL environment rewards a solution when its result set equals the gold query’s.
Running the unmodified environment with the constant answer SELECT 1 WHERE 0; earns full reward on
290 of its 653 training tasks (44%): 283 of the 540 SynSQL tasks and 7 of the 113 Spider tasks. Their gold queries
return no rows, and for 18 of them the result depends on the system clock through filters such as
year > strftime('%Y','now') - 5. We train RL only on the 363 tasks with non-empty gold results, and
reported the problem to the maintainers with a reproduction script
(SkyRL issue #2451).
BibTeX
@misc{lu2026know2act,
title = {Know2Act: What Knowledge Reaches a Tool-Using Agent's Actions?},
author = {Lu, Yiwen},
year = {2026},
note = {Manuscript}
}



