Predicting command outputs helps with later agent training
An agent working in a terminal tries to run a program. It sends a command, reads an error message and chooses its next action. In doing so, it must connect the output to the earlier command: recognize whether a file, memory or the right argument was missing. This determines whether the next attempt will improve the situation.
Such an agent can be trained on records of a more capable model's work. The record contains both commands and terminal responses, but training usually evaluates only the prediction of text written by the agent. Tool responses are available as context for subsequent decisions. The model does not, however, have to predict for itself what response will follow a given command.
After imitating demonstrations, the agent can continue learning through its own attempts. Yet two models that reproduce successful actions equally well may respond differently to new situations. An evaluation performed immediately after learning from demonstrations may therefore fail to show which model will be a better starting point for further training.
This relationship is examined in Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL. This preprint, dated 17 September 2026, concerns training agents based on language models. Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan, Vinayshekhar Bannihatti Kumar and Rashmi Gangadharaiah propose ActObs: during learning from demonstrations, the model also predicts environment responses, referred to below as observations. The subsequent algorithm for learning through rewards remains the same.
A terminal response can be context or a training target
Learning from existing demonstrations is a form of supervised fine-tuning, or SFT: the record shows the model the desired sequence of actions, and training adjusts its parameters. Text is divided into tokens, the units processed by the model. At each evaluated position, the model must predict the next token based on the preceding text.
Loss masking determines which positions affect the loss function, the numerical measure of error used to change parameters. In ActionSFT, the agent's tokens, including its reasoning and commands, are evaluated. Tokens from terminal responses remain in the sequence but are not prediction targets. Excluding them from the loss does not remove them from the context.
ActObs includes both groups in the loss. The model learns to predict a command and then the text of the tool response following that command. The initial task instruction is still excluded from evaluation. During actual operation, the response still comes from the tool; ActObs does not replace command execution with text invented by the model.
Prediction is evaluated using cross-entropy: the lower the probability the model assigned to the token present in the demonstration, the greater the penalty. In the basic ActObs variant, each action and observation token has the same weight. The average is calculated across all evaluated tokens, rather than across equally weighted groups of “commands” and “responses”. A longer terminal response therefore contributes more terms to the loss. Equation 1 in the paper specifies this normalization.
The same two commands provide three or five prediction targets
The following example was prepared to explain the method; it does not come from the experiment. The user asks to read a number stored in a file. A miniature tool supports two commands: list provides the filename, while read reads the specified file. The demonstration records:
agent: list
tool: report.csv
agent: read report.csv
tool: 42
For the calculation, let us assume a hypothetical way of dividing text into tokens (a tokenizer), in which list, read, report.csv and 42 are each a single token. We omit role markers. The sequence has five positions: three belong to the agent and two to the tool. report.csv appears twice, occupying a separate position each time.
ActionSFT evaluates the prediction of list, followed by read and the argument report.csv. Each of these three positions contributes 1/3 to the average loss. Both tool responses have zero weight. The model sees the first report.csv before predicting the command read, but it is not evaluated on predicting that name as the result of list.
ActObs processes exactly the same record but evaluates all five positions, each contributing 1/5. The first difference occurs at the response to list: the model now receives a training signal for predicting report.csv. The second occurs at 42, which it must predict after read report.csv. In both cases, it sees only the preceding part of the record, and during training each successive position receives the actual preceding text from the demonstration.
The commands correspond to the set of action tokens in the algorithm, and the tool outputs to the set of observation tokens. Their contribution to the loss changes; no part of the demonstration is resampled or removed. The signal at the output of read requires matching the prediction to the effect of the preceding command. This does not mean that the model can know arbitrary contents of an unread file: unpredictable information can still cause errors.
Five tokens and two calls are only a teaching simplification. In the study, sequences contain reasoning, full commands and long terminal responses. ActObs does not add separate demonstrations or model passes, because tool responses were already being processed as context. The entire training process still includes collecting demonstrations and the agent's subsequent execution of tasks; its budget is given below.
The advantage emerged after the same GRPO training
After SFT, the authors applied GRPO, a reinforcement learning (RL) algorithm that compares rewards within a group of attempts at the same task. Here, an attempt covers the entire interaction with the terminal, or rollout. An automatic verifier awards a reward for completing the task. At this stage, both models have the same algorithm, data and budget, and ActObs no longer adds a loss for predicting environment responses.
pass@k describes the chance of obtaining at least one correct execution among a specified number of attempts; a higher score is better. pass@1 concerns a single attempt, although its success rate can be estimated from many runs.
The clearest transfer result concerns Qwen3-4B on aider-polyglot, a set of code editing tasks in several programming languages, checked by unit tests. After the same GRPO schedule, the final model starting from ActionSFT achieved 9.7% pass@1, while the model starting from ActObs achieved 13.9%. The difference was 4.2 percentage points. The tasks in this test appeared neither in the SFT demonstrations nor in GRPO training. Before GRPO, ActionSFT had the higher score on this test. pass@1 was estimated from four attempts per task; the best of four responses was not selected.
A separate experiment with the larger Qwen3-8B on terminal tasks revealed a trade-off: ActObs reduced the single-attempt score but improved the chance of success over multiple attempts. It was therefore not a uniformly better initialization at every budget. This is an important distinction between a system that must perform a task once and a system that can retry and check the results.
Experiment details
Source of the configuration and results: §3.1, Table 1 and Appendices A–B of preprint v1. The main comparisons concern the final checkpoints, meaning saved parameter versions after the same training schedule. The authors describe them as final policies and compare the state after training has ended; the figures below are not maxima read from a training curve.
SFT. Qwen3-4B and Qwen3-8B were trained on 50 thousand multi-turn records of terminal work from the synthetic portion of Nemotron-Terminal-Corpus. DeepSeek-V3.2 generated the demonstrations. The corpus contained 0.71 billion tokens, approximately 45% of which were observations. Training lasted one epoch, 781 steps, with 64 examples per batch (batch size) and a maximum parameter update step coefficient (learning rate) of 0.00001, decreased according to a cosine schedule. Basic ActObs used an observation weight of 1; ActionSFT used 0. Normalization covered the number of action tokens and the weighted number of observation tokens.
GRPO. This was followed by 135 steps on 2392 Endless Terminals tasks, disjoint from the SFT corpus. The setup used 16 rollouts per task, batch size 32, learning rate 0.000001 and binary verifiers. The evaluation sets were disjoint from both training stages. These rollouts should not be confused with the two calls to the miniature tool in the example.
Terminal evaluation. Terminal-Bench 2.0 contained 89 tasks. Each main checkpoint was evaluated in 16 attempts per task, collected in four separately launched batches of four attempts. The setup used the terminus-2 agent, official time limits and the same model serving configuration: one vLLM engine on eight A100 GPUs, 12 parallel attempts, a context of 40 960 tokens and GPU memory utilization set to 0.90. Temperature, a parameter affecting randomness in token selection, was 0.6. Top-p restricted the sampling pool to tokens with a cumulative probability of at least 0.95, and top-k to the 20 most probable tokens. Model errors received zero; infrastructure errors were retried once. The model serving configuration was selected after comparing nine settings and then used for all methods.
Code editing evaluation. Aider-polyglot comprised 225 tasks in six programming languages, with four attempts per task. The model serving configuration mentioned above applies to Terminal-Bench 2.0; it should not automatically be attributed to the aider test.
Each row below compares ActionSFT→GRPO with ActObs→GRPO under the same setting. The table covers the main results for standard GRPO, without the ECHO variant that also adds observation prediction during RL.
| Model and test | Metric | ActionSFT→GRPO | ActObs→GRPO |
|---|---|---|---|
| Qwen3-4B, aider-polyglot, 4 attempts per task | pass@1 | 9.7 ± 0.6% | 13.9 ± 0.7% |
| Qwen3-4B, aider-polyglot, 4 attempts per task | pass@4 | 20.4 ± 0.9% | 25.3 ± 1.0% |
| Qwen3-4B, Terminal-Bench 2.0, 16 attempts per task | pass@1 | 5.6 ± 0.4% | 7.2 ± 0.4% |
| Qwen3-4B, Terminal-Bench 2.0, 16 attempts per task | pass@16 | 18.0 ± 1.4% | 19.1 ± 1.4% |
| Qwen3-8B, Terminal-Bench 2.0, 16 attempts per task | pass@1 | 12.3 ± 0.5% | 11.0 ± 0.5% |
| Qwen3-8B, Terminal-Bench 2.0, 16 attempts per task | pass@16 | 23.6 ± 1.2% | 27.0 ± 1.3% |
For Qwen3-4B on aider-polyglot before RL, pass@1 was 1.4 ± 0.3% after ActionSFT and 1.0 ± 0.3% after ActObs. The later ActObs advantage was therefore not simply a matter of retaining a better initial score.
Values after ± indicate one standard error estimated through bootstrapping: 20 thousand resamplings of attempts with the task set unchanged. They measure uncertainty arising from the limited number of agent runs. They do not measure variability between independent training runs or between different task sets. pass@k was calculated using an estimator that accounts for the number of successes in the available attempts.
Controls and diagnostics. The Obs→Act variant first trained only observation prediction, then only actions; it required two epochs and did not reproduce the advantage of joint training at larger attempt budgets. Shuffling responses between demonstrations reduced the scores after SFT. These are separate controls, not additional main GRPO results. Gradient analysis used 256 held-out demonstrations, and observation prediction evaluation used 300 records from the validation portion of the corpus. Measurements of entropy, a measure of the spread of next-token probabilities, on the final models covered 200 evaluation records each. In the entropy-matching control, the temperature of ActionSFT→GRPO was raised from 0.6 to 0.64; this did not eliminate the pass@16 difference on Qwen3-8B.
ActObs preserves prediction of effects and more command variants
The authors checked how the models predicted terminal responses in held-out demonstrations after SFT. ActionSFT worsened this ability relative to the model before SFT, whereas ActObs preserved and improved it. The effect covered the actual content of outputs, including errors and shell messages, not merely the repetitive header of a tool response.
Gradient diagnostics, meaning the directions of parameter changes arising from the loss, help explain why training on commands alone did not ensure this. During SFT, the direction that improved action prediction quickly stopped aligning with the direction that improved observation prediction. ActionSFT omitted the second signal. ActObs included it throughout training, although direct task execution scores after SFT remained similar.
After GRPO, ActObs retained higher entropy, meaning a greater spread of next-token probabilities. The difference was particularly visible in later command positions, where arguments, options and paths occur. More such variants remained available when sampling the next attempt. This does not mean that the agent discovered more entirely different strategies: analysis of successful executions primarily indicated differences in how similar procedures were carried out.
These measurements support an explanation involving the preservation of useful alternatives, but they do not prove that entropy alone accounts for the improvement. Increasing the randomness of the ActionSFT model to a similar level did not reproduce the ActObs result. The controlled change to the mask shows the influence of the SFT objective on later results; the complete chain from observation prediction to a particular successful action remains an interpretation supported by diagnostics.
Evaluate initialization after further training and at the intended attempt budget
The study concerns one model family in two sizes, demonstrations from one model and training in a terminal. Transfer to code editing broadens the scope of the test, but does not establish that the method works in a browser, graphical interface or robotics. The authors also disclose differences between the environments at different stages: during RL, commands were executed in a fresh, non-interactive shell, whereas the terminal test used a persistent tmux session. Comparisons between methods within each stage were controlled, but such differences may change absolute scores. The authors announce that code and detailed artifacts will be released after the paper is accepted.
SFT evaluation should account for the effect of subsequent RL, because a similar score after demonstrations does not guarantee a similar outcome from further training. A test of ActObs on your own setup should keep the data and budget the same and then compare the models after the full planned training. Because the larger model in the study gained over multiple attempts at the expense of a single attempt, both scores and the cost of retries need to be measured; pass@k does not solve the problem of recognizing which attempt ended correctly.
Source: Juzheng Zhang and coauthors, Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL, arXiv:2609.20715v1, 2026. Polish overview based on the full text released under the CC BY 4.0 license.
I use AI-generated content as part of my daily learning process.