# SageMaker multi-turn RL: what a search team must verify before training
> AWS reports gains on three of four held-out search benchmarks for its Qwen3.6-27B agent. This guide maps the SageMaker setup, costs and evaluation needed before a trial.
By BIG CHANGE Editorial
Published: 2026-10-04T17:55:25.736Z
Updated: 2026-10-04T17:55:25.736Z
Canonical: https://bigchange.ai/blog/fine-tune-search-agent-sagemaker-multi-turn-rl

BIG CHANGE graphic; data: AWS Machine Learning Blog
## The big change
In its [October 2 walkthrough](https://aws.amazon.com/blogs/machine-learning/fine-tune-a-search-agent-with-multi-turn-rl-on-amazon-sagemaker-ai/), AWS describes fine-tuning a Qwen3.6-27B search agent with multi-turn reinforcement learning (MTRL) in SageMaker AI. The chart above plots AWS's reported held-out nDCG@10 scores: they rose on WixQA, Wands and BrowseComp-Plus, and fell slightly on FreshStack. BIG CHANGE made the chart from AWS's figures; we did not independently test the agent or reproduce the results. The technique trains a model against the outcome of a sequence of tool-using decisions.
Before submitting a job, an ML engineer needs to establish that the model and region are supported, the agent can participate in SageMaker’s rollout loop, the prompts and rewards represent the intended search task, and a separate evaluation set can expose regressions. This guide is based on AWS documentation; BIG CHANGE did not run an AWS job or call its APIs.
## What MTRL changes in the training loop
In supervised fine-tuning, a team needs example trajectories that show the desired behavior. In AWS’s description of MTRL, SageMaker instead sends prompts to an agent, the agent calls the policy model and its tools over multiple turns, and the agent reports a reward for the completed rollout. The model is updated using that feedback. For search, a reward can score the final ranking, so early query choices are optimized in the context of the whole search episode.
The distinction matters when tool calls depend on earlier results. If a model makes a poor first query, sees weak results and should reformulate, an end-of-task reward can reflect whether the sequence ultimately retrieved relevant documents. It does not remove the need to define what “relevant” means or to build an agent that can call the team’s search tools.
## Check access, model and region first
The October 2 example ran in US West (Oregon), `us-west-2`, and fine-tuned Qwen3.6-27B. The [SageMaker MTRL supported-model table](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl.html), checked October 4, 2026, lists four models across six model-region pairings: Nova Lite 2.0 and GPT-OSS-20B in US East (N. Virginia), `us-east-1`, and US West (Oregon), `us-west-2`; Gemma-4-31B-it and Qwen 3.6 27B in Oregon only. Confirm the live table in the region where you intend to create the job; a model’s listing does not establish availability elsewhere.
A team also needs an AWS account and a SageMaker Studio domain if using the Studio interface, S3 access for datasets and outputs, and IAM roles configured for the newer MTRL job and runtime calls. AWS’s prerequisites document says the caller needs `CreateJob` and related job actions plus permission to pass the execution role to `job.sagemaker.amazonaws.com`. The execution role needs the `AmazonSageMakerJobFullAccess` policy and that service principal in its trust policy. The agent runtime role needs `AmazonSageMakerJobRuntimeAccess`; an AgentCore runtime also needs its own trust relationship. Studio’s runtime picker needs AgentCore list permissions. Existing broad SageMaker access alone does not include all these new job actions, according to AWS.
There are two documented agent paths. Deploy with Bedrock AgentCore for managed hosting, which AWS says works best with agents built using Strands, or host a custom agent and bridge it through a Lambda forwarder. The agent receives rollout prompts, calls the policy model through SageMaker’s job runtime, invokes its search tools, completes each rollout and reports a reward. AWS’s SDK decorator handles much of this integration; custom frameworks can call the runtime APIs directly. If the Lambda function uses a nonstandard name, AWS says its ARN may need to be added explicitly to the permission policy. A customer-managed VPC or KMS key adds further setup requirements.
## Prepare prompts and reward for the real search task
SageMaker’s asset guide accepts Parquet, JSON Lines, JSON or CSV. It looks for a column named `prompt`, falling back to the first column if there is none, and passes that value to the agent as-is. The service does not parse, validate or transform prompt content. For tool-using search, the prompt can carry conversation messages, task metadata, a reward specification and tool configuration in the format the agent expects. AWS warns that prompts pass through without inspection, so teams remain responsible for protecting sensitive data; encryption or another suitable protection should be designed into storage and the agent path.
A reward is the objective the model will optimize. AWS’s example uses nDCG@10 on the final retrieved documents and assigns a reward of -1 when the agent hits its turn or sampling-token limit. That is a concrete starting design for a ranked retrieval task, but it makes evaluation choices consequential: teams need reliable relevance labels or another defensible scoring method, and should check that the reward does not favor shallow ranking gains while harming answer quality, latency or tool cost.
The blog post says its training mix included FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique and MLQA, with five percent of training instances in each dataset reserved as validation. Its held-out test set comprised FreshStack, WixQA, BrowseComp-Plus and Wands. Those public or synthetic datasets are not a substitute for evaluation on the private corpus and query distribution a team intends to serve.
## Submit a small, inspectable job
The [October 2 AWS blog](https://aws.amazon.com/blogs/machine-learning/fine-tune-a-search-agent-with-multi-turn-rl-on-amazon-sagemaker-ai/) uses a different SDK import and argument names in its snippet. The current [SageMaker training-job reference](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-job.html) shows the following API shape for a Bedrock AgentCore runtime. This example uses AWS’s current documented GPT-OSS-20B model identifier; choose a model available in your target region and check the live supported-model table before adapting it. The current example also requires an MLflow app ARN, role ARN and EULA acceptance.
```python
from sagemaker.train.multi_turn_rl_trainer import MultiTurnRLTrainer
trainer = MultiTurnRLTrainer(
model="openai-reasoning-gpt-oss-20b",
agent_env="arn:aws:bedrock-agentcore:us-west-2:123456789012:runtime/my-agent-runtime",
training_dataset="s3://my-bucket/prompts/prompts.parquet",
mlflow_app_arn="arn:aws:sagemaker:us-west-2:123456789012:mlflow-app/mlflow-app-id",
s3_output_path="s3://my-bucket/output/",
role="arn:aws:iam::123456789012:role/SageMakerRole",
accept_eula=True,
)
trainer.hyperparameters.max_epochs = 1
trainer.hyperparameters.global_batch_size = 32
trainer.hyperparameters.max_steps = 12
job = trainer.train(wait=True)
```
This is the current AWS documentation example, not a tested BIG CHANGE command. AWS’s October 2 blog presents a different `sagemaker.modules.train` import, `model_id`, `agent_endpoint`, `training_dataset_s3_uri` and `output_s3_uri` arguments. Because these differ from the current job guide, follow the current reference for new implementation and verify the SDK version and model requirements in your own environment; we have not established whether the blog example reflects an older SDK revision or a documentation inconsistency. The Studio UI is another documented submission path: choose a supported JumpStart model, select Multi-Turn Reinforcement Learning, set the AgentCore runtime or Lambda forwarder, provide training data and submit.
For a first run, record the model, region, dataset version, reward, limits, hyperparameters and output path so a later evaluation can be compared to the same baseline. If you use a customer-managed VPC, AWS’s separate VPC guide specifies private subnets in two Availability Zones, VPC endpoints for S3 and CloudWatch Logs, plus an AgentCore or Lambda endpoint as applicable; MLflow also needs an endpoint when used. The execution role needs ENI-management permissions. The optional customer-managed KMS path requires additional permissions and key-policy setup for the caller, execution role and runtime; AWS-owned KMS encryption at rest is the default. See the VPC and encryption guides before using either configuration. Before submission, review account quotas: AWS lists one concurrent MTRL fine-tuning job and one concurrent evaluation job by default, with both adjustable through Service Quotas. Job-management API throttling limits are not adjustable.
## Evaluate on held-out prompts before deployment
AWS’s evaluation documentation says to keep evaluation prompts out of training, use the same prompt format, cover important tools and tool combinations, include previously failing cases and protect sensitive content. SageMaker’s evaluator can run a candidate against a prompt set and report reward, pass@k and trajectory metrics. The docs show evaluating a fine-tuned model and optionally evaluating the base model in the same comparison pipeline.
Set the evaluation protocol before training. Keep a fixed, held-out set representative of the queries, permissions and document changes the live agent will face. Compare base and tuned models on the same tools, limits and scoring code. Inspect not just average retrieval score but task failures, turn counts, latency, token use and cases where relevant documents disappear from the top ranks. The documentation establishes how AWS’s evaluation job can be submitted; it does not establish that any metric alone captures production search quality.
AWS’s post reports these held-out results:
| Dataset (questions) | Base nDCG@10 | Fine-tuned nDCG@10 | Base failure rate | Fine-tuned failure rate |
| --- | --- | --- | --- | --- |
| WixQA (400) | 0.5725 | 0.6781 | 0.67% | 0.17% |
| Wands (147) | 0.5762 | 0.6112 | 0.00% | 0.00% |
| FreshStack (672) | 0.4112 | 0.4089 | 0.20% | 0.05% |
| BrowseComp-Plus (830) | 0.5136 | 0.6354 | 22.89% | 0.68% |
The table shows three nDCG@10 increases and one small decline on FreshStack. AWS attributes BrowseComp-Plus’s lower failure rate to its penalty for hitting turn or token limits. The blog does not provide confidence intervals or a statistical significance analysis alongside these figures. It also reports that the fine-tuned model used 4.5 versus 4.3 average turns on WixQA and 2.9 versus 2.2 on Wands, so fewer turns were not universal. Treat this as a vendor-reported experiment, not independent validation or evidence that another corpus will improve.
## Budget the whole trial
AWS’s MTRL documentation describes three training charge dimensions: prefill tokens processed as inputs, sampled tokens generated during rollouts, and training updates. The pricing page provides model-specific rates; those rates and a team’s rollout volume, sequence lengths, epochs and runtime determine the estimate. The example article does not publish its total bill, and the documentation does not supply one universal trial price. Check the current rate for the selected model and region, then estimate from expected job usage rather than assuming a fixed cost.
Include adjacent costs in the estimate: S3 storage and requests for training, validation, checkpoints and output artifacts; the AgentCore runtime or Lambda and any supporting search infrastructure during rollouts; logging and MLflow; held-out evaluation; and any endpoint or Bedrock deployment used after training. AWS’s post specifically advises stopping or deleting a running MTRL job, deleting unneeded S3 model artifacts and removing evaluation endpoints to avoid continuing charges. Endpoint deployment is a separate decision and cost category from training.
## Decide whether to proceed
The workflow is most relevant when the search task needs multiple dependent tool calls and the team can score the final outcome. A trial is only decision-useful if it has a supported model-region pairing, working agent integration, representative prompt data, a reward that reflects the task, a held-out comparison and a cost estimate that includes surrounding services. A strong result on AWS’s four benchmarks is a reason to inspect the method, not to assume the same effect on a different index, query mix or permission boundary.
**Sources & further reading**
- [AWS Machine Learning Blog: “Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI” (October 2, 2026)](https://aws.amazon.com/blogs/machine-learning/fine-tune-a-search-agent-with-multi-turn-rl-on-amazon-sagemaker-ai/) — the vendor how-to, setup, dataset choices and reported evaluation results.
- [Amazon SageMaker AI: Multi-turn reinforcement learning](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl.html) — current supported models, regions and billing dimensions.
- [Prerequisites](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-prereqs.html), [Preparing your agent](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-agent.html), [VPC configuration](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-vpc.html) and [encryption at rest](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-encryption-at-rest.html) — permissions, runtime choices, integration and optional network/encryption setup.
- [Creating assets](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-assets.html), [training job submission](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-job.html), [evaluation](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-evaluation.html), [hyperparameters](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-hyperparams.html) and [quotas](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-quotas.html) — current implementation references.
- [SageMaker AI pricing](https://aws.amazon.com/sagemaker/ai/pricing/) — current model customization pricing details; rates can change and vary by model and region.
*Reporting note: Documentation review only. BIG CHANGE did not access an AWS account, submit or inspect a SageMaker job, call AWS APIs or independently reproduce the reported metrics. AWS authored the implementation guide and reports its own experiment.*
## Sources
- [AWS Machine Learning Blog: Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI (October 2, 2026)](https://aws.amazon.com/blogs/machine-learning/fine-tune-a-search-agent-with-multi-turn-rl-on-amazon-sagemaker-ai/) — AWS's implementation walkthrough, dataset choices and vendor-reported held-out evaluation.
- [Amazon SageMaker AI: Multi-turn reinforcement learning](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl.html) — Current model and region combinations, workflow overview and training billing dimensions.
- [Amazon SageMaker AI MTRL prerequisites and agent setup](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-prereqs.html) — Required permissions, runtime choices, agent integration and optional network and encryption configuration.
- [Amazon SageMaker AI MTRL training and evaluation documentation](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-assets.html) — Current data inputs and workflow; the article also links the training-job and evaluation references.
- [Amazon SageMaker AI MTRL training job submission](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl-job.html) — Current Python SDK import, argument names, AgentCore example and Studio submission steps; checked October 4, 2026.
- [Amazon SageMaker AI pricing](https://aws.amazon.com/sagemaker/ai/pricing/) — Current model-specific pricing information; rates vary by model and region.
The BIG CHANGE newsletter
The big picture. At your pace.
Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.
Sent at 09:00 Belgrade time: daily, Mondays or the first of the month. Your first edition arrives at the next scheduled send after you confirm.
Your privacy, your choice.
Necessary storage supports site security and remembers your choices. Optional Google Analytics stays off until you allow it. You can read every story with necessary storage only. Privacy details