Most agents ship with a hand-written prompt, and most prompts get improved the same way: someone tries a few conversations, spots a problem, edits the instructions, and hopes nothing else broke. It works for one agent. It does not work for ten. Microsoft Foundry's agent optimizer automates that loop. It scores your agent against a test set, generates alternative configurations, scores those too, and tells you which one is actually better.
Microsoft announced the optimizer as generally available in October 2026, so check the docs for current status in your region before you rely on it.
What it can improve
The optimizer works on two kinds of Foundry agent. What it can change depends on which one you have.
- Prompt agents: instructions, descriptions of function-calling tools, and the choice of model.
- Hosted agents: all of the above, plus skills.
It never retrains a model and never changes your agent's code. It only adjusts the configuration your agent loads.
How a run works
- The optimizer runs your current agent against a dataset of tasks and scores each answer against your criteria. That score is the baseline.
- It generates candidates, such as rewritten instructions or better tool descriptions.
- It scores every candidate on the same tasks.
- It ranks the results by a composite score between 0.0 and 1.0 and marks the best one. You decide whether to apply it.
The composite score is the average of all evaluator scores across all tasks, rescaled to a 0 to 1 range.
Path 1: optimize a prompt agent in the portal
This path needs no code. Open your project in the Foundry portal, go to Agents, pick the agent, and open the Optimize tab. A wizard asks for the agent version, a dataset, the evaluation criteria, and the models to use. You can generate the dataset from the agent's own traces, pick an existing one, or upload one. When the run finishes you see baseline and candidate scores, a before-and-after view of the prompt, and per-evaluator results. Promoting a candidate creates a new version of the agent.
Path 2: optimize a hosted agent from the command line
For a hosted agent, the optimizer needs to find your instructions and other settings in files it can swap at runtime. Your agent has to be made optimizer-ready: a baseline configuration with an instructions.md file and, if you want those targets, a skills/ folder with SKILL.md files and a tools.json file for function-calling tools. Your code loads its settings through load_config(). That call returns your baseline normally and supplies the candidate configuration during an optimizer run, with no flags in your code. Foundry Toolkit for VS Code can scaffold all of this for you.
Two models take part in a run, and both must be deployed in your project. The eval model scores answers, and any chat model works for it. The optimization model writes the candidates, and it is required for hosted agents. A more capable model here usually produces better candidates. The docs list the GPT-5 family and DeepSeek models as supported.
The whole workflow with the Azure Developer CLI looks like this:
azd ai agent init # scaffold the agent project
azd deploy # ship it to Foundry
azd ai agent eval init # generate a dataset and criteria from the agent's instructions
azd ai agent eval run # score the agent
azd ai agent optimize # generate and score candidates
azd ai agent optimize apply --candidate <candidate-id>
azd deploy # deploy the optimized agent as a new version
The eval init step solves the cold-start problem. If you have no test set yet, it builds one from your agent's existing instructions. Later you can replace it with a dataset built from real production traces. The CLI prints a results table with each candidate, its score, and what changed, and it marks the best candidate with a star. The apply command writes that candidate into your local configuration so you can review the diff before you deploy it as a new version. Some of these CLI commands appeared in preview earlier this year, so confirm the exact flags with azd ai agent optimize --help.
How to read the score
Microsoft gives a simple guide for how much a change in the composite score matters:
- Below 0.03 is noise.
- 0.03 to 0.10 is a moderate improvement and worth deploying.
- 0.10 to 0.20 is significant.
- Above 0.20 is major, and usually means the baseline was weak.
Things to watch
- The optimizer calls your agent on every task in the dataset. If your agent uses real tools, those calls really run. Point it at test endpoints or mocked tools so evaluation does not create orders, send emails, or hit rate limits.
- Better instructions are often longer, so responses can use more tokens. Compare the cost increase against the score gain before you apply a candidate.
- The score is only as good as your dataset and criteria. A dataset that does not look like real traffic produces an agent that is tuned for the wrong thing.
- Hosted-agent optimization needs the Responses protocol, and it is available in the hosted-agent regions except Norway East.
- Read the diff. The optimizer proposes changes, and a person should approve them.
Where this fits
The optimizer is one step in a loop that Foundry now supports end to end: watch production traces, find recurring problems, build evaluation data from real traffic, optimize, validate, and repeat. Start small. Take one agent, generate a dataset, record the baseline score, and run the optimizer once. Even a moderate gain on a single agent shows whether the loop is worth building into your release process.
Resources
- Agent optimizer overview: learn.microsoft.com
- What's new in hosted agents: devblogs.microsoft.com/foundry
- Foundry September 2026 update: azure.microsoft.com
Join the conversation! Your thoughts help the community grow.