Introduction
In this blog, I’m going to introduce a powerful new capability in Microsoft Foundry called
Rubric Evaluation. This feature enables developers and AI engineers to evaluate model
outputs using structured, criteria-based scoring instead of relying only on simple pass/fail or
similarity checks. Here’s a hands-on overview of how Rubric Evaluation works and why it’s
becoming an essential part of AI development workflows.
Public preview: Recently introduced in Microsoft Foundry (2026)
General availability: Available since June 2026
Problem Statement
As AI applications grow more advanced, evaluating their outputs has become increasingly
complex. Traditional evaluation methods such as exact match checks, keyword comparison,
or even basic semantic similarity often fall short when dealing with:
•Subjective responses (e.g., explanations, summaries)
•Multi-step reasoning outputs
•Creative or open-ended answers
•Context-aware conversational responses
This creates a challenge: How do we evaluate AI responses in a consistent, structured, and
human-like way while maintaining scalability and automation?
Solution: Rubric Evaluation in Microsoft Foundry
Microsoft Foundry introduces Rubric Evaluation, a structured evaluation approach where AI
outputs are assessed against predefined criteria (rubrics) rather than simple correctness
checks.
A rubric defines what “good” looks like by breaking evaluation into multiple dimensions
such as:
•Accuracy
•Completeness
•Clarity
•Relevance
•Tone or style
•Reasoning quality
Instead of just saying “correct” or “incorrect,” the model is scored across these dimensions,
providing richer insights into performance.
Key Benefits of Rubric Evaluation
1. Structured evaluation: Breaks down responses into measurable criteria instead of binary
outcomes.
2. Human-like assessment: Mimics how humans evaluate answers .
3. Better debugging: Helps identify why a response failed, not just that it failed.
4. Flexible and customizable: You can define rubrics tailored to your use case (chatbots,
summarization, Q&A, etc.).
5. Works well with LLM-based evaluation: Rubric scoring can be performed using LLM
evaluators for scalable automation.
How Rubric Evaluation Works
Step 1: Create an Evaluator
1. Go to the Microsoft foundry portal and then click on the evaluations section on the left side
panel and then click on the Create evaluator button as shown in the below picture.
2. Once Create evaluator window opens, fill in the required details, including the evaluator
name, display name, and description. Then, set the evaluator type to “Rubric”.
3. Next, select a suitable model from the available options. Choose the model that best
matches your evaluation requirements and is supported for rubric-based evaluation.
4. Select the agent you want to use for the evaluation. In this example, I have two agents
available, so I will choose the Leave Policy Agent.
5. After configuring the required settings, click Generate rubric to start creating the rubric
based on the selected evaluator setup. javascript:”javascript:”javascript:”
writer.editor.GO_TO_TOP
writer.editor.GO_TO_BOTTOM
Thank you, Swarangi for the valuable insights for this blog!
