◈ KROMALOCA ACADEMY · MODULE 39 (MASTERCLASS AUTONOMOUS SYSTEMS)
MASTERCLASS · TOPIC 39⏱️ 8 MIN READ⚡ 10-QUESTION SCENARIO CHALLENGE

Evaluation & Self-Correction Loops (AI Critiquing AI)

Eliminating sycophancy and hallucinations: building automated LLM-as-a-Judge rubrics and self-refinement harnesses.

← View Academy Curriculum HubCurriculum Track: Masterclass Autonomous Systems

The Grading Dilemma: Who Watches the Watchers?

In traditional machine learning, evaluating text output was notoriously difficult. Metrics like BLEU and ROUGE simply counted overlapping n-grams. If a candidate wrote "The weather is delightful" and the reference said "It is a lovely sunny day", BLEU scored it near zero despite having identical semantic meaning.

Today, the gold standard for automated quality assurance in complex pipelines is LLM-as-a-Judge, popularized by LMSYS and Zheng et al. (2023). However, if implemented naively, LLM evaluators suffer from severe biases: self-enhancement bias, position bias, verbosity bias, and sycophancy (flattering the input rather than offering rigorous criticism).

The Anatomy of a Calibrated Rubric

Never tell a judge model: "Grade this answer from 1 to 5." A model will give almost everything a 4 or 5. Rigorous evaluation requires an explicit, multi-dimensional rubric with anchor descriptions:

Example 5-Point Accuracy Rubric:
Score 1 (Critical Failure): Contains factually dangerous or completely fabricated claims. Misidentifies core entities.
Score 2 (Deficient): Identifies the correct topic, but makes major factual errors or misses primary user constraints.
Score 3 (Acceptable): Factually sound, but includes unnecessary filler, minor ambiguity, or weak formatting.
Score 4 (Strong): Completely accurate, concise, directly answers the prompt, adheres to all constraints.
Score 5 (Exemplary): Flawless precision, exceptional clarity, anticipates unstated edge cases, zero filler.

The Self-Correction Refinement Loop

When generating high-stakes content (legal briefs, software patches, financial summaries), single-shot generation is rarely sufficient. A closed-loop Critic-Refine cycle elevates output quality significantly:

  1. Generation: Agent A drafts the initial response based on the user prompt.
  2. Critique (Blind Evaluation): Agent B evaluates the draft against the rubric without knowing which model wrote it. It must output an explicit list of flaws:
    CRITIQUE:
    1. Violated constraint: Exceeded 100-word limit by 24 words.
    2. Missing entity: Did not mention the GDPR Article 17 requirement.
    3. Tone: Used casual colloquialism in paragraph 2 ("a bunch of stuff").
  3. Refinement: Agent A receives the critique and rewrites the draft, addressing every single numbered item.
  4. Pass/Fail Gate: If score >= 4, the response is delivered to the user; otherwise, the loop allows one final refinement pass.

Combating Evaluator Biases

  • Position Bias: When comparing Option A vs Option B, models favor Option A up to 60% of the time. Always swap positions and average the scores.
  • Verbosity Bias: Models naturally correlate longer answers with higher quality. Explicitly penalize verbosity in your scoring instructions.
  • Require Chain-of-Thought Grading: Never allow the judge model to emit a score first. Force it to write its detailed reasoning before outputting the final numerical rating.

TEST YOUR PROMPTING INSTINCTS

Topic 39 Scenario Challenge.

10 real-world scenarios designed to test how you apply the techniques from this lesson.

🎯
PASSING REQUIREMENT: 60% (6 OF 10 SCENARIOS)

You must achieve a minimum score of 60% on this challenge to unlock Lesson 40. Answers and technical rationales remain locked until all 10 scenarios are submitted.