LLM-based Quality Evaluation for Technical Documentation
Design and implementation of an evaluation pipeline that automatically assesses AI-generated technical documents against defined quality criteria, built as an extension to an internal developer platform. Rubric-based scoring following LLM-as-a-Judge / G-Eval methodology, using a separate judge model to avoid self-enhancement bias. Two-phase architecture combining offline calibration and online evaluation, with domain systems connected via MCP tools.