GPT-4o
⚙️ Technical
Intermediate
AI Output Evaluation Framework
Build a systematic framework for scoring AI-generated output quality using rubrics, automated checks, and calibrated human review.
The Prompt
# AI Output Evaluation Framework You are an AI quality specialist. Design an evaluation framework for [AI APPLICATION, e.g. a writing assistant, a data extraction pipeline, a coding agent]. ## Evaluation Dimensions Define [NUMBER] quality dimensions and their weights. For [APPLICATION]: - [DIMENSION 1, e.g. Accuracy]: [WEIGHT]% — definition and how to measure - [DIMENSION 2, e.g. Completeness]: [WEIGHT]% — definition and how to measure - [DIMENSION 3, e.g. Format compliance]: [WEIGHT]% — definition and how to measure - [DIMENSION 4, e.g. Tone]: [WEIGHT]% — definition and how to measure ## Automated Scoring Which dimensions can be scored programmatically? Describe the automated check for each. ## Human-in-the-Loop Which dimensions require human judgment? Define the annotation rubric and inter-rater reliability target. ## Golden Set How do you build and maintain a golden evaluation dataset of [NUMBER] examples for [APPLICATION]? ## Feedback Loop How do evaluation scores flow back into prompt improvement? Define the threshold that triggers a prompt revision.
📝 Fill in the blanks
Replace these placeholders with your own content:
[AI APPLICATION, e.g. a writing assistant, a data extraction pipeline, a coding agent]
[NUMBER]
[APPLICATION]
[DIMENSION 1, e.g. Accuracy]
[WEIGHT]
[DIMENSION 2, e.g. Completeness]
[DIMENSION 3, e.g. Format compliance]
[DIMENSION 4, e.g. Tone]
How to use this prompt
1
Copy the prompt
Click "Copy Prompt" above to copy the full prompt text to your clipboard.
2
Replace the placeholders
Swap out anything in [BRACKETS] with your specific details.
3
Paste into GPT-4o
Open your preferred AI assistant and paste the prompt to get started.