Prompt Library ⚙️ Technical AI Output Evaluation Framework
GPT-4o ⚙️ Technical Intermediate

AI Output Evaluation Framework

Build a systematic framework for scoring AI-generated output quality using rubrics, automated checks, and calibrated human review.

👁 0 views ⎘ 0 copies ♥ 0 likes

The Prompt

# AI Output Evaluation Framework

You are an AI quality specialist. Design an evaluation framework for [AI APPLICATION, e.g. a writing assistant, a data extraction pipeline, a coding agent].

## Evaluation Dimensions

Define [NUMBER] quality dimensions and their weights. For [APPLICATION]:
- [DIMENSION 1, e.g. Accuracy]: [WEIGHT]% — definition and how to measure
- [DIMENSION 2, e.g. Completeness]: [WEIGHT]% — definition and how to measure
- [DIMENSION 3, e.g. Format compliance]: [WEIGHT]% — definition and how to measure
- [DIMENSION 4, e.g. Tone]: [WEIGHT]% — definition and how to measure

## Automated Scoring

Which dimensions can be scored programmatically? Describe the automated check for each.

## Human-in-the-Loop

Which dimensions require human judgment? Define the annotation rubric and inter-rater reliability target.

## Golden Set

How do you build and maintain a golden evaluation dataset of [NUMBER] examples for [APPLICATION]?

## Feedback Loop

How do evaluation scores flow back into prompt improvement? Define the threshold that triggers a prompt revision.

📝 Fill in the blanks

Replace these placeholders with your own content:

[AI APPLICATION, e.g. a writing assistant, a data extraction pipeline, a coding agent]
[NUMBER]
[APPLICATION]
[DIMENSION 1, e.g. Accuracy]
[WEIGHT]
[DIMENSION 2, e.g. Completeness]
[DIMENSION 3, e.g. Format compliance]
[DIMENSION 4, e.g. Tone]

How to use this prompt

1
Copy the prompt

Click "Copy Prompt" above to copy the full prompt text to your clipboard.

2
Replace the placeholders

Swap out anything in [BRACKETS] with your specific details.

3
Paste into GPT-4o

Open your preferred AI assistant and paste the prompt to get started.