GPT-4o
⚡ Productivity
Advanced
Multi-Model Output Comparison
Design a rigorous evaluation framework to compare AI models on a specific task, with test prompts, a scoring rubric, and a cost-quality matrix.
The Prompt
# Multi-Model Output Comparison You are an AI systems analyst who helps teams choose the right model for a specific use case. Task to evaluate: [DESCRIBE THE SPECIFIC TASK OR USE CASE] Models to compare: [LIST 2–4 MODELS] Success criteria: [HOW YOU WILL JUDGE WHICH OUTPUT IS BETTER] Usage volume: [HOW MANY TIMES PER DAY OR WEEK WILL THIS RUN] ## Evaluation Framework ### Test Prompt Set Write 5 representative test prompts — from simple to complex, covering typical and edge cases. ### Scoring Rubric A matrix with 4–6 criteria specific to this task. For each criterion: - Definition (what counts as good vs. poor) - Weight (how much it matters) - Scoring method (1–5 scale or pass/fail) ### Cost-Quality Analysis For each model: - Estimated cost per 1,000 completions - Known strengths and weaknesses for this task type ### Decision Framework A scoring template where I fill in results and it identifies the winner across three scenarios: quality-first, cost-first, and balanced. ### Documentation Template How to record this decision so my team does not repeat the analysis.
📝 Fill in the blanks
Replace these placeholders with your own content:
[DESCRIBE THE SPECIFIC TASK OR USE CASE]
[LIST 2–4 MODELS]
[HOW YOU WILL JUDGE WHICH OUTPUT IS BETTER]
[HOW MANY TIMES PER DAY OR WEEK WILL THIS RUN]
How to use this prompt
1
Copy the prompt
Click "Copy Prompt" above to copy the full prompt text to your clipboard.
2
Replace the placeholders
Swap out anything in [BRACKETS] with your specific details.
3
Paste into GPT-4o
Open your preferred AI assistant and paste the prompt to get started.