- Tool usage: Does the model use tools correctly and consistently?
- Answer depth: Can the model handle medium to complex tasks beyond basic UI actions?
- Truthfulness: Does the model provide accurate answers without hallucinating?
🟢 Good 🟠 Partial 🔴 Poor