Skip to main content
The following LLMs have been tested against three criteria:
  • Tool usage: Does the model use tools correctly and consistently?
  • Answer depth: Can the model handle medium to complex tasks beyond basic UI actions?
  • Truthfulness: Does the model provide accurate answers without hallucinating?
🟢 Good   🟠 Partial   🔴 Poor