HyperAIHyperAI

Command Palette

Search for a command to run...

Résultats de performance de différents modèles sur ce benchmark

Metrics

Ours (%)
Binary (%)
Self-Rewarding (%)
All
Format
Length
Content
Keywords
Language
Startend
ChangeCase
Combination
Punctuation
Human
GPT-4o
Prompt
GPT3.5-turbo
LongFormer-Base-4096_3k
LongFormer-Large-4096_3k
Inst-level loose-accuracy
Inst-level strict-accuracy
Prompt-level loose-accuracy
Prompt-level strict-accuracy
4 lignes au total
IFEval | SOTA | HyperAI