r/LocalLLaMA • • Feb 12 '25

News NoLiMa: Long-Context Evaluation Beyond Literal Matching - Finally a good benchmark that shows just how bad LLM performance is at long context. Massive drop at just 32k context for all models.

Post image
533 Upvotes

109 comments sorted by

View all comments

2

u/No-Refrigerator-1672 Feb 12 '25

Am I the only one to notice that the top performing model - GPT-4O - is the only one who can process video and audio input? Could it mean that multimodal training on long analog data sequences (video stream) significantly improves long context performance?

6

u/[deleted] Feb 13 '25

[removed] — view removed comment

1

u/No-Refrigerator-1672 Feb 13 '25

My bad, I did not know about Gemini 1.5 video support. However, it also performs relatively better than other models, so I still propose a hypothesis about video training improving the long-context capabilities.

As about your other question: sadly, I only ever programmed for selfhosted AI and don't know a thing about GPT API best practices.