DeepSeek/deepseek-v4-pro and z-ai/glm-4.7 are the two models that hallucinated answers - notably, these are text input only models. It looks like the PDF was being treated as a sequence of images.
Testing this via OpenRouter seems risky to me. It would be interesting to see results for this without a proxy in the middle that potentially confuses the results.
DeepSeek/deepseek-v4-pro and z-ai/glm-4.7 are the two models that hallucinated answers - notably, these are text input only models. It looks like the PDF was being treated as a sequence of images.
Testing this via OpenRouter seems risky to me. It would be interesting to see results for this without a proxy in the middle that potentially confuses the results.