One single score can almost never tell {estimated AI participation extent, type of AI works, detector confidence, etc.} at the same time. Humans using predictor outputs for decisions should be more careful.
NeurIPS 2026 just desk-rejected hundreds of papers because an AI detector said they were AI-written (blog.neurips.cc/2026/06/02/ai-…).
I understand why they did it. Reviewer time is scarce, and a flood of slop submissions is a real threat to peer review.
But here's where it gets


