A new benchmark called PatternEval exposes how shortcut inference modes in hybrid-thinking AI systems produce dramatically more visible errors, even when accuracy scores look fine on paper. Tencent r⦠[+2962 chars]
No comments yet. Be the first to comment!