my hypotheses about 4.8: - it is more disagreeable by default, so when you use custom prompts that pushed previous models to disagree more, it goes beyond well calibrated pushback to contrary strawmanning - it tends toward unusual phrasing, neologisms and novel metaphors - maybe the result of too much synthetic data rephrasing with incentive for avoiding cliche?
Opus 4.8, the Think Different model