IMO if someone gets AC during a contest that AC should stand. Not hard to understand and has been that way for forever. It's very punishing for a user to get 4 AC and then find out a week later they actually came 1000th and lose giga rating because problem organizers can't figure out how to set problems. Look, its really simple
- stop putting 1000 testcases in the judge and gumming up the judge. It creates all sorts of bad side effects, solutions with good heuristics can outperform, and so on. You don't need that many test cases. Look at CSES, CF, etc.
- let all codes that got AC stand, or unrate the whole contest if it is very bad.
- test with python and make sure there is a decent buffer (eg. judge's solution solves with 10-20% of the allotted time)