"One way of doing this for planning tasks is to reduce the effectiveness of approximate retrieval by obfuscating the names of the actions and objects in the planning problem. When we did this for our test domains, GPT4’s empirical performance plummeted precipitously, despite the fact that none of the standard off-the-shelf AI planners have any trouble with such obfuscation. "
That's a great test. It shows they're matching prior patterns they saw, even down to what words were used, instead of thinking. We can match prior patterns, come up with the equivalences, and then plan that way. People often slow down when they do stuff like that, though. So, the A.I. would have to be able to do it but slowdowns would be acceptable.
That's a great test. It shows they're matching prior patterns they saw, even down to what words were used, instead of thinking. We can match prior patterns, come up with the equivalences, and then plan that way. People often slow down when they do stuff like that, though. So, the A.I. would have to be able to do it but slowdowns would be acceptable.