The Pelican comparison grid for Astra is pretty interesting
Posted by csmantle 3 days ago
Comments
Comment by suprjami 2 days ago
It's absolutely in the training data now.
Time to retire this (previously mildly amusing) benchmark as frontier models are now pelicanmaxxed.
Comment by conception 2 days ago
Comment by graemep 2 days ago
Comment by suprjami 2 days ago
Comment by orbital-decay 2 days ago
I actually think the pelican and similar tests are much less useful than it seems, but not because of pelicanmaxxing. They are supposed to serve as a vibe check of model's out-of-distribution performance, but do nothing to disentangle the generalization and memorization, which is the hard part. The combination of both is still useful though, and if you look at pelicans over time you'll see their quality is more or less correlated with that, with the exception of models specifically trained to produce vector graphics and 2D layouts.
Comment by williamDafoe 2 days ago
2. Bikes are almost always depicted going rightwards so you can see the drivetrain. 4 failures in this.
3. Only ~7 of the 23 models achieved a diamond bicycle frame shape
4. Almost all the models have the bird directly over the crankset; as little as 4 or 5 have the birds sitting on a set-back bicycle seat.
Comment by amluto 2 days ago
I assume that these have been in the training set for quite a while, but that they’re still a bit hard for the models because it’s not covered by the RL pipelines.