Caltech Mathathon – first hackathon ever devoted to research level mathematics
Posted by astroanax 1 day ago
Comments
Comment by brian-bfz 19 hours ago
- We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors.
- We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants.
- Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Comment by danielmarkbruce 15 hours ago
Comment by brian-bfz 14 hours ago
We allow prior work as long as it's labelled. We record chat logs, so it's easy to verify what's prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
Comment by ajkjk 14 hours ago
just something to think about
Comment by brian-bfz 14 hours ago
Comment by danielmarkbruce 12 hours ago
Good going, hope it goes well.
Comment by maypop 18 hours ago
I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
Comment by brian-bfz 18 hours ago
Comment by fred123123 19 hours ago
Comment by brian-bfz 19 hours ago
Comment by fred123123 19 hours ago
Comment by jegutman 18 hours ago
Comment by queuebert 17 hours ago
All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
Comment by ajkjk 14 hours ago
Now, I think it's the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I've been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I'm an AI skeptic in many other ways; it's not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been something of a disaster for a long time (thanks, largely, due to the academic incentive structure which heavily favors novel results, no matter how esoteric).
Comment by suriyaG 17 hours ago
why is that a concern in this context? would you have asked the same about steam engines and horses?
this is a really cool concept, organized very well. and that is very commendable.
Comment by lovasoa 17 hours ago
Comment by tim-kt 16 hours ago
Comment by suriyaG 12 hours ago
Comment by hatsix 15 hours ago
Comment by suriyaG 12 hours ago
do you mean,
> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.
is good argument?
of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc. everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.
why is that not worthy enough to pursue?
Comment by tacomonstrous 14 hours ago
Comment by brian-bfz 16 hours ago
Comment by reasonableklout 15 hours ago
Comment by brian-bfz 14 hours ago
Comment by reasonableklout 10 hours ago
Comment by danielmarkbruce 15 hours ago
Comment by ktallett 17 hours ago
Comment by a2ff6eeb0 17 hours ago
Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.
We can't put this genie back in the bottle.
Comment by mattmcal 16 hours ago
Comment by a2ff6eeb0 16 hours ago
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?
Comment by mattmcal 15 hours ago
Comment by a2ff6eeb0 15 hours ago
Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?
Comment by tacomonstrous 13 hours ago
Comment by a2ff6eeb0 5 hours ago
Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.
Comment by ktallett 7 hours ago
Comment by a2ff6eeb0 5 hours ago
Comment by nill0 8 hours ago
Open Problems in Computational Geometry Listed by Erik Demaine, Joseph Mitchell, Joseph O'Rourke in 2024
And also the popular list below, which contains some of the frontier problems and undefeated beasts that have remained unsolved for decades, some even for centuries.
https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_m...
Caution: Solvability is not guaranteed!
Comment by Semkas 1 day ago
More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
Comment by brian-bfz 19 hours ago
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
Comment by falcor84 23 hours ago
Comment by Ey7NFZ3P0nzAe 6 hours ago
So far humans failed at those problems. Also IIRC there was a guy that proved a substantial problem 2-3 months ago by basically pasting over and over "keep looking for a solution" or something like that for 2 days with little formal math background.
Comment by Donald 23 hours ago
It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
Comment by rookienumbers30 23 hours ago
I wonder how models perform on finding analogies between analogies
Comment by Aboutplants 1 day ago
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
Comment by a2ff6eeb0 1 day ago
I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
Comment by charlieyu1 1 day ago
Comment by a2ff6eeb0 1 day ago
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
And left it for a long time. Jarred isn't a mathematician, he's the maintainer of a janky JavaScript environment.
Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...
Comment by hgoel 1 day ago
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
Comment by charlieyu1 22 hours ago
Comment by hgoel 21 hours ago
Comment by a2ff6eeb0 20 hours ago
7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.
Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2
You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.
Comment by ewjt 13 hours ago
>"illegible reasoning in a few reinforcement-learning environments over long rollout"
Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.
This video presentation of the paper you linked was interesting: https://www.youtube.com/watch?v=hUp3zh23aHw
Comment by hgoel 20 hours ago
Comment by a2ff6eeb0 20 hours ago
Comment by hgoel 19 hours ago
Comment by charlieyu1 17 hours ago
Comment by MostlyStable 20 hours ago
Comment by a2ff6eeb0 19 hours ago
You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.
Comment by youoy 18 hours ago
Comment by jhonof 1 day ago
Comment by a2ff6eeb0 19 hours ago
Comment by loloquwowndueo 1 day ago
Comment by falcor84 23 hours ago
Comment by thatseasy 1 day ago
Because the companies that run frontier models are malevolent by every metric.
They are destroying the environment, especially those in neighborhoods of low income people.
They are empowering their owners who are some of the most deplorable and duplicitous people living.
They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.
They stole the entire creative output of humanity and are trying to sell it back to us.
They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.
They are being used to kill in war and for surveillance.
Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.
I for one, am one who walks away from Omelas.
Comment by Fraterkes 1 day ago
Comment by groundzeros2015 21 hours ago
Comment by corinthia 21 hours ago
caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
Comment by brian-bfz 19 hours ago
Comment by chaoxu 1 day ago
Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.
Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
Comment by deeznuttynutz 9 hours ago
Comment by youoy 1 day ago
Comment by brian-bfz 19 hours ago
Comment by youoy 18 hours ago
I am also a bit frustrated seeing maths go in the direction of prompt enginnering. I am afraid of a world were a math phd student cannot go one week thinking about a problem without prompting an LLM to give him/her an invented answer. Something is lost along the way.
For me maths is not Lean, or formal systems, or an agent reasoning about formal systems to join literature from different fields. I see the value of it, but i think it will make it way more difficult for students (and profesional mathematitians) to see beyond that. And i see us heading into a reality were those who think like me will in practice remain a minority for quite a few years/decades because the low hanging fruit of LLMs will be to vast to ignore.
Comment by brian-bfz 17 hours ago
Mathathon's goal is to reshape rather than stop LLM use. Can we set high standards for LLM use? Can we highlight the roles of a mathematician beyond proof generation? Can we redesign our incentives to promote these standards and roles?
I'd love to hear your thoughts on how to improve this event. We're very open to criticisms.
Comment by youoy 8 hours ago
I am mathematitian that is working as a software engineer. I have see first hand what these models are doing to SE. Its not that writting well thought, and compact code is not a good idea anymore, or that it doesnt beat LLM code, but that the people that see the value are a minority. If you write a piece of old school code in an LLM repo, it doesnt really make a difference, because old school code requires a team effort.
In maths the situation is not exactly the same. Probably reading a good piece of well thought math inside a book of LLM assisted proofs will stand out so much for the carefull reader that there will be no question about the value.
However, if the only way to get a position at a university is to print as much papers as possible, then who would risk printing 1 paper instead of 5 for doing old school maths?
So the solution to this is building a culture and community effort around these topics. And for that the topics need to be openly discussed.
Good luck with the organisation! And thanks again for the conversation!
Comment by isotypic 21 hours ago
Comment by bwfan123 20 hours ago
I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.
AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or humans being duped as reverse-centaurs.
Comment by a2ff6eeb0 19 hours ago
Comment by ianm218 19 hours ago
This seems like a strawman. It’s certainly not consensus that AGI has been achieved and I don’t think the people participating in this event feel like there is no value in human input or steering the AI.
Comment by tzs 23 hours ago
> This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster?
and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.
Comment by blondie9x 1 day ago
Comment by sb10128 1 day ago
Comment by charlieyu1 1 day ago
Comment by boothby 23 hours ago
Well, that's pretty damned ignorant; I was attending William Stein's hackathons on the BSD conjecture and the Sage Math project nearly 2 decades ago.
Comment by amelius 23 hours ago
Comment by viccis 21 hours ago
Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.
Comment by ondrejdvorak 12 hours ago
am I eligible?
Comment by xqcgrek2 20 hours ago
Comment by milkshakes 8 hours ago