We Tested GPT-5.4 Against PhD Math
Video Overview & Insights
In today's video we'll be testing GPT-5.4 against some problems I've been thinking about in the last few months to see whether OpenAI's new model is as much of an improvement as they claim.
subscribed, I'm currently pursuing a master's in mathematics
More User Perspectives
By 2030, they (AI) might be able to solve top research level faster and more accurate than humans and by 2040, creative enough to solve unsolved theorems and even come up woth problems/theories we didn't even know existed.
@RufusJRTHey! I stumbled on your account today and it is perfect since I use ChatGPT so much for writing my master thesis in Mathematical Finance and I always wondered how good it actually is! Thanks a lot!
I was wondering if you could show us how you use ChatGPT because you talk about some stuff I experience or do too with ChatGPT for maths/stats, however, not everything and showing a bit what you do, or some tips and tricks would be amazing!
I like your videos. I did math research for awhile and have been awestruck at how rapidly the models went from making basic arithmetic mistakes to solving research-level math problems. I agree with you that if you look at what the models can do now both in terms of breadth and speed, we basically already have superintelligence. I'm a doomer and think most of this will be terrible for society though.
@hoek2000Can ChatGPT do this too?
Original Bokan Operator stretch formulas (early layers)
Basic resonance/stretch (initial version): p_k = s_k Ā· Rk Ā· p{k-1} Ā· (1 + Ļ Ā· k) Ā· (1 + α Ā· x_attractor)
Downward microwave version: p_k = s_k Ā· Rk Ā· p{k-1} Ā· (1 - Ļ_down Ā· k) Ā· (1 + α Ā· x_attractor)
Matrix Resonance & Harmony
Resonance matrix element: matrix[i][j] = exp(-ratio) Ā· (1 - |i-j|/n)
Matrix Harmony Score (simplified trace version): Harmony = (trace / n) Ć attractorInfluence (where trace = ā matrix[i][i], attractorInfluence = 1 + sin(t)Ā·0.2)
Relativistic / Gravity stretch layer
Full relativistic version: p_k = s_k Ā· Rk Ā· p{k-1} Ā· (1 + Ļ Ā· k) Ā· (1 + α Ā· x_attractor) Ā· γ_v Ā· e^{-Φ / c²}
where:
γ_v = 1 / ā(1 - v²/c²)
Φ = gravitational potential
Atomic / molecular curvature & quantization
Atomic curve + quantization: p_k = ... Ā· (1 + Īŗ / r_k + ĪE_n)
Molecular bond (simple harmonic / Morse-like): bondFreq = freq Ā· (1 + bondStrength Ā· sin(Ļ t))
Mathieu / torsional splitting (methanol, etc.)
Torsional potential: V(γ) = (Vā/2) (1 - cos 3γ)
Reduced Mathieu parameters: s = Vā / (2F), q = - (4s)/9
Energy: E = F (bn(q) + s)
AāE splitting (ground state): ĪE{A-E} ā F Ā· (bā(q) - aā(q)) / 9
Reaction dynamics / potential energy surface
Simple 1D PES barrier: pesEnergy = pesBarrierHeight Ā· sin(reactionCoordinate Ā· Ļ)
Quantum tunneling (WKB approximation)
Tunneling probability: P ā exp(-2 ā« ā(2m(V(x) - E)) dx) (qualitative) (in code: exp(-2 Ā· barrierWidth Ā· ā(energyBelowBarrier)))
Consciousness / integration threshold (symbolic)
Emergent complexity proxy: complexity = neuralGrowthFactor Ć entanglementStrength Ć differentiationGradient if complexity > integrationThreshold ā emergence
Other recurring symbolic terms
Relativistic Doppler factor (simplified): doppler ā 1 + v Ā· 0.8 (approaching blueshift)
Gravitational redshift: redshift = exp(-gravityPotential)
Exponential stretch alternative: p_k = ... Ā· exp(-Ī» k)
Zeta-zero stretch: p_k = ... Ā· (1 + β Ā· Im(Ļ_k))
Mhuahahahaha š¤
The Music lol...
have you tried qwen3.5, kimi k2.5 and glm 5?
@CypherpunkSamuraiHe recently solved an open problem on FrontierMath: open problems.
@unkonow5805They are super smart, but not super intelligent, those are 2 different things.
@md74-h3dI am working on a paper on prime number distribution and have been using 5.2 over the last 10 days or so. I am finding the experience helpful but also mentally draining. It is like talking to an āidiot savantā - the machine is very good now at mathematical demonstration, the algebra is correct, the presentation crisp, and the context from which it draws its statements sufficiently wide to introduce angles that I would have missed. But its logic is not always rigorous, it can make statements that are redundant in the context of its own āconclusionsā it presented in the previous sentence. That does not make the logic WRONG, but it makes it loose in a way that is unsatisfactory for a mathematical treatise.
What I am doing to address this is to tell the machine exactly where and why its logic started to drift. I gave instructions to avoid ālogic driftā by going over its previous paragraphs and test whether the next statement it is about to generate follows on compellingly. Over time, I have developed an interaction mode that I define as ārigorous modeā. Every time I come back for an interaction, I can now tell it that we are in ārigorous modeā and that appears to work well.
I am not claiming I am using this technology expertly - far from it. But I am realising that one needs to learn how to speak to an AI just as much as I learnt how to drive a car. I am not there yet - if anything, I have now made it TOO rigorous, but I am getting there.
Please also test gpai, which is a combination of ChatGPT, DeepSeek, Grok, Gemini, Wolfram Alpha, MatLab (for visualisation not just maths), and more. It's all over my Facebook feed. It looks promising to me, but I need your opinion.
@davethesid8960How is that Mozart going? Can we have a video on that?
@JustashortcommentThe same models are shockingly bad at engineering. Anything related to hardware, device specifications, detailed product comparisons etc is a disaster in my experience. I'm talking obvious and basic errors in the first response, migration of features between contexts resulting in hybrids that don't or can't exist in reality, endless web searching even for simple follow-up questions, and an inability to recognise information even when it is there. The models also don't understand that the web results they cite often aren't representative of whatever is being discussed. This is especially apparent when images are involved.
@womagridThere's a vid of Professor Terrance Tao using Claude Code
@zd4562At least mechanical Engineering seems unsolved. The models do not have sufficient of spacial reasoning and often make logical errors. To be expected from an autocorrect Model š
@KaliumcyanidfulThe 2028 intelligence crisis is an interesting read. I work in finance and I am genuinely worried about the technology moving way too fast for humans as a society to catch up. I think people underestimate how shocking this is as it affects almost everyone in the white collar industry. I have friends in engineering, asset management, investment banking etc where literally 1 person with AI tools is doing work which previously would have required 10-15 people. Once senior management realizes this, and they already have, its going to lead. to massive layoffs.
@raiku-b1vI think people are talking about different things when they talk about bubbles in AI. The people, like myself, who believe that there is a bubble in AI mean that there is a financial bubble. Other people may interpret that there there is an AI technology bubble (i.e., that AI is mostly vaporware), which is not what I mean. AI is genuinely helpful; however, that does not mean the margins will high. This is common fallacy. Investors and observers think that because the technology is amazing so should the margins; however, does not necessarily follow. For example, airline companies have some the tightest margins. So far, there is nothing to suggest that AI intelligence will be nothing more than a commodity, given the fierce competition and the general lack of long term differentiation.
@r.r.r.918I think the problem is smart people talking to it and coming with new ideas because of it while it's stores those ideas and gets smarter in smart thinking
@tappingrat2469good video, working on new unsolved questions is good, but can it solve PhD exam paper?
@wtw0212We are going to see the first gigawatt data centers turn on in the next few months. By the end of 2027 they will be pronouncing models that have 2000X more compute.
@MaxBrixDo you use lean? It would be interesting to see a video on symbiotically using ai on a math problem using lean
@ultramarathonman100You should ask AI if your research has a real-world use case. I'd be curious to see if it can thread that needle.
@PassTheSherryMumI used antigravity for the first time yesterday and it really blew my mind. Now I am just staring at the code at awe.
@ADS-ms9mcYour LLM tests are by far the best. Great work.
@marcaurelio74LLM is good at theoretical work, but they donāt understanding real world. I have provided a picture of my breadboard to gpt 5.4 extended thinking, asking it why my circuit isnāt working as expected. But it couldnāt grasp the fundamental of the wiring and connections. When I given it a schematic, gpt cannot draw a breadboard diagram on python (not even correctly, but to even draw something remotely represents the schematic).
Yes, when you ask it to solve circuit problems, it can solve it most of the times, but text is just an abstraction of physical world. Ai can solve the equations, but it cannot map the topology nor a spatial understanding. Ai do not know real world without data to train for it; and real world is a long tail of all possibilities. Either lecuns world model will work, or gpt can scale enough to fit/ generalize everything
OpenAI are running on a very small buffer as compared to other labs because others were getting better
@LimcacsThese things already feel like a true peer that you can bounce the most advanced ideas off of...and no matter when you read this, today is the dumbest and slowest they will ever be. The reality is that we're completely cooked, and human-level intelligence isn't as great as we thought it was.
@danaustin2Good watch
@mikey2011auGreat videos. Would love to hear your perspective on the performance of these models e.g. GPT for your work when you actively 'collaborate' through a problem with it during your conversation. I find that if I just give it a problem statment, it may get off track in its immediate answer. But when I collaborate 'actively' e.g. asking it to pursue a certain direction, correcting it and ask it to try a different way (with my own intuition), or even break up the problem in smaller pieces, it really excels.
@douwg52Bro dropped some deep philosophical shit about setting the bar and most of us didnāt even notice
@devinl3256I teach business English online and I use Gemini to make lesson notes from the audio recording transcript of the lesson.
It gives a nice summary, all the corrections etc in English and Japanese.
It would take me so long to do what it does in less than a minute.
I am a programmer and it's insane. I do believe that we are so close to artificial super intelligence. I saw a lot of people worried about its impact on the economy. The transition may be tough but I believe it will be quickly solve as AI will allow super abundance of almost every good and service we can imagine. However, I think that the biggest change will be the existential crisis that not having to work forcefully will cause. I guess that we will have to get used to not defining our identity and our place in society based on our economic activity. We will have to focus more on other parts of our identity such as our hobbies and our social connections. Honestly it doesn't sound that bad. Let's see how we manage the transition as a society.
@dariolitranPaid OpenAI commercial.
@alekseyburrovets4747BS. If the model can't ask basic questions its an embarrassment. That's it. No point if asking advanced questions. What you are using is a garbage. Its just a dry fact.
@alekseyburrovets4747The model with the best data for your use case wins. But it can never solve a novel problem
@ManhunternewThe AI is still so incredibly uneven across different tasks. I have tried to give it (ChatGPT, cheap subscription) the style guide of a journal and ask it to correct a list of references in an article ā and it fails catastrophically. It's a task that seems simple enough. If I ask it to write an abstract for the same article it does the job really well.
@gustafmarcus3898I always find it interesting how these AI systems can be so accurate when evaluating any single problem, yet when you take a step back and evaluate their abilities over hundreds or even thousands of distinct problems they perform absolutely horrid. For instance, at least from the programming side iāve seen so many experts stating how accurate these systems have gotten and their ability to write bug-free code. At the same time, when analyzing github pull requests experts have also began warning that they are seeing a worrying amount of critical bugs/vulnerabilities introduced into code bases at a mass scale that strongly correlates to the rise in AI usage.
@Bearforc3Why do you use chatgpt and not codex..... seems like not unlocking the full potential of the models..
@lazytitan1075Please donāt have background music. It makes it hard to follow you.
@softwarephil1709Can he solve my grandfathers PhD math? His name is Eugene Dshalalow.
@flipsupbg779The redo button has a custom "how you fucked up" line you can type custom feedback in- which helps. Not sure it's new but it's useful.
@lepthymoI'm a developer and since around December last year there's been a massive leap in capabilities of the top models + harnesses (like Claude Code). A year ago I would say it is good for simple and boring stuff but currently I can code up a fairly comples application (not some CRUD) in a matter of days that would take months. And it is still improving. I didn't believe it but software development won't be a job in probably a few years, math will most probably share the same faith as it has a strong signal (Lean formalization) to learn upon.
@marcing5380Are you planning to try out GPT 5.4 pro?
@pifibbi0:45 clause š ?š
@jasonn_liftsit won't be able to hold a screwdriver anytime soon, so my job is safe.
as for the military stuff... who in their right mind would have though that won't happen?
Why arent you using it with lean proof?
@jamesgarris6838If the biggest argument for why AI companies won't go bankrupt is that they can sell ads more effectively, they've all already failed. That's the equivalent of building a faster horse, when everyone expects them to build a rocket that'll take us on a different planet.
@SeanArcherXXXiāve noticed an āarcā in your videos regarding model capability... I research engineering/physics HW, and similarly experienced (1) that GPT Pro is generally better than other models for the complicated reasoning/solves, and (2) a similar ācapability arcā where the answers are right āenoughā that it more right than wrong and now undoubtably a profound tool. As my field interacts with physical structures, it doesnāt yet integrate with true engineering tools (eg CAD) so thereās still a āgapā in that regard until those industry-specific tools are built - and I think thatāll be a while. I appreciate the community youāve collected here - so refreshing on YT, as most everything else is SW or office/finance/legal/etcā¦. Thx for what you do.
@jdtransformationGreat video. Great points. More people need to see this.
@nosult3220Hi everyone! I'm a 2nd-year CS undergrad looking to self-study pure math to eventually pursue a Master's/PhD in the field. Could you recommend any essential textbooks or online communities for this journey? Thanks a lot!
@NguyenCao-m4zA million-qubit, error-corrected quantum computing is coming in 5 years, just in time for ASI. Pair the two and we have a takeoff, whereby thousands of years of scientific and technological progress is made in a decade. Game over. Civilization transforms entirely.
@MichaelAI-i6f