GPT-5.6 just made itself CHEAPER
Video Overview & Insights
Hyperagent is giving away $500 to the first 500 people who sign up! https://www.hyperagent.com/forwardfuture500
โTask completedโ matters. But โtask completed correctlyโ may matter even more.
That one word can mean the difference between a finished job and hours of correctionsโor, in some cases, critical failures.
On small projects, the difference may be easy to miss. On large projects with complex codebases, forgetting โcorrectlyโ can become very visible, very quicklyโand very expensive.
Join My Newsletter for Regular AI Updates ๐๐ผ
https://forwardfuture.com
Anthropic is like driving a monster truck to buy milk. ๐ค
My Links ๐
๐๐ป X: https://x.com/matthewberman
I used to be a claude fanboy, then I reached my weekly limit in 1 day and it made me test codex, and omg it's so much better. Claude codes like a crazy dev that doesn't really verify but codex seems like a full team of devs.
๐๐ป Forward Future X: https://x.com/forwardfuture
๐๐ป Instagram: https://www.instagram.com/matthewberman_ai
Why is everybody correlating price with cost? They can lower the price as a business decision even if inference costs rise. RAM is pricier, compute is pricier, energy is more expensive. None of that has been solved. So no, dont assume any of these lower prices come inevitably from higher eficiencies. They are not releasing the cost structures so we can't know (they cannot demonstrate) they got more efficient.
๐๐ป Discord: https://discord.gg/u7wTTGWhuJ
๐๐ป Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8
Does it mean that you get more usage from your chstgpt plans or are the discounts only applied to api based usage?
Media/Sponsorship Inquiries โ
https://bit.ly/44TC45V
That is crazy, I find myself switching my anthropic max to gpt pro sub back and forth more times then a single month of the subscription. AI race is crazy
Links:
https://x.com/OpenAI/status/2082878156483219672
Could you research how big/small Luna and Tera are? Are these newer versions of GPT-OSS 20B and 120B ?
https://artificialanalysis.ai/
Cost per task is king. Glad they are doing this metric
More User Perspectives
ChatGPT says the Max setting is a waste. Use high.
@jpw5820You really must prioritize talking about Chinese open source model EFFICIENCY. The cost gains per task efficiency out does Americans models and many many of your subbs will gravitate towards these Chinese models. Js saying
@paulmuriithi9195Improved itself? They probably used the same method Deepseek did and distilled the larger model.
@drob9673I just hope competition does not result in safety minded actors being acquired by megacorps.
Competition can not just be about price and potency of a product if we care about safety and other concerns.
Fable to Pro and 100% next?
@MarkoTManninenI want to see Fable 5 price drop!
@kalloh5519I agree. Iโm surprised more people arenโt talking about this. It feels like the beginning of real RSI, and if it delivers, it could give American AI labs a meaningful commercial edge over Chinese labs. If competitors donโt catch up quickly, this advantage could last for quite a while.
@haggle196You assume price == ecffeciency. Im guessing luna is under used and scam altman needed to distract you.
@brianhansen6481Makes sense. Distill from your own models.
@amazingjoe76need to be comparable to your own pc to prevent open source to clame top spot.
@benjamindemontgomery6317That's why we should keep Chine on the loop of everything. They will keep the leeches from sucking too much blood
@MarioMartinez_Is Luna basically like haiku though? also does this mean I will also be able to get more out of my codex usage?
@jfpicturesRecursive self improvement is now necessary to remain competitive...
@raccoon351๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ๐ฉ
@_Jax_55Yeah I've already decided to stay with OpenAI Chatgpt for now... The way they did this make it more efficient made it more accessible. That is a lot better than the elitist exclusivity of 'Anthropic' and Dario's arrogant statements saying we are now in the era of 'expensive tokens'.. I know, intelligence is a private good but considering he made money out of the collective knowledge of he internet, its unfair for Dario to make this access highly 'inaccessible' by making claude use soo expensive!
@realme72onlyThe only reason OpenAI and Anthropic are ahead is because the US prohibits sales of advanced chips to China, and China doesn't have EUV lithography yet.
Once China even gets halfway to catching up in EUV and starts making remotely on-par chips, it's game over for OpenAI and Anthropic. But that might not be for several more years.
The usage on the subscription didn't improve tho
@eletrohitsbrluna and terra is terrible. Charts are just bs.
@gordonfreimannMusk might be able to reach them because he might throw more compute. He has largely caught up, if he can continue have ore compute, more spend (and he can) and doesn't do too badly in the talent department... The other one that cold is Google, but they are currently a long way behind and would have to make a lot of use of the massive compute they have.
@JonathanBerry-u6lThis isn't recursive. Models are trained from data, and the algorithm hasn't changed. They aren't just a code to optimize. These are just task running platform upgrades or hosting upgrades. Escape velocity is only hit once it can optimize and reduce the training part, and not even but find a way to keep improving without more data, which is a real-world physical limit. Not the runtime.
@mh72videosWhen comparing the work of Luna Terra with models even such as Opus 4.6, you will be surprised that on paper they should work better but in reality they suck.
@saturnfrakso more slop for cheap
@Danish-its-meWaiting for china to do some "efficiency improvements" and 1 up the price drop
@aes9217this is why we need regulation
@byronjuarez5244Anthropic should use fabel to do the same and become at least a little efficient.
@arghyashrivastav8559wtf it's free.!
@Morwicyouโre running open claw ๐ and youโre nervous about open ai making improvements??? ๐๐๐๐
@deestortIt's hilarious how everyone counts OpenAI out, done, finished, prepping for bankruptcy, etc.
THEN BOOM.
5:40 I thought they are not supposed to train on our data?
@mikedodger7898I have been running codex 20hrs per day for the last 10 days on Sol Extra High, Im only on the 100$ plan
@wn352How many sponsors per video, lol. I stopped watching
@tube5607-i4nIts worth noting the graph looks like logarithm on x axis - not linear. Does not change your argument - but worth noting
@iainmackenzieUKChina is paying you. NBC is now pushing media narratives of "America needs an open source solution" - that would be cancer to persue harder than it is
@Digitalnomad-o9bNecesitamos que habilites la traducciรณn a espaรฑol.
@patojp3363How are they going to pay back the VCs?
@alan83251this is 'recursive self improvement" with diminishing returns. Next iteration, it'll find far less places to improve, but faster, and then even faster to discover even less.
This is not the singularity. Until it can start coming up with actual breakthroughs, it's just finding errors fast. And, no, programs that you've found errors in don't just magically have more errors.
4:51 they wrote up a blog post .... "Sol wrote up a blog post" ๐คฃ
@blissweb4:30 after using Sol the past few months I would say that not many OpenAI engineers know what that means. Sol is great at doing stuff which is way over people's heads and making up a legit, engineer sounding, name for it. ๐๐คฃ
@blisswebโฆ anyone else remember where we were at in January?
Canโt imagine where weโll be by 2027, the rate of progress is mind boggling.
I don't get it. Even the best frontier models screw up so often. Why would I want to save some money using a cheaper model that will certainly screw up even more and add to my manual error checking and verification work? ๐
@grahamashe9715How can I trust that they are giving us unquantized model APIs, i randomly see that the model quality goes nose dive and then gets back up at random hours. i'm unable to prove or quantify it though.
@user-tk7sc4gz2vLooks like someoneโs trying to maneuver themselves out of bad press
@SFJayAnt