Sorry, but there’s a very big difference between using my graphics card that is there for my entertainment to run an LLM using standard air cooling (aka a small fan with a shaped brick of aluminium) and:
building huge complexes that can’t be used for anything else
use evaporative cooling, especially in dry areas, because the republicans there make it easier to simply ignore local resistance
use fucking gas turbines poisoning the air in the area.
increase the power bill of everyone around them because fuck the common people
The amount of resources that are used for the 2 scenarios are not even comparable. I paid for my single GPU with Vram to be able to play games, running small specialist local models is simply an plus this enables. I pay for my power bill - my GPU maxes out at around 250W, which is 1/4th of my microwave - and that bill isn’t as egregious as what the US power companies are currently pushing on their private clients to expand their power generation.
And the most important difference: local open models serve me and not the interests of someone else.
If you can’t see the difference, i am very sorry for you, because a future where we are in control of local models would be very much more preferable to the shit the hyperscalers are trying to push currently.
Training of models is a one-time cost. If that model gets widespread use, then the cost of creating it in the first place becomes less and less relevant. This means we all would profit from creating “creative commons”-models that can be used privately free of charge, while licensing them out to businesses to finance the costs of licensing material for the next “update” of the CC model. Someone more involved might use his calculator to know the ideal update frequency (6 months? 1 year? if its really expensive maybe every 2 years).
It’s really a shame that in order to stimulate adoption to increase development of LLMs they have to do so much damage first. It’s not unlike with cars, scaling up from the very first ICE models has pretty much doomed us all and even now vehicular transport is still at a net loss.
By the time LLMs or even AI become efficient enough that they can run on the more and more advanced tech of handheld devices (maybe quantum?), it’ll have done so much damage, not to mention all those datacenters will need a new purpose. I hope to see it in my lifetime because of advancement life I that ever happens, I’m holding out hope AI will help us do stuff so much more efficiently we can actually reverse all of the damage done.
Either that or it’ll just kill all humans deciding we are the problem.
Using and talking positively of self hosted LLMs still promotes LLMs and convinces people to use them (most likely using services owned by large corporations)
Its still something that can and has caused people to not trust their parents, doctors, believing what the chatbot says instead. Its still something that can and has caused people to kill others or themselves
If used on the internet, it still contributes to it becoming worse (even without considering my first point)
It isnt as bad, but it still has problems. Whether the use is worth it is something that i have not seen discussed at all (im really curious about the uses for llms, the technology seems cool with actual uses, but every example i have seen was terrible)
(Also, isnt the energy issue with the training, not the usage?)
LLMs are not the devil, they are useful tools. Talking positively about local models is simply that - informing the surroundings that you can have a model that is NOT bound to a hyperscaler. We should talk MORE about local models to make people aware that there is a way forward without OpenAI and Anthropic.
Anyone who will put in the time to actually work with a local model will also be aware of the limitation of those things - another point for “make people own their own shit”. Anyone who involves themselves in running a model will also be keenly aware that they are NOT TALKING TO A SENTIENT BEING (I fault Sam Altman and Dario Amodei personally for spreading this bullshit, they should be sued and thrown in jail)
The energy cost of training a model for local use is not even comparable to the huge amount of resources that are used for training something like Mythos - these things don’t scale linear, but exponentially. And it is a one time cost, which means the longer a model is in use the less relevant the training costs become (just as buying a car -the upfront costs are high, but in the long run you will pay more in gas and maintenance than you paid for the car itself). Even if you train one large model on the scale of Mythos, you then have the possibility of distilling those models into smaller ones for local use, specialized for specific areas at a fraction of the cost.
I had a long reply written, but got deleted because my phone ran out of battery; i doubt any of our opinions about the existence and relevance of the discussed downsides would change if this was argued further anyways (and i am too lazy to rewrite the reply)
Your reply doesnt answer the question asked in my comment. You say its a tool, but didnt say what you use it for
(Also, the first sentence seems like a generic reply to any comment anti-ai, it doesnt say anything and its not even relevant in this situation; it only makes your arguments feel worse)
I disagree regarding the “generic reply to Anti-ai”. Discourse nowadays simply ignores locally run models - they do not exist outside of a sphere of technically adept users. Making sure more people know that you don’t have to finance OpenAI, Anthropic, Meta or the Neoclouds should be way more accessible information.
This tech and what is happening currently should not be conflated. If Big Tech used up all global stockpiles of bubble gum because they wanted to create the worlds largest chewing gum bubble, it doesn’t mean that you with your single pack of gum is even comparable to what they are doing.
The only question in your comment i can see is the concern about energy cost while training a model. Those costs are a static amount - not ignoreable, but finite. The longer you use that model, or increasing the amount of people using it, reduces the cost per instance. Use a model long enough and the energy cost of creating in in the first place converges towards zero. It’s another point for local LLMs - nobody needs the aggressive update cycle Anthropic and OpenAI are running, and those can’t stop because the moment they do, investor money will stop flowing.
For what it can be used: well obviously it is a great tool in all areas where you can automate checking if what it did meets a specific criteria. That’s the reason why Mythos has Business entities like Microsoft scrambling like mad just right now - checking to see if Mythos managed to escalate privileges or run remote code is trivial, and Mythos is showing up security issues all over their software, up to the point that MS is scared that even if they patch the really ugly shit, Mythos might still chain multiple low level vulnerabilities to fuck them over, and all that is needed from Joe Shmoe are the words “go hack Onedrive”.
I personally like to use them to calculate stuff where i know it can be calculated because of my experience, but where i do not recall the specific formulas needed. I also use it for creating regular expressions on the fly, teaching myself how to create useful bash scripts (all my knowledge about sed and awk comes from a local LLM), among other things that are not the business of anyone - that’s why I am using a local model in the first place.
I honestly forgot that i also put a question about training cost; by the question i meant when i asked about uses for ai, i didnt realize i made it unclear whether it was a question or not
I dont think a model that you have said is bad for the environment is a good example
Regarding the first sentence, i have said that this seems like a generic reply because i have not said that llms are the worst thing ever (even saying that their uses could offset the downsides) or that they do not have uses (“AI is just a tool” is also an argument i see so often, never with any explanation behind)
For such impressive technology, i doubt the best use of it is a better search engine that you have to fact check or writing help
I have seen someone use a local model to control home assistant
Oh sorry, i didn’t realize that. OK, so looking at the thread, i would say my reply would fit the inane FUD @[email protected] is spouting. Klear, if you wanna reply, i am open for discussion.
Sorry, but there’s a very big difference between using my graphics card that is there for my entertainment to run an LLM using standard air cooling (aka a small fan with a shaped brick of aluminium) and:
The amount of resources that are used for the 2 scenarios are not even comparable. I paid for my single GPU with Vram to be able to play games, running small specialist local models is simply an plus this enables. I pay for my power bill - my GPU maxes out at around 250W, which is 1/4th of my microwave - and that bill isn’t as egregious as what the US power companies are currently pushing on their private clients to expand their power generation.
And the most important difference: local open models serve me and not the interests of someone else.
If you can’t see the difference, i am very sorry for you, because a future where we are in control of local models would be very much more preferable to the shit the hyperscalers are trying to push currently.
Iirc, the local models you run had to be developed in the hige datacenters first
Models can be trained on consumer hardware. They’ll just be limited in size and scope compared to larger models.
Training of models is a one-time cost. If that model gets widespread use, then the cost of creating it in the first place becomes less and less relevant. This means we all would profit from creating “creative commons”-models that can be used privately free of charge, while licensing them out to businesses to finance the costs of licensing material for the next “update” of the CC model. Someone more involved might use his calculator to know the ideal update frequency (6 months? 1 year? if its really expensive maybe every 2 years).
It’s really a shame that in order to stimulate adoption to increase development of LLMs they have to do so much damage first. It’s not unlike with cars, scaling up from the very first ICE models has pretty much doomed us all and even now vehicular transport is still at a net loss.
By the time LLMs or even AI become efficient enough that they can run on the more and more advanced tech of handheld devices (maybe quantum?), it’ll have done so much damage, not to mention all those datacenters will need a new purpose. I hope to see it in my lifetime because of advancement life I that ever happens, I’m holding out hope AI will help us do stuff so much more efficiently we can actually reverse all of the damage done.
Either that or it’ll just kill all humans deciding we are the problem.
Using and talking positively of self hosted LLMs still promotes LLMs and convinces people to use them (most likely using services owned by large corporations)
Its still something that can and has caused people to not trust their parents, doctors, believing what the chatbot says instead. Its still something that can and has caused people to kill others or themselves
If used on the internet, it still contributes to it becoming worse (even without considering my first point)
It isnt as bad, but it still has problems. Whether the use is worth it is something that i have not seen discussed at all (im really curious about the uses for llms, the technology seems cool with actual uses, but every example i have seen was terrible)
(Also, isnt the energy issue with the training, not the usage?)
LLMs are not the devil, they are useful tools. Talking positively about local models is simply that - informing the surroundings that you can have a model that is NOT bound to a hyperscaler. We should talk MORE about local models to make people aware that there is a way forward without OpenAI and Anthropic.
Anyone who will put in the time to actually work with a local model will also be aware of the limitation of those things - another point for “make people own their own shit”. Anyone who involves themselves in running a model will also be keenly aware that they are NOT TALKING TO A SENTIENT BEING (I fault Sam Altman and Dario Amodei personally for spreading this bullshit, they should be sued and thrown in jail)
The energy cost of training a model for local use is not even comparable to the huge amount of resources that are used for training something like Mythos - these things don’t scale linear, but exponentially. And it is a one time cost, which means the longer a model is in use the less relevant the training costs become (just as buying a car -the upfront costs are high, but in the long run you will pay more in gas and maintenance than you paid for the car itself). Even if you train one large model on the scale of Mythos, you then have the possibility of distilling those models into smaller ones for local use, specialized for specific areas at a fraction of the cost.
I had a long reply written, but got deleted because my phone ran out of battery; i doubt any of our opinions about the existence and relevance of the discussed downsides would change if this was argued further anyways (and i am too lazy to rewrite the reply)
Your reply doesnt answer the question asked in my comment. You say its a tool, but didnt say what you use it for
(Also, the first sentence seems like a generic reply to any comment anti-ai, it doesnt say anything and its not even relevant in this situation; it only makes your arguments feel worse)
I disagree regarding the “generic reply to Anti-ai”. Discourse nowadays simply ignores locally run models - they do not exist outside of a sphere of technically adept users. Making sure more people know that you don’t have to finance OpenAI, Anthropic, Meta or the Neoclouds should be way more accessible information.
This tech and what is happening currently should not be conflated. If Big Tech used up all global stockpiles of bubble gum because they wanted to create the worlds largest chewing gum bubble, it doesn’t mean that you with your single pack of gum is even comparable to what they are doing.
The only question in your comment i can see is the concern about energy cost while training a model. Those costs are a static amount - not ignoreable, but finite. The longer you use that model, or increasing the amount of people using it, reduces the cost per instance. Use a model long enough and the energy cost of creating in in the first place converges towards zero. It’s another point for local LLMs - nobody needs the aggressive update cycle Anthropic and OpenAI are running, and those can’t stop because the moment they do, investor money will stop flowing.
For what it can be used: well obviously it is a great tool in all areas where you can automate checking if what it did meets a specific criteria. That’s the reason why Mythos has Business entities like Microsoft scrambling like mad just right now - checking to see if Mythos managed to escalate privileges or run remote code is trivial, and Mythos is showing up security issues all over their software, up to the point that MS is scared that even if they patch the really ugly shit, Mythos might still chain multiple low level vulnerabilities to fuck them over, and all that is needed from Joe Shmoe are the words “go hack Onedrive”.
I personally like to use them to calculate stuff where i know it can be calculated because of my experience, but where i do not recall the specific formulas needed. I also use it for creating regular expressions on the fly, teaching myself how to create useful bash scripts (all my knowledge about sed and awk comes from a local LLM), among other things that are not the business of anyone - that’s why I am using a local model in the first place.
I honestly forgot that i also put a question about training cost; by the question i meant when i asked about uses for ai, i didnt realize i made it unclear whether it was a question or not
I dont think a model that you have said is bad for the environment is a good example
Regarding the first sentence, i have said that this seems like a generic reply because i have not said that llms are the worst thing ever (even saying that their uses could offset the downsides) or that they do not have uses (“AI is just a tool” is also an argument i see so often, never with any explanation behind)
For such impressive technology, i doubt the best use of it is a better search engine that you have to fact check or writing help
I have seen someone use a local model to control home assistant
you see to not have read my responses to the other comments.
my comment here was a bit, trying to replicate the average anti ai stance I think exists.
Oh sorry, i didn’t realize that. OK, so looking at the thread, i would say my reply would fit the inane FUD @[email protected] is spouting. Klear, if you wanna reply, i am open for discussion.