With the capture of HuggingFace (#1561625) and the continuous seeking of regulatory capture by the LLM circlejerk industry (#1569731), we may need to protect the interests of the people against these actors. Since it is election season in the US and AI panic has been made an agenda item™ (#1567663), the probability of outcomes that are restrictive for the freedom of the individual, small lab and small startup, sovereign or slave.
This morning, I've decided that there is a possible way to beat this capture, in pretty much the same way we have fought capture in the past. But I need input from the stackers, at least the ones that actually run LLMs locally, to validate my idea before I waste everyone's time.
Questions:
- What models do you use now, and which ones have you used in the past?
- Do you use mostly vanilla models, or finetunes?
- Can you list the ones you've downloaded from huggingface and/or modelscope?
Thank you.
Deepseek v4 pro/flash have been my bread and butter.
A dash of Kimi and GLM when something calls for it.
What kind of quant are you running these at? Or are you going to put me in awe and explain that you actually own that B300 rack that I want but absolutely cannot afford for the foreseeable future?
Not sure what you mean by quant in this context, but I don't run local LLM much.
The best I can run on my hardware is deepseek coder or Qwen 3.8 Q4... And I just realized quant might mean quantization here lol. So Q4.
When I finally get around to dropping a bunch of money and make my setup, it would be for more than LLM. Nodes, torrent seedbox, etc. so it's not a rabbit hole I've gone deep into yet.
I am keeping an eye on the open weights and huggingface scene though. Apparently there's a torrent based clone out (pirate face) to be censorship resistant of they ever crack down on models but I'm not sure how reputed it is atm. Just glad to see torrent still being used.
Yeah the problem with the recent flash models is that they're huge so you got me puzzled there for a moment. I wish I were able to run Kimi or even GLM-5.2 locally. I'd deal with living in the noise if that were an option available to me.
The torrent based clone is largely unseeded. And it removes the git metadata in the torrent. So needs some work. But by itself the torrent idea is already cool, if anyone seeds it. I've put down a todo item to play a bit with it and see if I can integrate it with git+LFS or git-annex. The fallback URL isn't decentralized so you can't depend on that if it ever comes to legislative bullcrap (I really don't know how deep the fear factor will be milked, it's hard to say), but we can do something with it in this state too and harden it further.
If you had asked me last week, I'd have said only proprietary models...
But this weekend, ended up running Ollama with qwen2.5:14b on my entry level Macbook Pro, just to get familiar with what it is to run things locally and privately. Budget for new hardware will free up in coming months, so I hope to choose the right specs based on this early experience, so that I can run something more powerful.
So probably can't give you very useful feedback or input for now.
qwen2.5:14bis a valid answer! I wonder one thing tho: why Qwen 2.5? It's from September 2024.Wanted something really lightweight to play around with for a 16gb machine.
Anything heavier you recommend?
Turns out it was much easier to set up than I had anticipated, so eager to try a few other ones.
And yes, I can feel a world of difference using such an old model as I got used to the proprietary flagship models. Feels almost retro, even though it's only been 2 years.
Ollama has these in their library:
gemma4:e4b(gemma4:12bcould work but e4b works better on my 16GB macbook than any 12b ones)qwen3.5:9bNeat. Will report back after trying.
My list:
Chat
Non-Chat
TIL Muse Glimmer - I should pay more attention to the releases again because I've completely missed this and it's "already" a month old!
I liked GLM-4.7-Flash too but I wasn't able to find a quant I really liked on my Mac, so I threw all that out in favor of ppq. One of my knockoff-openclaws is still configured with that, but it's been switched off for months.
Muse-Glimmer is a good chatter. Good for things like talking about specs. Pros / cons of this vs that...etc.
In fact, on the public api's - I find Muse Spark (the big brother) to be my favorite spec / feature chatbot. This is surprising because metas recent models have sucked, but not so surprising considering chat is what put lama on the scene....
Funny related point, but Cursor has moved to heavy grok promotion given the acquisition and while grok 4.6 code seems fine. I don't honestly like chatting with it. It has the obvious snark that its known for, but a bigger issue is that I think since its trained so heavily on "twitter speak" it can be painful to hand in-depth tech discussions with. Sort jumps around topic to topic in disjointed way or something. Can't put my finger on it, but I do recognize it....
https://pirateface.co/
Saw this earlier today.
I literally just started running an agent the other day. Currently running Qwen 3.8 and have one of the latest Deepseek models that I'll try out soon.
With all the negativity around open source AI I figure I'd better download some before they get banned. Not that I really think anyone can stop it, they just might make it more difficult.
That's interesting, thank you. Days like this, I feel lucky that I am not the only one seeing the issue w/ HF.
Straight torrents are almost as awesome as
git-annex(which was literally built for this before LFS was built into GitHub), and can be a reasonable compromise. It still relies on DNS, but I'll see if I can go seed some.Add: top seeded is very poor haha
It's probably less than a week old.
Yeah. Ship fast. I guess they don't have the funding to boostrap the torrents yet.
I've made a reminder to play a bit with one of my dev servers that has a TB free disk and see if I can bridge the git+LFS from HF with the torrents/magnet from PF. To not lose the metadata from git in the torrenting process because after checking
MiniLM-L6seems that it does not include the git data.Not sure this will help you but here is my answer:
Unfortunately most of the time I used it my computer ended up shutting down unexpectedly because of the temperature of the motherboard so I gave up.
It has been now a year I have been waiting for an economic crisis hopping prices would drop to buy a GPU or a MacBook. So I have been mainly using models on Kagi in the meantime (Quick, Qwen, models used for translation, etc).
Here's my list. Almost all are gguf, and this is what I currently have on disk across 2 machines. There are things I deleted (mostly: reapers/dolphins) that I don't remember and also was mostly disappointed by.
Chat models:
Qwen/Qwen3.8-27B<--- in useQwen/Qwen3.6-27BQwen/Qwen3.5-27BQwen/Qwen3.5-4BQwen/Qwen3-Coder-Next-Q4_K_MOpenGVLab/InternVL3_5-8BOpenGVLab/InternVL3_5-4B-Pretrainedgoogle/gemma-4-31B-it-qat-q4_0-gguf<--- in useggml-org/gemma-4-E4B-it-GGUFmradermacher/gemma-4-E2B-GGUFgoogle/gemma-3-4b-it-qat-q4_0-ggufmeta-llama/Meta-Llama-3-8B-Instructmeta-llama/Llama-3.1-8B-Instructmeta-llama/Llama-3.2-3Bmxmcc/xLAM-2-32b-fc-r-mlx-8BitMenlo/Jan-nano-ggufjanhq/Jan-v3-4B-base-instruct-ggufjanhq/Jan-v3.5-4B-gguf<--- need to eval stillbartowski/dolphin-2.9.4-llama3.1-8b-GGUFbartowski/nvidia_Orchestrator-8B-GGUFNon-chat models:
opendatalab/MinerU2.5-2509-1.2B<--- in usehandy-computer/nemotron-3.5-asr-streaming-0.6b-ggufhandy-computer/parakeet-unified-en-0.6b-gguf<--- in useds4sd/SmolDocling-256M-preview-mlx-bf16Qwen/Qwen3-Embedding-4B<--- in usemicrosoft/VibeVoice-1.5Bmlx-community/whisper-large-v3-mlxmlx-community/whisper-large-v3-turbo<--- in usesentence-transformers/all-MiniLM-L6-v2<--- in useI just started messing with qwen2.5-coder:7b
deepseek v4.1 flash
No idea; I've never used those models, so I can't help you.
As of maybe 12 months ago, I downloaded anything that might conceivably work well on 6GB of VRAM. Experimenting with a lot / discovery is HF's value to me.
(through Ollama, but I believe they use Hugging Face as CDN)
For actual use, GLM/Flash through providers... may investigate newest DeepSeek. I'm also hopeful for the next Nemotron.
Got the itch to experiment with edge device models (like Cactus Needle, Lumma .6B) so those would be the next most likely downloads
I think Jensen buying HF is the best of all probable outcomes, incentives aligned as they ship chips and the rift with Dario. But, it's ultimately the kind of resource that would be better if decentralized. Was actually just ruminating on our Lightning Video infrastructures applicability to weights and clanker news (using webtorrent to decentralize CDN and nostr for cards)