

It’s kind of insane that integrated wikis with citations/documentation aren’t the centerpiece of (Reddit) communities.
Instead, we got… Discord?


It’s kind of insane that integrated wikis with citations/documentation aren’t the centerpiece of (Reddit) communities.
Instead, we got… Discord?


It’s a bit misleading.
Qwen 27B has way less “world knowledge” than GPT-5. Ask it random trivia without internet search access, and GPT would know waaay more.
This is generally true of small vs large models.
…But honestly, Qwen 27B is better at tool use or agentic stuff. It’s hyper optimized for just that and coding assistance, basically.
This is often true of old vs new. Most newer models have hyper focused on agents/coding, often to the detriment of other use cases.
Quantization for practically running Qwen 27B also has an impact. A off-the-shelf Q4_K_M is not the same as the unquantized weights in real-world use, or even an “optimized” quantization like a custom exl3.


You could though. You could run a knockoff gambling site that doesn’t actually cost (or give) you any money.
The irony of gambling is that it’s the opposite of other media; the addictive part is paying for it.


In other words, the average American adult placed roughly $1,000 in legal bets on sports last year. Handle is gross throughput, not consumer expenditure: over 90% of what is wagered gets returned to bettors in the form of winnings. That $1,000 in bets translates to roughly $100 in average losses per adult.
But it does point out the vast majority of the losses are absorbed by 5% of gamblers.


Be aware, q4_0 KV quantization really borks models.
But the K is far more sensitive than the V. Try q5_1 for the K while leaving the V at q4_0 or q4_1; vram usage will be almost the same, but it should work dramatically better.
As for throttling, try disabling turbo on the 8700.
It doesn’t actually need turbo clocks for these models. The t/s loss I get from doing that on my rig is very modest.
And like others suggested, try the QAT release. You might try the ik_llama.cpp for while you’re at it, at it should be faster with MoEs like this.


…What backend are you running?
With what settings?
Pretty much every one should allocate the kv cache up-front; if it’s gonna OOM, it will do it when the backend spins up.
The only one I know of that still doesn’t is plain huggingface transformers, but no one should be using that for actual inference. I’m unsure of MLX, though.
It’s also odd that 8K is your max context on a 24GB card, if that’s indeed your situation. And that rope scaling is even a consideration; that’s an artifact of the distant past. 64K+ native should be quite doable. What model are you running?


Also, AI or not is really irrelevant for this case.
They said as much:
Saflor uses AI tools and does not believe that AI itself is the problem. Goldman pointed out that his arguments against Memes Apps would be largely the same even without the AI aspect. But Saflor considers Memes Apps’ platforms to be examples of irresponsible AI products, where the operators problematically advertise that you can “fire your ad agency” and replace all creative work with a meme generator.
Which sounds to me like effectively abolishing the copyright system for anyone who can’t afford paying lawyers to hunt down every little infringement, only leaving it up for wealthy corporations.
And yeah, that would be a really terrible precedent…


Well, I tried to have a touchy discussion about China with Gemini 3.1 Pro, in their UI. Specifically it was about a very questionable cultural conformance law. It straight up refused.
GLM 5.1 went right into it, going out to read the original Mandarin text itself, and straight up criticized the law as “incompatible with the internationally accepted definitions of human rights.”
Xiaomi MiMo got censored from their portal, but run the weights locally, and it is a completely different animal. It knows all about Tiananmen Square too, if you coax it out with raw completion syntax instead of chat formatting.
This has been my experience with Chinese models going back to the original Yi 34B. Those Chinese devs like to have their cake and eat it; they leave the models pretty open and uncensored, but then they tighten the screw in the user facing portals, to give the aura of govt compliance I suppose.
…What I’m saying is this person has no idea what they are talking about.
They are interacting with them over API (or web portals), where they’re censored to heck.
They are interacting with them via chat APIs, not raw syntax (or at least a clever system prompt) to test what’s deep in their weights and get past the random refusals.
They are talking about censorship mostly, which is distinct from a models openness. Yes, the training data of basically all Chinese models is closed, and they probably filter some politically sensitive stuff out. But some prominent ones (like Nvidia Nemotron) have open datasets.
And the author is leaving out the elephant in the room: “frontier lab” models are the worst of everything. They’re heavily censored in the weights, heavily censored in their front ends, and it’s impossible to even access raw completion syntax or logprobs to dig into what they’re really thinking. There’s zero public information on their architecture or training regime. They “closed” on a whole different level than any Chinese model.


To be clear, I’m not a fan of LW either. I’ve got communities blocked there, and I’m still considering leaving.


Because when I bring this up, I get replies calling me a fascist sympathizer, and the whole thread barrels towards Godwin’s law.
You can indulge in that all you want, again, if you wish. I’m not interested in that rabbit hole.


I’m just citing my own personal observations interacting with the db0 community. The calls for violance and death were not subtle.
It has little to do with whatever happened to tesseract; I noticed this trend before that, but all that controversy just bubbled it to the top of my conscious (as it seems to have with that startrek community).
Curiously, you’re doing it with the exact same language that AdmiralPatrick used, where the exact nature of the targets of the “instance’s views” are.
If you’re accusing me of being AdmiralPatrick or related to them, I don’t really know how to respond to that. But I’m not, and I’ve been on Lemmy over 2 years so far.


Well, the current state is that there are a couple of “best” models, but literally hundreds of independent providers serving them. As an example, one can get GLM 5.1 from its trainer, or one can get it from DigitalOcean, or Baidu, or SiliconFlow ASICs, or get it at very high speed from Cerebras ASICs, or AMD providers, or finetune it from a number of services, or rent the self hosting…
The companies aren’t keeping the models to themselves, and that blows the marketplace open to a boatload of competitors.


+1.
They were both trying to do all at once, albeit with some distinctions.
And an addendum: transformers LLMs are not leading to AGI anyway. That was a lie.


I ranted about this in another thread some, but I’m with startrek.website here.
I almost switched to db0 when I first found Lemmy.world and joined it; you guys seemed, ideologically, closest to me. I loved the vibe.
Communities about Anarchism, Generative AI, Copylefts, Neurodivergence, Filesharing, and Free Software. (And Math!) Follow the Anarchist Code of Conduct and honor the Disengage Rule. Don’t be shitty to each other. Keep it SFW. Obey the spirit of The Golden Rules. Fuck around and find out.
This is me. Every word. I love this.
…But holy hell.
The bar for calling for explicit violence is basically zero, as long as the target fits the instance’s views. And I’ve watched the admin come in and personally validate that policy in the threads in question.
I wouldn’t want to touch this instance, either.
Again, I’m otherwise sympathetic to everything about this instance, yet I think I’d vote for Lemmy.world to defederate from db0 if it ever comes up.
And I’m not here to debate how I’m apparently not antifascist enough, or shoot down accusations that I’m a pacifist, or whatever.
I’m sick of that.
And I’m clearly not moving the needle here. I’m not trying to; do what you do. I just thought I’d provide a perspective of someone who was otherwise quite attracted to this community.


They don’t have to host it themselves. They could use a number of providers for the same model, and basically keep doing whatever they were doing with OpenAI/Anthropic via the exact same APIs.


Bonsai? Or whatever it’s called? It’s a con, so far; it’s not better than smaller models quantized to 3-4 bits.
I love, love the idea of bitnet, but it only seems to work with models trained from scratch, which no one has done at scale yet.


You want this one:
https://huggingface.co/turboderp/Qwen3.6-27B-exl3_3.30bpw/tree/main
Or maybe the 3.5bpw one if you don’t mind less context, or 3bpw if you need more:
https://huggingface.co/turboderp/Qwen3.6-27B-exl3
For faster inference at the cost of a little more VRAM usage:
https://huggingface.co/turboderp/Qwen3.6-27B-DFlash-exl3
And you run those in:
https://github.com/theroyallab/tabbyAPI
And FYI, if you have 64GB of RAM or more, you might consider hybrid inference instead.


I also forgot to emphasize this, but Xiaomi’s plan, in my opinion, is undiscovered fruit.
GLM had a similar coding plan, but once it got in the news and popular, it got WAY more expensive and limited. I’m grandfathered into 6 more months of a GLM plan you literally cannot buy now.
And I think Xiaomi is in the same situation GLM was 6+ months ago. It’s a fantastic model series, but unlike Kimi/GLM no one knows about it yet, which is how it’s still $60 for a year.


+1 for Hermes.
If you have a newer Nvidia GPU, you can run Qwen 27B via exllamav3 and get good quality/speed in 16GB. And it’s worth the trouble, as 27B is an amazing model.
If it’s AMD, yeah, a3B is a good bet, depending on how much spare CPU RAM you have.
I use an exl3, with 4 bit MLPs but higher bit depth attention layers. And I force some custom sampling so I can lower the temperature a bit while keeping it out of loops.
This won’t work in LM Studio though. You have to run such a thing in TabbyAPI or some other backend that supports exllamav3.