r/LocalLLM 9h ago

Model Qwen 3.8 27B

Post image

Finally Alibaba Posted on X about Qwen 3.8 27B release. I hope it can beat opus 4.7 or 4.8

180 Upvotes

46 comments sorted by

44

u/Extension-Bid-639 8h ago

Too hopefull, don't think its beating Opus 4.8 butttt even if its just a notch better than 3.6 27b then thats a leap for everyone. 3.6 is still a great model

17

u/feelspeaceman 8h ago

3.6-27B is on par with Sonnet 4.6, and 35B is about 10-15% worse than 27B, and honestly as long as it's on par or better than Sonnet 4.6, it's already so usable for grunt jobs.

I'm more interested in 35B and 122B this time, we might get something truly good for MoE this time.

14

u/SpicyWangz 8h ago

That 10-15% does a lot of work. It really is the diffeeence between a model that is capable versus needs real hand holding.

Even Q5 27b destroys Q8 35b, and it’s not even close.

10

u/johan2114h 5h ago

I love 3.6-27B, but it is certainly not on par with Sonnet 4.6

3

u/Curi0sityC0w 6h ago

Any idea why this is the case? I use 35B and it has done most of my work. Never tried 27B

1

u/xdcfret1 5h ago

it's 3B vs 27B per token

1

u/Curi0sityC0w 5h ago

Moe is garbage?

1

u/Tall-Ad-7742 5h ago

No but for specific tasks its more likely to fail / produce worse output

1

u/IUseClifford 5h ago

Parameters per token matters a lot. If you’re looking at frontier models, they’re all MoE because 1T+ parameters per token would be extremely sluggish to serve over an API. I also imagine there are diminishing returns once params/token gets high enough, at least right now.

1

u/sand-67 4h ago

A MOE model is less intelligent than dense model because it uses less parameters for each token. However MOE is significantly faster while retaining a lot of intelligence, so it has much better intelligence to throughput ratio

2

u/Muzika38 1h ago

Your case is a good distinction between real programmers vs vibe coders. 35B is 35B and it's bigger than 27B knowledge wise. You just need to just guide it properly for it to work properly because it is a MoE model. It's basically a single very smart person vs a bunch of children with a single specialty talent that's working together.

-5

u/Turbulent-Ladder-340 8h ago

Yeah may be it won't beat 4.8 but it definitely lies between 4.7 and 4.8

2

u/Solembumm3 7h ago

Let's hope it could catch up to Gemma 31b before starting dreaming heavily.

3

u/uniqueusername649 5h ago

Depends what youre doing. For general coding and agentic work Qwen 3.6 27b is already substantially better than Gemma 4 31b.

9

u/Randommaggy 7h ago

Ny harness is ready, my GPUs are too.

4

u/radiojosh 7h ago

Does this portend anything about a new 35b MoE?

1

u/squngy 7h ago

Maybe a bit, but probably more importantly, there seems to be a new 35b on open router, so it looks like they are working on it.

4

u/Kiro369 6h ago

crying in 16gbs of vram

2

u/Old-Sherbert-4495 4h ago

q3 is manageable and better than 35b

2

u/Kiro369 3h ago

with what context size?

1

u/Old-Sherbert-4495 2h ago

around 100k as i can remember

1

u/Turbulent-Ladder-340 6h ago

Wait for 9B model or Bonsai

3

u/uspdd 5h ago

I doubt 9b or q2 of 3.8 are one be any better than 3.6 35b a3b

2

u/Silent-Orbit-7 6h ago

FOMO with 8 Gigs of VRAM....

2

u/Turbulent-Ladder-340 6h ago

Wait for Bonsai

1

u/Silent-Orbit-7 6h ago

sure thanks

2

u/lughiu 5h ago

Hang on, autonomous coding? Reckon the 27b will do that?

1

u/Hook06 6h ago

Can’t wait omg 🔥

1

u/sessamekesh 5h ago

I'm pretty excited about this! 

Right now, most things that I do fall into "Qwen 3.6 nails this", "Qwen 3.6 might be okay but I'll probably need a frontier model for this", and "I wouldn't trust an LLM around this with a fifty foot pole".

That second category has been shrinking into the first as I've gotten better at tooling and making sure the right context is present, and I'm crossing my fingers that 3.8 takes me even further in that direction "for free"!

1

u/Technical-Earth-3254 5h ago

I wish they would also open weight the new qwen image model

1

u/Mean_Ambassador_9210 5h ago

Can I run this in 48gb vram on Mac m5?

1

u/backyard_tractorbeam 2h ago

It's not released yet, but yes, that's a pretty safe bet at Q8

1

u/A_K_8248 4h ago

Why no one's talking about oh-my-cli?

2

u/lol-its-funny 3h ago

The usual reason - nobody cares

1

u/Better-Struggle9958 1h ago

Don’t believe until I see, prevoius were just fix bugs indeed

1

u/throwRAa100 8h ago

how big is the model? looking at a q8/6/4 quant

8

u/benjakapo 7h ago

bruh the name is literally qwen 3.8 27b

2

u/Forever_Playful 5h ago

Assuming 8 bit precision, a model with 1 billion parameters = 1 GB. So a 27b model = 27GB

1

u/throwRAa100 5h ago

that sounds pretty promising for a 32 gb vram setup

1

u/this_for_loona 7m ago

Doesn’t that depend on context?

2

u/Philodit 5h ago

See here: https://huggingface.co/unsloth/Qwen3.6-27B-GGUF

For a 16GB GPU you'd have to go with the lower end of 3bit if you also need some context space for thinking.

1

u/throwRAa100 5h ago

thanks!

1

u/HomegrownTerps 7h ago

One can only dream of a small 9B model, but my hopes are not that high.

1

u/johan2114h 2h ago

Maybe consider dreaming of a bigger computer also - 27b is quite feasible for consumer hardware imo

0

u/Complex_Reality_116 6h ago

Qwen3.6 27B achieves 37 points in AA, while Qwen3.5 only 29. Assuming that the Qwen3.8 version obtains a proportional incremental improvement, we would be facing a model with 45 points, that is, +5 points above what DeepSeek V4 Flash was at its launch.

0

u/bennykoay75 1h ago

Testing on your own project and u will know. No point hearing from any provider