Qwen3.8-Max: A New Bar for Coding and Cowork

(qwen.ai)

211 points | by ai2027 2 hours ago

18 comments

  • toshinoriyagi 1 hour ago
    They've also announced Qwen3.8-27B being released open-weight next week. Qwen3.6-27B is widely regarded as one of the best local models, especially since nothing else comes close to it, that isn't benchmaxxed, without being significantly larger. If 3.8 truly improves upon it that would be awesome.
    • nozzlegear 1 hour ago
      Qwen3.6-35B is my daily driver for AI, and what convinced me to cancel my Claude subscription back in April. The Qwen3.6 line is easily the best local model I've tried, and I've tried a lot. I've got it diligently grinding away on my laptop right now, reviewing and fixing some bugs in my F# code.
      • pettijohn 58 minutes ago
        35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter, so I daily drive that. I wish there was a ~100B MoE with maybe 10B active. It would be super smart and fast!
        • mattnewton 54 minutes ago
          There was a 3.5 122B 10A release -

          https://huggingface.co/Qwen/Qwen3.5-122B-A10B

          • kanemcgrath 15 minutes ago
            I tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now
        • nozzlegear 37 minutes ago
          I've heard 27B is smarter! I tried it some time ago but couldn't get it working with my oMLX. I need to try it again.
      • neumann 1 hour ago
        compared to claude - how 'fast' is it in terms of throughput on your laptop?
        • brucehoult 26 minutes ago
          On my SpacemiT K3 SBC with 32GB RAM (where models run on the eight A100 RISC-V cores with 1024 bit vectors) doing the same task I got 5, 5.8, 6.5 tok/s using gemma-4-26B-A4B-it-QAT-Q4_0.gguf, Qwen3.6-35B-A3B-Q4_K_M.gguf, Qwen3.5-35B-A3B-Q4_K_M.gguf. The corresponding dense models are more in the 2.5-3 tok/s range.

          Kind of slow, but using only 14W of electricity so the Wh per task is twice as good as using my i9-13900 laptop with 4060 GPU.

        • syntaxing 48 minutes ago
          I use it with a strix halo server. 35B runs stupidly fast. 27B is about 700 TPS prefill and 30 TPS token generation. Which interestedly is about what Kimi K3 gives me depending on provider.
          • dionian 38 minutes ago
            what hardware do you use or recommend for this? never heard of it until today.
        • nozzlegear 39 minutes ago
          It's pretty fast, faster than I could type anyway, but not as fast as Claude of course. My oMLX dashboard says I get about 50 tokens per second from the Qwen model I'm running (I host it on my M1 Mac Studio, not on my laptop).
      • ufish235 56 minutes ago
        What laptop?
        • nozzlegear 45 minutes ago
          It's just a MacBook Air with an M4, cheap and nothing special. I host Qwen on my Mac Studio, an M1 with 64gb ram. The model uses around 20-25gb ram depending on what it's doing.
    • icelancer 1 hour ago
      This is what I've been waiting for. We are still using fine-tuned deployments of Qwen3.6-27B with a lot of success but could use a bump in intelligence. Here's hoping.
    • mathieudombrock 26 minutes ago
      Qwen 3.6 27b has been the sweet spot for me in terms of local models. I've had good luck using it with Pi harness. Looking forward to this.
    • XCSme 54 minutes ago
      If they trained it well, and can do computer use, it will be a new era. Companies can keep PCs, put Qwen 3.8 27b on it and get rid of the employees, lol...
  • storus 27 minutes ago
    I think the window for a ban of open weight models is closing fast so let's hope US administration is going to miss it and we get Fable-level models (at least in some aspects) with open weights without infringing any newly introduced law as a long-term local baseline.
  • aliljet 12 minutes ago
    I'm trying and failing to find value running this on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU tells for help, but the problem is that no math works to upgrade this machine even losing a ton vs one of the mega subscriptions with the big model providers...

    How would you all run a potentially better qwen 3.8 model locally?

  • simonw 1 hour ago
    > Today, we are officially releasing Qwen 3.8-Max, the most capable model in the Qwen family to date. This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week.

    I don't understand. That's dated today, but:

    https://twitter.com/alibaba_qwen/status/2078759124914098291

    > Qwen3.8 is launching and going open-weight soon! [...] You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork.

    That was on July 19th. I used it to draw this pelican: https://simonwillison.net/2026/Jul/20/afraid-of-chinese-mode...

    So what are they releasing today?

    • bloomsa 1 hour ago
      July 19th post mentions “Max-Preview” vs. today’s post dropping the “Preview”. Unclear what changed if anything though.. Maybe broader availability or it’s a slightly improved checkpoint
    • Jowsey 1 hour ago
      My understanding is that these "preview" models are usually earlier RL checkpoints, and that "official release" happens when they're happy with the training run?

      I believe they mentioned around the preview announcement that they'd be releasing improvements to capability, which I assume means continued training.

    • simonw 1 hour ago
      ... other comments were right, this is the full qwen3.8-max model, two weeks ago was the qwen3.8-max-preview release.

      Here's a pelican I just got out of the new model. It took 11 minutes and forgot the wheels! https://tools.simonwillison.net/markdown-svg-renderer#url=ht... (scroll to bottom)

      The reasoning trace is pretty great:

      > More additions: basket with fish in it? Cute detail — a fish poking out of a basket on the handlebars! This adds charm and pelican context.

      If the price is $2/$6 that cost me 17 cents: https://www.llm-prices.com/#it=90&ot=29734&ic=2&oc=6

      • ComputerGuru 41 minutes ago
        It gave the svg for the wheels in the reasoning trace then forgot to include them in its final answer. Lol.
        • codedokode 29 minutes ago
          It has a "definition" for wheel in SVG inside <defs>, but did not use it in the picture.
      • pettijohn 56 minutes ago
        Wow, bike geometry is really good! Except for the missing wheels lol
        • applfanboysbgon 53 minutes ago
          Do pelican bikes need wheels? They've got wings, after all... I think Qwen is on to something here.
    • telemaxs 1 hour ago
      they releasing Max.
  • adi2907 1 hour ago
    Once OpenAI and Anthropic are public, every such announcement will become a reliable sell signal
    • gr_norm 28 minutes ago
      Agree, I don't necessarily see a strong argument favoring OpenAI or Anthropic here. In the interest of perspective, can anyone (perhaps playing devil's advocate) give one?

      The open models are now good enough for what I want to do with them, let alone any future improvements. And factoring in efficiency gains, a model in the ~70b range starting to satisfy my needs would completely obviate the need to pay others for inference. This does not seem far-fetched to me, comparing with where open models were at this time last year. What am I missing?

    • ycui7 41 minutes ago
      Can they still go public ? MiniMax M3 Pro is also coming, then DeepSeek-v4-Pro GA, then GLM5.5. There will only be bad news for them in the coming few weeks/months.
      • wmf 14 minutes ago
        Fable 5.1 is coming, then GPT-6...
    • MangoCoffee 1 hour ago
      US AI labs really rub me the wrong way, especially with the doom and scare tactics they use. Both Altman and Dario keep talking about how AI will replace workers and how we should regulate LLMs for national security, Dario’s main point.

      LLMs are useful. We can all see that in agentic coding. But replacing everyone’s job? Hardly. And what’s with the scare tactic of trying to get the US government to ban foreign models?

      LLMs are useful, and dare I say they’re on par with the internet. Making them cheaper and affordable is good for everyone. The fear mongering from Anthropic and OpenAI looks like an attempt to corner the US market into using only US models so they can keep the profits, especially since China has proven that LLMs are a commodity. US AI labs should work on making LLMs cheaper or better harness. Altman and Dario are not trustworthy.

      • EMIRELADERO 39 minutes ago
        You are right to feel that way about the frontier labs, especially Anthropic. From https://stratechery.com/2026/anthropics-safety-superpower/

        > "Anthropic believes that they are the ones who should have final say over how Anthropic is used; given that they think only they should be developing leading edge AI, they by extension think that only they should have final say over AI generally. When you further combine this realization with the company’s pronouncements about AI’s ability to conduct all economic activity, you realize that Anthropic’s leadership effectively wants to have power over everything and everyone."

        • usef- 29 minutes ago
          To be fair, we're simultaneously mocking anthropic for believing in safety so much and also for them thinking they're the only ones that care enough about it. It's true that no one else seems to care as much. Judging by reactions from everyone, all their safety talk is very bad PR.
      • usef- 33 minutes ago
        If there are genuine society risks in a tech I don't want to discourage CEOs from talking about them. I feel like we've spent decades talking about how evil chemical companies (etc.) were about covering up issues in the 20th century. But yes, that's different to being a reason to ban external models.
      • dmix 44 minutes ago
        Sam drank the "superintelligence" kool aid early on and said 30-40% of jobs could be impacted by AI, but recently admitted he was wrong

        > “My scorecard, at the highest level, would be we’ve been roughly right on technological predictions and pretty wrong on the social and economic implications” https://www.cxtoday.com/ai-automation-in-cx/sam-altman-softe...

        I agree re: Dario quietly pushing for government control. He also said LLMs would replace a lot of entry-level information jobs, doubling the unemployment rate from 4-5% to 10%.

        Yale did a study recently showing little impact on employment in high-AI exposed jobs https://budgetlab.yale.edu/research/ai-probably-not-yet-reas...

        • conception 21 minutes ago
          I imagine it will be a long tail. Most companies won’t fire people for AI but probably won’t immediately replace a person that leaves, if at all.
    • int32_64 53 minutes ago
      It's not so simple, if such a headline can get them closer to the regulatory capture they want to lock in American businesses and forbid them from using Chinese AI.
    • _jayhack_ 1 hour ago
      Only the ones that beat expectations
  • boredatoms 55 minutes ago
    3.8 27b is the real news here
  • kopirgan 42 minutes ago
    Can a model be stripped off anything not relevant to coding and get a lot lighter? Or is that impossible?

    Just like we have professors with specialisation wondering if AI models can also be so.

    • htrp 30 minutes ago
      You can... but the trick is to do so without killing performance. Turns out a lot of random things help make coding performance good.
    • ReptileMan 22 minutes ago
      I guess it can but it will be useless. After all the model superpower is awareness and ability to guess and infer some stuff. Right now a model saves you time not only by coding faster, but that it can figure out some stuff about the shape of the data and its purpose.

      If you throw general purpose model at a codebase - it will look at the table and data logical connections beyond what is explicitly declared. It will figure out on its own that Salaries should be displayed on SalariesTable.php and it will "know" that your prices should include vat and so on.

      A human knows that VAT and price go together and are related, full size LLM does too, stripped one - doesn't.

    • sp1982 4 minutes ago
      [dead]
  • ddxv 1 hour ago
    It seems this is the only mention of cost?

    > Qwen3.8-Max comes with the official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:

    > xhigh (default): for complex tasks demanding thorough analysis

    > medium: balancing accuracy and speed

    > low: efficient reasoning optimizing for speed and cost

    I hope this is significantly cheaper. I've been loving Deepseek for it's nearly free usage costs, hard to justify switching from cents per day.

  • wxw 1 hour ago
    > This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week.

    Nice!

  • ComputerGuru 37 minutes ago
    Does the page actually load for anyone? I get stupid spa skeleton spinners.
  • jofzar 1 hour ago
    Lmao I love their video with the idea that people will be able to do their hobbies while ai does their job.

    Surely Alibaba is leading by example here by reducing work hours per week while keeping pay the same right? Right?

    • SyneRyder 59 minutes ago
      > I love their video with the idea that people will be able to do their hobbies while ai does their job...

      Are you not already experiencing this? I think this is fairly common for people using AI now, though the time may not always go into hobbies or sports. It's common for me to setup Claude with an hour+ task while I catch up on housework, or while I'm getting ready in the morning.

      In the last couple of weeks I've unfortunately had multiple family illnesses - it has been helpful to have Claude keep up with much of my product development programming work while I visit my mother in hospital and check on my father's recovery. I'm able to give more time to family without worrying that business progress isn't keeping up. The overnight Claude sessions while I'm asleep have been particularly helpful.

      • jofzar 41 minutes ago
        No I haven't had time to spend my afternoon rock climbing while ai generates documentation.

        It's infinite work, I just did more work while codex was doing it's thing in the background.

    • mlmonkey 1 hour ago
      That's the thing. Wny are companies like OpenAI/Anthropic/Alibaba/Kimi/Deepseek still hiring SWEs if their models have become so good?
      • rrix2 51 minutes ago
      • BetterThanSober 1 hour ago
        The models are good even by skeptics standard, it's just that evangelists are overselling the capabilities. If you understand the limits of LLMs not using them as a business is shooting yourself in the foot.

        However, they are not at the point where they can effectively train themselves, nor did they are capable of researching their own method of learning. SWEs in mid-corps on my country are right now relegated to reviews and sanity check, basically babysitting the LLMs and making sure they're not spouting nonsense. If you think about it, that's basically QA and can also be delegated to another AI. If Bun's rust rewrite that they tout as fully LLM-led can pass the test of time in a year or so I think that's it.

        I believe all that is now constrained by compute and capital, not tech.

      • wmf 1 hour ago
        There's infinite work to be done, so higher productivity makes people worth more. (Obviously this doesn't apply if AI can do everything but we're not there yet.)
      • cute_boi 1 hour ago
        The world never runs out of problem. There is so much work to do.
      • Mythorian 1 hour ago
        I mean its pretty obvious right? This models are not flawless and sometimes reach stupid conclusions so there needs to be some one who watches it. Thought i must say u are right. Every one of them pretends that this new model is gonna finally take ur jobs lol
  • esafak 17 minutes ago
    Does anyone know how token- and reasoning efficient it is? The charts don't show how many tokens were used in any benchmark.
    • wmf 11 minutes ago
      The imminent third-party benchmarks will cover that.
  • luciana1u 1 hour ago
    the benchmark I trust most is whether the model can explain its own pricing page without getting confused
  • BeriV2 1 hour ago
    We will eventually need a self evolution benchmark to see where these large models can create recursive solutions that improve
  • whateveracct 23 minutes ago
    ah so they distilled fable and sol, eh?
  • TacticalCoder 51 minutes ago
    > In this case, Qwen3.8-Max was asked to create the oh-my-cli project from scratch and, over a 10+ day long-horizon autonomous coding run, build a self-evolving harness.

    They don't explain how successful that went but it's a bit hilarious seen that an Anthropic dev explained that it's been 15 days Claude was hard at work --with nothing to show yet-- trying to rewrite itself in another language.

    "You rewrite Claude Code, we rewrite oh-my-pi."

    "You're nowhere after 15 days, we do it in 10."

    Sure, it's apples to oranges and all that. But part of me thinks they know fully well what they did there.

  • choppaface 1 hour ago
    “self-evolves through feedback loops”

    Does this mean they distilled Claude? Sounds like what Claude Code will often do.

    • charcircuit 1 hour ago
      It's meaningless. Models have always been able to do this and this capability is strengthened during RL since being able to explore the solution space to figure something out will give it a reward.

      What is important is how long it can go without requiring human intervention. Not just that it's possible to run on its own for a time.

    • Art9681 1 hour ago
      Of course they did.
  • VladVladikoff 1 hour ago
    Are these latest Qwen models still open weights or has Qwen moved away from that?
    • a2dam 1 hour ago
      The second sentence of the page: "This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week."
      • VladVladikoff 1 hour ago
        Page won’t load for me it’s just grey bars fading back and forth forever.
        • Larrikin 59 minutes ago
          You can always wait until the page loads before posting your thoughts on the Internet