Qwen 3.6 27B

Dense agentic coding with thinking control

Chat
Multi-Modal
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.10/ M
Output
$0.40/ M
Cache read
$0.04/ M

About Qwen 3.6 27B

Qwen 3.6 27B is Alibaba's dense open-weight coding model under Apache 2.0: all 27B parameters active on every token, a 262K-token context window, and native multimodal input — text, images, and video. It was the first open model to ship thinking preservation, carrying reasoning traces across conversation turns.

Dense weights make its quality per parameter stand out: on agentic coding it beats MoE models more than 10× its size — 53.5 on SWE-bench Pro against 50.9 for Qwen3.5-397B-A17B — while hybrid linear attention and multi-token prediction keep serving efficient.

With thinking enabled it posts 77.2% on SWE-bench Verified and 59.3 on Terminal-Bench 2.0, with strict function calling, JSON output, and per-request thinking control to trade reasoning depth for speed.

27Bdense parameters

Every parameter is active on every token — consistent quality with no routing variance.

262Ktoken context window

Holds full repositories and long coding sessions in a single context.

1stopen model with thinking preservation

Reasoning traces persist across turns, cutting redundant thinking tokens in agent loops.

Key capabilities

Qwen 3.6 27B is strongest where step quality matters more than raw throughput — repository-level coding, visual front-end work, and multi-turn agents that build on earlier reasoning.

Agentic coding

77.2% on SWE-bench Verified and 59.3% on Terminal-Bench 2.0 with thinking enabled — plans, edits, and verifies real repositories.

Dense quality per parameter

53.5 on SWE-bench Pro — ahead of MoE models more than 10× its size, from 27B dense weights.

Thinking preservation

Chain-of-thought carries across conversation turns, reducing redundant reasoning and KV-cache pressure in multi-turn agents.

Native multimodal input

Screenshots, diagrams, and video frames are first-class inputs — front-end bugs and visual tickets stay in one model.

Where this model fits

Best for
  • Repository-level coding agents and refactors

    Resolves 77.2% of SWE-bench Verified issues with thinking enabled — consistent step quality with no expert-routing variance.

  • Front-end work from screenshots and recordings

    Screenshots, UI recordings, and diagrams are native inputs, so visual bugs and design-to-code tasks stay in one model.

  • Multi-turn agent loops

    Thinking preservation keeps chain-of-thought across turns, cutting redundant reasoning tokens and KV-cache pressure.

  • Multilingual engineering teams

    71.3 on SWE-bench Multilingual — issue-to-patch work on non-English repositories and tickets.

Avoid for
  • Cost-first, high-volume agent fleets

    Every token runs all 27B parameters. When per-token cost dominates, use Qwen 3.6 35B-A3B — near-peer quality at 3B-active pricing.

  • Tight output-token budgets with thinking on

    Thinking mode is verbose — roughly 4× the median output tokens in Artificial Analysis testing. Cap or disable thinking for cost-sensitive replies.

  • Audio input or speech understanding

    Inputs are text, image, and video only. Transcribe audio upstream before sending it to the model.

Estimate your Savings

Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.

Performance and
benchmarks

Entrim performance for Qwen 3.6 27B under production load in the past few days.

Throughput (tokens / sec)

avg. 105tokens/s

Time to first token (TTFT)

avg. 780ms

Public model benchmarks

77.2%SWE-bench Verified

Resolves real GitHub issues end-to-end in agentic coding runs.

53.5SWE-bench Pro

Harder, contamination-resistant software engineering tasks.

59.3Terminal-Bench 2.0

Autonomous work in a real terminal — builds, debugging, ops tasks.

Run Qwen 3.6 27B in minutes

Test this model with $25 free credit. Use Entrim's OpenAI-compatible API to call Qwen 3.6 27B, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "Qwen/Qwen3.6-27B",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

Tool-heavy coding at 3B-active cost

262K
FP8
Vision
Multimodal
Video
Input / M
$0.05
Output / M
$0.25
Cached / M
$0.025

DeepSeek V4 Flash

Version 0731

High-throughput agentic work with 1M context

1M
FP4
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Multimodal long-context work with thinking control

256K
FP8
Vision
Multilingual
Input / M
$0.10
Output / M
$0.30
Cached / M
$0.04

Run Qwen 3.6 27B with the Most Cost-Effective LLM Inference

Use $25 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service