# Run Mellum2.1 locally and check its first response

> JetBrains’ current GGUF repository documents a llama.cpp path for serving Mellum2.1 locally. Install the runtime, request a completion, and check it before trying an agent.

By BIG CHANGE Editorial

Published: 2026-10-09T17:10:04.232Z
Updated: 2026-10-09T17:10:04.232Z
Canonical: https://bigchange.ai/blog/run-jetbrains-mellum21-locally-llama-cpp

![An unbranded desktop computer tower with a small orange indicator stands on a warm ivory floor.](https://bigchange.ai/api/media/file/mellum21-local-inference-hero-v1.png)
Conceptual illustration of a computer that can host a local model server; no Mellum2.1 installation or test is depicted. AI-generated illustration by BIG CHANGE.

JetBrains’ Mellum2.1 is available as a 12-billion-parameter model with 2.5 billion active parameters. The BF16 repository and a separate GGUF repository are both live on Hugging Face. For a first local run, the GGUF repository documents a llama.cpp route: download a quantized file on demand, start a local server, send it a prompt, and inspect the answer before giving any coding agent access to a repository.

This is a documentation-based setup guide. BIG CHANGE did not install Mellum2.1 or run the commands below. The success check is a local completion returned by the server; it does not establish coding quality, agent reliability, or performance on your hardware.

### What you need

- A Windows, macOS, or Linux computer with enough memory and storage for the selected model and its runtime. JetBrains’ card lists the original model as BF16 with a 131,072-token context. The GGUF repository lists its recommended Q4\_K\_M file at 8.1 GB. That is the file size, not a complete estimate of runtime memory: the runtime, context and other processes need additional resources. JetBrains does not publish a minimum memory requirement.
- An internet connection for the initial software/model download. The inference request itself can go to the local server.
- llama.cpp and a terminal. The repository documents `winget install llama.cpp` for Windows and a `llama serve` command for local serving.

The model is released under Apache 2.0. The weights have no listed usage fee; hardware, electricity, storage and any infrastructure you choose to rent have costs that JetBrains does not quantify. The BF16 model card currently says no inference provider serves that repository. The GGUF repository is a separate quantized artifact with its own llama.cpp quickstart.

### Step 1: Install llama.cpp and start the local server

In PowerShell on Windows, install llama.cpp using the [official Mellum2.1 GGUF quickstart](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF):

```powershell
winget install llama.cpp
```

Open a new terminal if the `llama` command is not yet on your PATH. Start the recommended Q4\_K\_M build and explicitly bind it to the local machine on port 8080:

```powershell
llama serve -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M --host 127.0.0.1 --port 8080
```

The GGUF repository documents this model ID and quantization for `llama serve`; it also documents the Windows install route. The [llama.cpp server reference](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) documents `--host` and `--port`; the repository’s example endpoint is `http://localhost:8080/v1`. On macOS and Linux, the same GGUF repository documents installing llama.cpp with `curl -LsSf https://llama.app/install.sh | sh`, followed by the same serve command.

The process must download the model before its first response. The Q4\_K\_M file is listed as 8.1 GB. Keep the server bound to your own machine for this test; do not expose it to a network or put credentials into the prompt. The repository also lists a smaller 7.0 GB MXFP4\_MOE file and larger Q6\_K, Q8\_0 and BF16 variants. Quantization changes the model artifact; the listed file size alone does not tell you whether a particular computer can serve it at a useful context length or speed.

If the command fails to start, first check the exact model ID, available disk space, the llama.cpp version and the full error text. A memory allocation failure is a reason to stop and check runtime options and context settings against the current llama.cpp documentation; it is not evidence that the model is defective. Do not assume the published 131,072-token context will fit on your machine.

### Step 2: Send a smoke-test prompt

Leave the server running. In a second PowerShell window, send a short request to its local API:

```powershell
$body = @{
  model = "JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF"
  messages = @(
    @{ role = "user"; content = "Reply with exactly: MELLUM21-LOCAL-OK" }
  )
  # Author-selected budget for this short smoke test; not a JetBrains recommendation.
  max_tokens = 512
  temperature = 0.6
  top_p = 0.95
  top_k = 20
} | ConvertTo-Json -Depth 5

$response = Invoke-RestMethod -Uri "http://localhost:8080/v1/chat/completions" -Method Post -ContentType "application/json" -Body $body

$choice = $response.choices[0]
[pscustomobject]@{
  finish_reason = $choice.finish_reason
  has_reasoning_content = -not [string]::IsNullOrWhiteSpace($choice.message.reasoning_content)
  content = $choice.message.content
}
```

The model ID, local endpoint and sampling values follow the repository’s API example. `max_tokens = 512` is a bounded value chosen here for this short check, not a model-card recommendation. Mellum2.1 is a thinking model; the GGUF card says it emits reasoning in `<think>...</think>` blocks. The code reports whether the separate `reasoning_content` field is nonempty without printing that field. Depending on the runtime’s reasoning format, `content` can still include `<think>` text. The prompt is a smoke test, not a model-quality benchmark.

The basic check passes only if the request returns a completion instead of a connection or server error, `finish_reason` is not `length`, and the final `content` contains `MELLUM21-LOCAL-OK`. An HTTP success alone is not enough: a thinking model can spend a small output budget before producing the requested final text. If the output is empty, the marker is missing, or `finish_reason` is `length`, treat the check as inconclusive; raise the bounded output budget and try again rather than diagnosing a model failure. If PowerShell reports a connection failure, confirm the first terminal still shows a running server and that the request uses port 8080. If the server returns an error, save the exact message and resolve that before testing a coding agent. A successful response confirms that this local inference path can answer one request; it does not confirm the model can safely edit code.

### Step 3: Check it before connecting an agent

Keep the first evaluation separate from a live project. Use a disposable repository with a known, small task, and record the model artifact and quantization, llama.cpp version, operating system, context setting, prompt, output, elapsed time and resource use. Ask for a bounded change, inspect the proposed diff, and run the repository’s existing tests yourself. Repeat the same task with your current baseline if you want a comparison. One prompt or one successful test run is not a benchmark.

A model server returns text and, depending on the runtime and request format, may support structured tool-call workflows. It does not by itself decide what tools an agent can use or safely execute them. The surrounding agent controls repository access, shell commands, file edits and test execution. Start with read-only or tightly scoped access, require approval for writes and commands, review every diff, and keep secrets out of the test repository and prompts. Stop if the agent tries to exceed the task, alters unrelated files, or cannot explain a failing test.

For private use, confirm both where inference runs and how your agent is configured. A local model process can keep prompts on your machine when requests stay on the local endpoint, but an agent may still call external services for other features. Check network access, logging, telemetry and tool permissions in that application before using private code or data.

### Step 4: What the published evidence does and does not show

JetBrains describes Mellum2.1 as a 12B mixture-of-experts model with 2.5B active parameters, BF16 precision and a 131,072-token context. Its model card says the release is Apache 2.0 and describes post-training that included reinforcement learning in sandboxed software environments. JetBrains also publishes benchmark results, including agentic coding evaluations. Those scores are JetBrains’ self-reported results; they are not measurements from BIG CHANGE, and they do not predict your hardware’s throughput or your project’s outcomes.

One release detail changed between publication surfaces. JetBrains’ October 8 launch post said GGUF builds were “coming soon.” On October 9, the official Hugging Face GGUF repository is accessible and includes llama.cpp, Ollama and other quickstarts. This guide uses that currently published GGUF repository. The BF16 repository remains a separate artifact; check both repositories for changes before repeating these steps.

### The big change

You can now follow a published llama.cpp quickstart for a quantized Mellum2.1 build and query it through a local OpenAI-compatible endpoint. That makes a first self-hosted inference check practical without relying on an inference-provider deployment for the BF16 repository. A response proves the serving path works; it does not prove agent quality, fit, privacy across an entire toolchain, or readiness for production.

### Sources & further reading

- [JetBrains: “Mellum2.1 Gets to Work” (October 8, 2026)](https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/) — launch description and the original GGUF availability statement.
- [JetBrains Mellum2.1 BF16 model card](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) — model details, vLLM and Transformers instructions, benchmarks and license.
- [JetBrains Mellum2.1 GGUF repository](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF) — current quantizations, file sizes, llama.cpp commands and local API example.
- [llama.cpp project](https://github.com/ggml-org/llama.cpp) — runtime source and release information.

## Sources

- [Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents](https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/) — JetBrains launch post; reflects Oct. 8 announcement including its then-current statement that GGUF builds were coming soon.
- [JetBrains/Mellum2.1-12B-A2.5B-Thinking model card](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking) — Official BF16 artifact details, model-card quickstarts, benchmark table and inference-provider status; benchmark values are self-reported by JetBrains.
- [JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF model card and quickstart](https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF) — Current official GGUF artifact, quantization file sizes, Windows llama.cpp quickstart and local OpenAI-compatible endpoint.
- [llama.cpp](https://github.com/ggml-org/llama.cpp) — Runtime project reference; no runtime version is asserted.
