JetBrains’ Mellum2.1 is available as a 12-billion-parameter model with 2.5 billion active parameters. The BF16 repository and a separate GGUF repository are both live on Hugging Face. For a first local run, the GGUF repository documents a llama.cpp route: download a quantized file on demand, start a local server, send it a prompt, and inspect the answer before giving any coding agent access to a repository.

This is a documentation-based setup guide. BIG CHANGE did not install Mellum2.1 or run the commands below. The success check is a local completion returned by the server; it does not establish coding quality, agent reliability, or performance on your hardware.

What you need

  • A Windows, macOS, or Linux computer with enough memory and storage for the selected model and its runtime. JetBrains’ card lists the original model as BF16 with a 131,072-token context. The GGUF repository lists its recommended Q4_K_M file at 8.1 GB. That is the file size, not a complete estimate of runtime memory: the runtime, context and other processes need additional resources. JetBrains does not publish a minimum memory requirement.
  • An internet connection for the initial software/model download. The inference request itself can go to the local server.
  • llama.cpp and a terminal. The repository documents winget install llama.cpp for Windows and a llama serve command for local serving.

The model is released under Apache 2.0. The weights have no listed usage fee; hardware, electricity, storage and any infrastructure you choose to rent have costs that JetBrains does not quantify. The BF16 model card currently says no inference provider serves that repository. The GGUF repository is a separate quantized artifact with its own llama.cpp quickstart.

Step 1: Install llama.cpp and start the local server

In PowerShell on Windows, install llama.cpp using the official Mellum2.1 GGUF quickstart:

PowerShell
winget install llama.cpp

Open a new terminal if the llama command is not yet on your PATH. Start the recommended Q4_K_M build and explicitly bind it to the local machine on port 8080:

PowerShell
llama serve -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M --host 127.0.0.1 --port 8080

The GGUF repository documents this model ID and quantization for llama serve; it also documents the Windows install route. The llama.cpp server reference documents --host and --port; the repository’s example endpoint is http://localhost:8080/v1. On macOS and Linux, the same GGUF repository documents installing llama.cpp with curl -LsSf https://llama.app/install.sh | sh, followed by the same serve command.

The process must download the model before its first response. The Q4_K_M file is listed as 8.1 GB. Keep the server bound to your own machine for this test; do not expose it to a network or put credentials into the prompt. The repository also lists a smaller 7.0 GB MXFP4_MOE file and larger Q6_K, Q8_0 and BF16 variants. Quantization changes the model artifact; the listed file size alone does not tell you whether a particular computer can serve it at a useful context length or speed.

If the command fails to start, first check the exact model ID, available disk space, the llama.cpp version and the full error text. A memory allocation failure is a reason to stop and check runtime options and context settings against the current llama.cpp documentation; it is not evidence that the model is defective. Do not assume the published 131,072-token context will fit on your machine.

Step 2: Send a smoke-test prompt

Leave the server running. In a second PowerShell window, send a short request to its local API:

PowerShell
$body = @{
  model = "JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF"
  messages = @(
    @{ role = "user"; content = "Reply with exactly: MELLUM21-LOCAL-OK" }
  )
  # Author-selected budget for this short smoke test; not a JetBrains recommendation.
  max_tokens = 512
  temperature = 0.6
  top_p = 0.95
  top_k = 20
} | ConvertTo-Json -Depth 5

$response = Invoke-RestMethod -Uri "http://localhost:8080/v1/chat/completions" -Method Post -ContentType "application/json" -Body $body

$choice = $response.choices[0]
[pscustomobject]@{
  finish_reason = $choice.finish_reason
  has_reasoning_content = -not [string]::IsNullOrWhiteSpace($choice.message.reasoning_content)
  content = $choice.message.content
}

The model ID, local endpoint and sampling values follow the repository’s API example. max_tokens = 512 is a bounded value chosen here for this short check, not a model-card recommendation. Mellum2.1 is a thinking model; the GGUF card says it emits reasoning in <think>...</think> blocks. The code reports whether the separate reasoning_content field is nonempty without printing that field. Depending on the runtime’s reasoning format, content can still include <think> text. The prompt is a smoke test, not a model-quality benchmark.

The basic check passes only if the request returns a completion instead of a connection or server error, finish_reason is not length, and the final content contains MELLUM21-LOCAL-OK. An HTTP success alone is not enough: a thinking model can spend a small output budget before producing the requested final text. If the output is empty, the marker is missing, or finish_reason is length, treat the check as inconclusive; raise the bounded output budget and try again rather than diagnosing a model failure. If PowerShell reports a connection failure, confirm the first terminal still shows a running server and that the request uses port 8080. If the server returns an error, save the exact message and resolve that before testing a coding agent. A successful response confirms that this local inference path can answer one request; it does not confirm the model can safely edit code.

Step 3: Check it before connecting an agent

Keep the first evaluation separate from a live project. Use a disposable repository with a known, small task, and record the model artifact and quantization, llama.cpp version, operating system, context setting, prompt, output, elapsed time and resource use. Ask for a bounded change, inspect the proposed diff, and run the repository’s existing tests yourself. Repeat the same task with your current baseline if you want a comparison. One prompt or one successful test run is not a benchmark.

A model server returns text and, depending on the runtime and request format, may support structured tool-call workflows. It does not by itself decide what tools an agent can use or safely execute them. The surrounding agent controls repository access, shell commands, file edits and test execution. Start with read-only or tightly scoped access, require approval for writes and commands, review every diff, and keep secrets out of the test repository and prompts. Stop if the agent tries to exceed the task, alters unrelated files, or cannot explain a failing test.

For private use, confirm both where inference runs and how your agent is configured. A local model process can keep prompts on your machine when requests stay on the local endpoint, but an agent may still call external services for other features. Check network access, logging, telemetry and tool permissions in that application before using private code or data.

Step 4: What the published evidence does and does not show

JetBrains describes Mellum2.1 as a 12B mixture-of-experts model with 2.5B active parameters, BF16 precision and a 131,072-token context. Its model card says the release is Apache 2.0 and describes post-training that included reinforcement learning in sandboxed software environments. JetBrains also publishes benchmark results, including agentic coding evaluations. Those scores are JetBrains’ self-reported results; they are not measurements from BIG CHANGE, and they do not predict your hardware’s throughput or your project’s outcomes.

One release detail changed between publication surfaces. JetBrains’ October 8 launch post said GGUF builds were “coming soon.” On October 9, the official Hugging Face GGUF repository is accessible and includes llama.cpp, Ollama and other quickstarts. This guide uses that currently published GGUF repository. The BF16 repository remains a separate artifact; check both repositories for changes before repeating these steps.

The big change

You can now follow a published llama.cpp quickstart for a quantized Mellum2.1 build and query it through a local OpenAI-compatible endpoint. That makes a first self-hosted inference check practical without relying on an inference-provider deployment for the BF16 repository. A response proves the serving path works; it does not prove agent quality, fit, privacy across an entire toolchain, or readiness for production.

Sources & further reading