Dedicated model · Available as managed deployment
A vision-capable open model with a 262,144-token context window — the largest in the AxForge catalogue — validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. An OpenAI-compatible endpoint on your own machine, operated by AxForge in Málaga, Spain.
Why AxForge
| 262,144-token context | The largest context window in the AxForge catalogue — whole documents, long transcripts and large tool outputs in one request. |
|---|---|
| Text + vision | Gemma-4 26B accepts image input alongside text: the same endpoint reads screenshots, scans and photos, served under the model name gemma on your own OpenAI-compatible /v1. |
| EU-hosted, zero prompt retention | A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) in Málaga, Spain (eu-es-1). Prompts and completions are processed in memory — not logged, not retained, never used to train. |
Specifications
| Model | Gemma-4 26B — open model, Gemma family |
|---|---|
| Served model name | gemma |
| Context window | 262,144 tokens |
| Modalities | Text + vision — accepts image input |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the Gemma-4 26B page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys Gemma-4 26B on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with model gemma — text and image input. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
262,144 tokens — the largest context window in the AxForge catalogue.
Yes, the model is vision-capable: it accepts image input alongside text.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Gemma-4 26B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.
Two parts: the DGX Spark hardware rental — by the hour, week, month or year, with longer terms earning the lower rate — and the managed service, quoted per deployment. Both are confirmed in writing before anything is billed.