Private · European · flat price

A private AI endpoint
with a flat price,
not a meter.

NEURA Brain, hosted in Europe, behind an OpenAI-compatible API. You buy slots, not tokens: €69 a month each, or €59 on an annual plan, and just €39 during early access. Nothing is billed per token, per seat or per retry, and what you send is never used for training.

  • ∞ One price per slot. No token meter, no seat per person.
  • ⊘ Private by contract: hosted in the EU, never used for training.
  • ↻ Change one URL. Open WebUI, Cursor, your agents and scripts keep working.

What runs underneath, said plainly: at the end of the page.

The product

One endpoint.
Three promises.

NEURA Brain Private Endpoint is the private brain behind the software you already use. Point your tools, agents and scripts at one URL and they run on a private model, in Europe, at a price that does not move with the work.

01

A price, not a meter.

A slot is one answer at a time, all day, every day. It costs €69 a month, or €59 a month on an annual plan, and €39 during early access. Input, output, retries and agent loops are not counted. Finance writes one line in January, and it is still true in December.

02

Private by contract.

What you send is used to answer you, and for nothing else: not for training, not for a dataset, not for anyone. The work runs in the European Union and is operated by IANUSTEC S.R.L. in Italy. Conversation memory is kept for you alone and deleted on request.

03

Context that does not run out.

The model reads long files in full. When a conversation outgrows what the model can hold, IANUSTEC keeps the whole history and brings back the passages each answer needs, instead of cutting the oldest turns.

Measured

Numbers from the real stack.
October 2026.

We measure the service with real, long chats and reasoning switched on. Here are the figures, and the limits that come with them.

≈ 1 s

The model starts working.

On a conversation that is already open, with a long history behind the question, thinking begins in about a second (median 0.9 s, 95% of requests under 1.7 s). The written answer follows once the reasoning is done: typically 5 to 8 seconds.

60–120

Tokens a second, faster than anyone reads.

Between 60 and 120 tokens a second, about 45 to 90 words, depending on how busy the pool is. The answer scrolls as fast as you can follow it.

14 s

The first message of a long conversation.

Reading a 200-page conversation from scratch takes about 14 seconds. After that it stays warm, and every further message starts in about a second.

10,320

Answers under load. Recall 100%.

Earlier test, August 2026: 10,320 requests against a 166,000-token case file, every detail recalled, zero errors. On a paired statistical test (HotpotQA) the result was no different from giving the model the whole file.

What happens What to expect
A conversation you already have open Thinking starts in about a second. Written answer in 5 to 8 seconds with reasoning on.
The first message after a pause or a restart The conversation is read again: about 14 seconds for 200 pages.
More requests at the same moment than you have slots They wait in a queue and start on their own. Past a maximum wait you get a clear retry response, never a silent failure.
Agents running back to back, without pauses Tested: they keep writing at about 20 tokens a second or better.
The full service, with the IANUSTEC memory layer on top Not in these figures yet. We re-run them on the complete service before launch.

Serving engine measured 2–10 October 2026. Raw data under NDA, or we re-run it live on your own file, on a call. The method stays closed. The outcome does not.

Cost

A bill that does not move.
Whatever the agents do.

On a meter, an agent that loops is a bill that grows: every step reads the whole context again, and every retry is billed. On a slot it is the same invoice. Estimate your own workload and see what a meter would add up to.

Per-token API €0 a month, list price
NEURA Brain, 1 slot €177 a month, annual plan

Per-token prices are the OpenRouter list for Qwen3.8-27B on 11 October 2026: typical provider $0.15 per million tokens in, $1.875 out, $0.0375 cached; cheapest listed $0.04, $1.35, $0.018. Converted at €0.92 per $1. With caching on, 80% of the context is billed at the cached price. Slots are estimated at 15 seconds per request, a peak three times the average, and a slot busy at most 60% of the time. Your numbers will differ: this is arithmetic, not a quote.

What the price buys besides the arithmetic.

  • A contract, not a terms page. Hosted in the EU, a DPA, no training. Most per-token routes do not tell you where a request lands.
  • Capacity that is yours. A slot is reserved for you. No rate-limit surprises when the shared pool is busy.
  • The same price when an agent loops. A retry storm at 3am is work, not an incident on the invoice.
  • Memory that does not cut. The whole history is kept and brought back, not trimmed.

How many slots do you need?

A slot is one answer at a time. It is not a person, and it is not a seat.

  1. Ten people is not ten questions at once. They read, they think, they type the next one. For most of the minute nobody is waiting on the model. When one answer lands, the next colleague can ask.
  2. The busy second is about one in three. In a professional office the peak we plan for is 25–30% of the room sending at the same moment. The rest are still working, just not hitting the endpoint together.
  3. A slot is not a wall. A slot answers one request at a time. Send three together and they queue: the second starts when the first lands. At about ten seconds each, three requests together are all back within half a minute. If the wait shows up in the room, you add a slot.
  4. The number follows from that. 10 people × 30% = 3 questions in flight. Round up for a bad Monday: 4 slots: €236 a month on the annual plan, €156 during early access.

10 people×30% peak=3 in flight→4 slots

On the annual plan a desk of ten is about €24 a person a month, or about €16 during early access. Typical chat-assistant seats list at $20–25 a person (August 2026 list prices), and they are not private by contract.

Agents are the exception: they can stay on a question for hours. If the desk is agents, plan one slot per agent that stays hot, or one for every two that take turns.

Who it is for

Not a vertical.
A kind of work.

If the work is continuous, if the files live for months, if the data cannot go to a public API, or if finance needs the number before anyone types, you are the customer. A studio, a ward, a workshop or a desk of agents: the job is the same.

This is for you if

  • Your agents run all day: coding agents, OpenClaw, Hermes, n8n workflows.
  • The conversation is the work: a case, a patient, a project that cannot be cut.
  • The data cannot go to a public API or be used for training.
  • You need a monthly figure before anyone sends a prompt.
  • You already have the software, and only need the AI behind it.

This is not for you if

  • You want to buy an app, not an endpoint. Your software has to be able to call an OpenAI-compatible URL.
  • You want a finished chat app with files and automations. That is NEURA Cloud or Rack.
  • You need the largest frontier model for every single request.
  • You want to try it on Friday and forget it on Monday.
Agents

Leave your agents on.
The loop is work, not a line item.

An agent is not a prompt. It reads, retries, keeps memory, writes back on WhatsApp at 3am. On a token meter every loop costs money, so you learn to switch it off at six. On a slot it is just work. If the agent has a base URL field, it works with us.

Always on · your chats

OpenClaw

Your assistant in the rooms you already use: WhatsApp, Telegram, Slack, Discord, Signal, iMessage. One gateway on your machine with memory, cron and tools. Add NEURA Brain as a custom OpenAI-compatible provider.

openclaw.ai
Always on · one container each

NanoClaw

The same job, built to be left running. Each agent lives in its own container, not in one process with the keys to the house. Point it at an OpenAI-compatible host when you want a private brain.

nanoclaw.dev
Always on · memory and skills

Hermes

Nous Research. Persistent memory, terminal, files, skills. A self-hosted agent that already talks to any /v1/chat/completions host. Set the base URL and it keeps working.

hermes-agent.nousresearch.com

They already take an OpenAI URL. That is enough.

Coding desk
  • Cursor
  • Cline
  • Continue
  • Aider
  • OpenHands
  • Goose
  • OpenCode
Teams and workflows
  • CrewAI
  • LangGraph
  • n8n
  • AutoGen
The rule

If it has a base URL field, paste https://api.neurabrain.io/v1 and use model neura-brain. Plan one slot for each agent that stays hot around the clock.

Three fields. Same as any client.

OpenClaw: a custom provider, openai-completions, base URL with /v1. Hermes: a custom endpoint in config.yaml. Cursor, Cline, Continue: the OpenAI-compatible host you already see in settings. We do not ship the agent. We are the brain it runs on.

Base URL
https://api.neurabrain.io/v1
API key
nr-sk-…
Model
neura-brain
Real work

Tasks we run every day.
On our own systems.

These come from IANUSTEC's own assistant: Open WebUI with tools connected to Odoo, Nextcloud, email, calendar and a browser. They share a shape: many tool calls, results much longer than a chat message, and data that must stay private. That is the work the endpoint is built for.

Odoo · Nextcloud calendar

Mileage claims, in one request.

The assistant lists the month's calendar events, finds the company vehicle in Odoo, works out the trips and writes the odometer entries. One real conversation: about 90 tool calls, each result longer than the question that started it.

Real task, run in our assistant
Email · Nextcloud

Find the message, the attachment, the folder.

It reads the inbox, opens a forwarded email and its attachment, and searches the shared folders for the matching practice. Documents that would never be pasted into a public chat.

Real task, run in our assistant
Odoo CRM

Every lead, one by one.

It reads the whole pipeline, counts by stage and revenue, and flags what has no activity. Measured in September 2026: 167 leads analysed in about two minutes, with correct totals in 9 runs out of 9.

Measured, September 2026
Tickets · OpenProject

Which tickets are late, and whose.

Overdue tickets, who has the most, the five most urgent with priority and due date. In our tests the dates came out right in 3 runs out of 3, on a list of 120 tickets.

Measured, September 2026
Browser agent

An agent that works a web page.

It navigates, reads the elements on the page, clicks and fills in forms, one step after another. A single browser test in our assistant took 136 tool calls: exactly the kind of loop that is expensive on a meter.

Real task, run in our assistant
Open WebUI · LibreChat

An assistant that remembers the practice.

One private conversation per client or case, growing for months: files, emails, decisions. The model reads the file directly. When the history outgrows it, IANUSTEC keeps the whole conversation and brings back what each answer needs.

Law firms, accountants, consultancies, local government
Privacy

Answered in Europe.
Never used for training.

European rules are not a page in the footer. They are how this is built. The work stays in the European Union. What you send answers your question and builds your own memory, and nothing else: no training, no dataset, no resale.

01

In the Union.

The answer is made in the EU. The country can change from one member state to another. The Union does not.

02

AI Act, by design.

IANUSTEC runs this, from Italy, and a named company is accountable. That is the starting point of the Act, not a sticker added after launch.

03

Memory that is yours.

Your conversation memory is kept for your account only, never used to train a model, and deleted on request or when the contract ends.

Where the work happens

Green is a country in the Union.

Every pin is in the Union. The work can move from one EU country to another. It does not leave. The United Kingdom is on the map, and it is not a pin: we do not leave the Union to save money.

Where NEURA Brain answers in the European Union Map of Europe. Green marks are countries in the Union where the answer can be made. The United Kingdom is on the land. It is not a pin. IE FR ES LU NL DE CZ AT PL SE EE SI IT HR RO BG
Trust

Four ISO certificates.
So you do not have to worry.

IANUSTEC holds them so security and privacy stay on our side of the desk. You run the work. NEURA Brain is a product of IANUSTEC S.R.L., Italy.

The certificates are the audit. This is the clause.

You should not have to think about where a file sits, who can train on it, or which country it woke up in. That is IANUSTEC's job. ISO is the audit. The contract is the rest.

  1. Europe, not a sticker. Processed in the Union. The data does not leave because a region was cheaper. Where the work happens.
  2. GDPR, on paper. A DPA is available. Deletion on request. The memory is yours to take back.
  3. No training, in the contract. Access, not a data business. What you send is not a product we sell.

Europe·GDPR·no training→IANUSTEC S.R.L.

Conegliano (TV), Italy. The operating seat. The company you sign with.

Pricing

One price per slot.
You pick the number.

Early access is €39 a month per slot, paid upfront for twelve months and locked. After that it is €59 a month on an annual plan, or €69 month to month. No tiers, no minimum seats, no per-token billing: a slot costs the same whether you take one or twenty. Users are unlimited. Tokens are not counted.

Not users. A slot is one answer at a time. In an office about one person in three is asking at any moment, so ten people usually need three or four slots. An agent that stays hot around the clock needs one of its own. Type any number; the slider is a shortcut.

slot
164 · type more
Early access, 1 slot
€39/mo

€468 paid upfront for twelve months. Price locked.

After twelve months, annual plan
€59/mo
After twelve months, month to month
€69/mo
Request early access

€39 a month per slot, prepaid for twelve months. A person replies within 24 hours.

  • No seats. Ten operators or a hundred. You do not pay per person.
  • 1 slot. The only thing you buy: how many answers run at the same time. Extra requests wait their turn.
  • No token meter. Input, output, retries and agent loops are not counted. A fair-use policy protects the pool from abuse.
  • Reasoning on. Every answer thinks before it writes.
  • Memory that does not cut. The whole history is kept and brought back when an answer needs it.
  • 24/7. Nights, weekends, December. No overtime surcharge.
  • Private. Hosted in the EU, a DPA, never used for training.
  • One number. The figure you see is the figure finance writes down.
Bring your own client

The clients are open source.
The endpoint is ours.

We do not ship a chat app and we do not sell a UI. These clients are free software, yours to install. NEURA Brain is the private European endpoint they talk to. Three fields: URL, key, model neura-brain.

Base URL https://api.neurabrain.io/v1
API key nr-sk-…
Model neura-brain
Open WebUITeam chat

The most deployed open chat UI. RAG, tools, multi-user roles. Add NEURA Brain as a custom OpenAI connection.

LibreChatTeam chat

Multi-provider team chat with agents, MCP and LDAP/SSO. One UI for NEURA Brain and the rest of your stack.

AnythingLLMDesktop · server

Document RAG without the setup. Drop PDFs, point at NEURA Brain, chat with citations. Desktop app or Docker.

JanDesktop

Native Mac, Windows and Linux app, privacy first. Any OpenAI-compatible endpoint next to local models.

ChatboxDesktop · mobile

Lightweight client for Windows, Mac, Linux, iOS and Android, built to plug in a custom API.

Cherry StudioDesktop

Multi-model desktop assistant with agents, documents and an OpenAI-compatible provider field.

OpenClawAlways on

A gateway on your machine: WhatsApp, Telegram, Slack, Discord. Custom OpenAI-compatible provider, base URL with /v1.

HermesAlways on

Nous Research. Memory, skills, terminal. A custom endpoint in the config, any host that speaks /v1/chat/completions.

MIT · self-hostedDocs
NanoClawAlways on

One agent, one container, built to be left running. Point its OpenAI-compatible path at NEURA Brain.

ClineCoding

VS Code, JetBrains and CLI. Plan and act, MCP, bring your own host. Paste the OpenAI-compatible URL.

OpenHandsCoding

Autonomous runs in a sandbox: web UI, SDK, CI. Long jobs that cannot live on a meter.

Your stackAPI

LangChain, LlamaIndex, CrewAI, n8n, or a script you already wrote. OpenAI SDK, one base URL, the rest stays the same.

We recommend these because they are free to run and speak OpenAI. The client is yours. The endpoint is NEURA Brain.

Developers

One line changed.
Everything else identical.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.neurabrain.io/v1",
    api_key="nr-sk-...",
)

resp = client.chat.completions.create(
    model="neura-brain",
    messages=[{"role": "user",
               "content": "What did we decide about the Rossi contract in March?"}],
)
print(resp.choices[0].message.content)  # same SDK, same call
curl https://api.neurabrain.io/v1/chat/completions \
  -H "Authorization: Bearer nr-sk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "neura-brain",
    "messages": [{"role": "user", "content": "What was the penalty clause we agreed in March?"}]
  }'
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.neurabrain.io/v1",
  apiKey: "nr-sk-...",
});

const resp = await client.chat.completions.create({
  model: "neura-brain",
  messages: [{ role: "user", content: "Pick up where we left off" }],
});
Early access

Request early access.
A conversation, not a list.

Tell us who you are and what you will run. A person at IANUSTEC writes back within 24 hours. Early customers pay €39 a slot a month, prepaid for twelve months and locked, and can re-run our measurements on their own files.

  • A person, 24 hoursNot a queue. Someone reads this and writes back.
  • Early access price€39 a slot a month, prepaid for twelve months and locked.
  • Your file, liveWe re-run the measurements on your own data, on a call.

Or write hello@neurabrain.io. Same promise: a person, 24 hours.

Want more

This page is the endpoint.
The infrastructure is NEURA.

IANUSTEC also builds NEURA, a complete private AI infrastructure: software and hardware, on-premises or in our cloud, three redundant nodes. The applications are already there and there are no SaaS seats on top. When the endpoint is not enough, you take the house.

01

Software and hardware

One system: chat, files, automations, admin and the machines they run on. Not a chat bolted onto someone else's cloud.

02

On-prem or cloud

NEURA Rack in your office: data never leaves the building. NEURA Cloud if you want the same platform hosted by IANUSTEC. Same stack, you choose the perimeter.

03

Redundant. Always on.

Three identical nodes, no master. One fails, the others keep the company running. IANUSTEC watches the cluster, never the conversations.

04

No application subscriptions

Open WebUI, Nextcloud, n8n and the admin panel are included, integrated and maintained. Odoo, OpenProject and Chatwoot when you want them. No per-app seats.

In the building

NEURA Rack

Hardware on your floor: three nodes, a dedicated router, one Ethernet cable. Processing stays inside the perimeter.

Managed

NEURA Cloud

The same applications and the same admin, with no rack in the office. IANUSTEC runs the metal. You run the company.

Under the hood

What NEURA Brain
runs on today.

Said plainly: today NEURA Brain is Qwen3.8-27B, an open-weight model from the Qwen team. It is the best balance of cost and quality we have found for everyday operations, which is exactly the job of this service. And it is only the beginning: we have just started, and the model behind the name will change as better ones arrive.

What runs today

Qwen3.8-27B.

A 27-billion-parameter model that reasons before it answers, reads long files and calls tools. It is the same open model you could download yourself.

  • Open weights from the Qwen team
  • Reasoning on, built for long files and tool calls
What we add

The service around it.

A model is a file. A service is a price, a queue, a memory and a contract. That is the part you buy, and it stays the same when the model changes.

  • Flat-price slots, with a queue instead of refusals
  • Memory that keeps the whole history, uncut
  • EU hosting, a DPA, no training, ISO-certified operator
  • An OpenAI-compatible API your tools already speak

Tested before we ship it: on our agent test battery (CRM analysis, ticket triage, coding tasks, 18 episodes) the build we serve scored no worse than the full-precision original. When a better model replaces this one we say so first, and we run the same tests.