Back to Discover

model-council

connector

Totti0135

Consult other LLMs as tools — ask them, compare their answers, synthesize a conclusion.

View on GitHub
0 starsSynced Aug 17, 2026

Install to Claude Code

/plugin marketplace add Totti0135/model-council

README

Model Council

English | 简体中文

An MCP server that seats other LLMs at your table. Your assistant asks them, reads their answers as tool results, relays those answers back and forth for critique, and gives you one merged conclusion — inside a single normal conversation, with no copy-paste.

Your assistant chairs the council. Any number of members, from any mix of OpenAI-compatible and Anthropic-compatible endpoints — a hosted API, a self-run gateway, a local server, or several of each.

Tools

ToolWhat it does
ask(model, prompt)Ask one member by id
ask_all(prompt, models?, rounds?)Ask everyone (or a named subset) the same prompt in parallel, answers side by side. rounds=2 turns it into a discussion
list_council()The roster: ids, endpoints, weights, call budget, and whether each member is ready. No network calls
probe_models(model?)Ask a provider's /models route what ids it really exposes

Members are stateless and cannot see your conversation, so the chair passes everything they need in each call. That is exactly what makes cross-review work: it puts one member's answer inside another's prompt.

Rounds

ask_all(prompt, rounds=2) runs that cross-review for you. Round 1 is the usual parallel ask. Round 2 goes back to each member carrying the question plus every answer from round 1 — its own and the others', verbatim — and asks it to revise: take what is right, correct what is not, and say where it still disagrees and why. The transcript comes back round by round, so you can see who moved and who held their ground.

Carrying the previous round back is the whole mechanism. Members remember nothing between calls, so without it a second round is just the same question asked twice. Up to 3 rounds; each one costs another call per member and a longer prompt than the last, so 1 is right for a survey of opinion and 2 for a question where the disagreement is the interesting part.

Members can also carry different weights, for the common case where the council is not a council of equals.

Retries

A call that fails transiently is retried before it is reported: HTTP 429 and 5xx, and dropped or timed-out connections. Backoff is exponential from 1s with jitter, and a Retry-After header wins over that curve — unless it asks for longer than 30s, in which case the call stops and says so rather than sitting on it. Failures that will not change on a second look — 401, 404, a malformed response body — are reported immediately; retrying them only spends the same quota to be told the same thing. Defaults to 2 retries; set retries: 0 for the old single-shot behaviour.

A member that exhausts its attempts returns its error as that member's answer, so the rest of the council still answers.

Install

The server is listed on the official MCP Registry as io.github.Totti0135/model-council, so a client that browses the registry can find and add it there. To wire it up by hand instead, read on.

It runs from PyPI with no clone and no virtualenv. You need uv:

curl -LsSf https://astral.sh/uv/install.sh | sh

Claude Desktop

Easiest is the desktop extension. Download model-council-<version>.mcpb from the latest release and drag it onto Settings → Extensions. The app asks for the endpoints and keys in a form and keeps the keys in your OS keychain, so no file on disk holds them. The form seats two models; for a larger council, point its "Config file" field at a JSON config (see below).

To wire it up by hand instead, edit claude_desktop_config.json (Settings → Developer → Edit Config), add the block below, then fully quit and reopen the app — Cmd-Q, not just closing the window. You will know it worked when the tools menu lists model-council.

{
  "mcpServers": {
    "model-council": {
      "command": "uvx",
      "args": ["model-council-mcp"],
      "env": {
        "COUNCIL_MODELS": "gpt5,glm",
        "GPT5_BASE_URL": "https://your-openai-compatible-host/v1",
        "GPT5_API_KEY": "sk-xxxxxxxx",
        "GPT5_MODEL": "gpt-5",
        "GLM_BASE_URL": "https://open.bigmodel.cn/api/anthropic",
        "GLM_API_KEY": "xxxxxxxx",
        "GLM_MODEL": "glm-4.6",
        "GLM_FORMAT": "anthropic"
      }
    }
  }
}

Claude Code

claude mcp add model-council -e GPT5_BASE_URL=... -e GPT5_API_KEY=... -- uvx model-council-mcp

Other MCP clients

Anything that launches a stdio server works: run uvx model-council-mcp and pass the same environment variables.

Serving a team over HTTP

Everything above runs one copy of the server per person, launched by their own MCP client, configured with their own keys. --http is the other shape: one deployment holds one set of keys and answers a whole team, who configure a URL and no secret at all.

model-council-mcp --http --host 0.0.0.0 --allow 10.20.0.0/16

Colleagues then add it as a remote server, with nothing sensitive in the config:

claude mcp add --transport http model-council http://council.internal:8000/mcp
// Claude Desktop and other clients
{ "mcpServers": { "model-council": { "type": "http", "url": "http://council.internal:8000/mcp" } } }

deploy/ has a Dockerfile, a compose file and a systemd unit.

Who may call

Over stdio the operating system answers this: whoever launched the process already had the keys. HTTP removes that guarantee — one port now stands in front of shared provider quota — so the server refuses to start on a non-loopback address until --allow says who may reach it. Loopback needs no flag, and is always admitted.

Flag
--allowCIDRs, bare addresses, or the names private (RFC1918), loopback, any
--trust-proxyreverse proxies whose X-Forwarded-For may be believed
--allow-originpermit a browser Origin; repeatable

--trust-proxy is the one worth reading twice. Without it X-Forwarded-For is ignored entirely and the peer address decides — behind nginx that is nginx, so the allowlist matches everyone or no one. With it, the client is the rightmost hop in the chain that is not a trusted proxy, which is what stops a caller from writing X-Forwarded-For: 10.0.0.1 and walking straight through.

Requests carrying an Origin header are refused by default. MCP clients are not browsers and do not send one; a web page always does. An allowlist admits every machine on the office network, and each of those runs a browser that will issue requests on behalf of whatever page it has open — Origin is what tells the two apart.

Every flag has an environment variable (COUNCIL_ALLOW, COUNCIL_TRUST_PROXY, COUNCIL_HTTP_HOST, …); --help lists them.

What this does not do

The allowlist is a network boundary, not an identity. Anyone inside it calls without a credential, so usage cannot be attributed to a person, rate-limited per person, or revoked for one person. That is a deliberate trade — it is what makes the client config a bare URL — but it means the network has to be a boundary you actually trust, and it does not survive contact with a VPN that admits contractors, or a CI runner on the same subnet.

If you need per-person attribution or quota, put an LLM gateway (LiteLLM, one-api, or whatever your organisation already runs) behind this server and give each caller their own virtual key there, or run this behind a reverse proxy that does SSO.

One more thing worth knowing before you announce the URL: this service forwards whatever text it is given to an external provider. A shared endpoint with no credential is a data-egress path for everyone who can reach it.

Behind a reverse proxy

ask_all with rounds=2 is a long request — several models, several attempts each, at up to COUNCIL_TIMEOUT (180s) per call. Default proxy timeouts will cut it off well before the server is finished:

location /mcp {
    proxy_pass http://127.0.0.1:8000;
    proxy_http_version 1.1;
    proxy_buffering off;          # the transport streams; buffering defeats it
    proxy_read_timeout 900s;
    proxy_set_header Host $host;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}

Then start the server with --trust-proxy set to the proxy's address, or the allowlist will only ever see the proxy.

Configuring the council

Two layers, so several models can share one endpoint without repeating its credentials:

  • provider — an endpoint: base_url + api_key + which wire format it speaks
  • member — one model on some provider, addressed by a short id

Configuration comes from whichever source is most explicit: the file COUNCIL_CONFIG points at, else the roster COUNCIL_MODELS names, else a config file at ~/.config/model-council/config.json, else the built-in default roster. Explicit beats discovered on purpose — a config file you left lying around must not silently override settings a client just handed the server. list_council() always reports which source won.

Environment variables

COUNCIL_MODELS lists the ids; each id gets variables named after it, uppercased with non-alphanumeric characters turned into underscores (my-modelMY_MODEL_BASE_URL).

COUNCIL_MODELS=gpt5,glm
GPT5_BASE_URL=https://your-openai-compatible-host/v1
GPT5_API_KEY=sk-xxxxxxxx
GPT5_MODEL=gpt-5
GLM_BASE_URL=https://open.bigmodel.cn/api/anthropic
GLM_API_KEY=xxxxxxxx
GLM_MODEL=glm-4.6
GLM_FORMAT=anthropic

Per member: _BASE_URL, _API_KEY, _MODEL, _FORMAT, _LABEL, _WEIGHT, _MAX_TOKENS, _TEMPERATURE, _TIMEOUT, _RETRIES, _RETRY_BACKOFF, _HEADERS (a JSON object), _PROXY, _ENABLED. Globally: COUNCIL_TIMEOUT, COUNCIL_RETRIES, COUNCIL_RETRY_BACKOFF, COUNCIL_CONFIG, COUNCIL_ENV_FILE.

Omit COUNCIL_MODELS and the roster defaults to chatgpt,glm, reading CHATGPT_* and GLM_*.

A config file

Better once you have more than a handful of members, or when several share an endpoint. Set COUNCIL_CONFIG=/path/to/config.json, or drop the file at ~/.config/model-council/config.json where the server finds it on its own.

{
  "providers": {
    "my-relay": {
      "base_url": "https://your-openai-compatible-host/v1",
      "api_key": "${MY_RELAY_KEY}",
      "format": "openai"
    },
    "zhipu": {
      "base_url": "https://open.bigmodel.cn/api/anthropic",
      "api_key": "${GLM_KEY}",
      "format": "anthropic"
    }
  },
  "members": [
    { "id": "gpt5",  "provider": "my-relay", "model": "gpt-5", "label": "GPT-5",
      "weight": 2 },
    { "id": "codex", "provider": "my-relay", "model": "gpt-5-codex", "temperature": 0.2 },
    { "id": "glm",   "provider": "zhipu",    "model": "glm-4.6" },
    { "id": "kimi",  "base_url": "https://api.moonshot.cn/v1",
      "api_key": "${KIMI_KEY}", "model": "kimi-k2" }
  ]
}

${ENV_VAR} is expanded from the environment, so the file carries no secrets and can be shared or committed. See examples/config.json for a fully annotated version.

A member gets its connection one of two ways, never a mix of both: name a provider and take that endpoint whole, or omit provider and supply base_url + api_key + format yourself (as kimi does above). Naming a provider and overriding one of those three is refused — that member is disabled and list_council says why. The reason is that a partial override would pair one endpoint's credentials with another endpoint's URL, quietly sending your key to a host it was never issued for. Per-member headers, timeout, temperature, max_tokens and label are not part of that identity and stay overridable.

Fields

The first three travel together as one unit — see the rule above.

FieldApplies toNotes
base_urlprovider, or a member with no providerRoot the route hangs off — /chat/completions for openai, /v1/messages for anthropic. Usually ends in /v1 for OpenAI-compatible hosts
api_keyprovider, or a member with no provider
formatprovider, or a member with no provideropenai (default) or anthropic
modelmemberThe model id sent to the endpoint
labelmemberDisplay name in answers; defaults to the id
weightmemberHow far this member's opinion carries. Default 1, max 10, 0 for advisory only. See Weights
max_tokensmemberAnthropic format only, where it is required. Default 8192
temperaturememberSent only when set
headersprovider, memberExtra HTTP headers
timeoutprovider, memberSeconds, per attempt. Default 180
retriesprovider, memberExtra attempts a transient failure gets. Default 2, max 5, 0 to disable
retry_backoffprovider, memberSeconds before the first retry, doubling from there. Default 1
proxyprovider, memberOmit to follow HTTP_PROXY/HTTPS_PROXY; false to connect directly; a URL to use that proxy
enabledmemberfalse parks a member without deleting its config

timeout, retries and retry_backoff can also be set at the top level of the config file, as the default every member inherits.

Weights

A council is rarely made of equals. weight says how far a member's opinion carries — everyone is 1 until you say otherwise, and only the ratios mean anything, so 2 and 1 is the same council as 10 and 5.

{ "id": "gpt5", "provider": "my-relay", "model": "gpt-5", "weight": 2 }

It changes nothing about the call. When the weights differ, ask_all labels each answer with its weight and ends the transcript with the ranking:

===== GPT-5 (gpt-5) · weight 2 =====
...
===== Local (qwen3-8b) · weight 0.5 =====
...

[WEIGHTS — GPT-5 2, GLM 1, Local 0.5]
These are this council's standing priors on its members, not votes. ...

Weight belongs to the seat, not the endpoint, so a provider cannot set it: two members on one relay may be a frontier model and a small fast one. 0 means advisory — the member answers and is read, but its agreement counts for nothing. Anything unusable (a negative, a word) falls back to 1 rather than corrupting the ranking silently; list_council prints the effective value.

The members are never told each other's weights. A model informed that it is outranked stops arguing and starts agreeing, which costs exactly the independent dissent a council is assembled to produce — so round two carries the other answers and not their standing. The weights are for whoever reads the transcript, and they are a prior, not a vote: they break ties and decide who carries the burden of proof. A specific, checkable reason from the lowest weight still beats a bare assertion from the highest.

Wire format notes

  • format is not inferred from the URL. Pointing base_url at an Anthropic-style endpoint without also setting format: "anthropic" leaves the member on the OpenAI format, and every call fails. This is the single most common misconfiguration.
  • Anthropic endpoints: the server posts to {base_url}/v1/messages, so base_url should not already include the /v1.
  • OpenAI-compatible endpoints: the server uses /chat/completions, never /responses. Some gateways expose both, but /responses may inject a provider-chosen system persona, which is wrong for a general-purpose advisor.
  • A system proxy is followed by default. If a member sits on a network your proxy cannot reach — an internal gateway, typically — it fails with a bare ConnectError that never mentions a proxy. Give that member or provider "proxy": false and it connects directly, while everyone else keeps using the proxy. The error message says so too when a proxy is in play.
  • Model ids move fast. Run probe_models to see what an endpoint actually offers today.

Using it

Things worth typing to the chair:

  • "Answer this yourself, then ask_all and give me a table of where you all agree and disagree."
  • "Ask gpt5 and glm this, then critique both answers and tell me which is more correct and why."
  • "Run two rounds of ask_all on this, then tell me who changed their mind and what actually settled it."
  • "Ask only glm — I want a second opinion on this one file."

Local development

uv sync

Copy .env.example to .env, fill in real values, then:

uv run python tests/test_smoke.py

The smoke test checks both configuration paths offline; with a usable .env it finishes with a live round-trip. To point a client at your working copy, use the model-council-mcp script inside your environment instead of uvx.

Troubleshooting

  • Server doesn't appear — check the client's MCP logs (Claude Desktop: ~/Library/Logs/Claude/mcp*.log). The server writes configuration warnings to stderr at startup.
  • A tool answers [... is not configured] — that member is missing base_url, api_key, or model. Run list_council for a per-member breakdown.
  • HTTP 401 — wrong key, or a key the provider has disabled.
  • HTTP 404 — wrong base_url, or the wrong format for that endpoint.
  • The model id is rejected — run probe_models.
  • A member says gave up after N attempts — it failed transiently every time. The error text is the last one the endpoint gave. list_council shows each member's budget as attempts × timeout.
  • A call takes far longer than the timeout — retries multiply it: three attempts at 180s each is a worst case of ~9 minutes plus backoff. Lower timeout, or retries, for a member you would rather have fail fast.

License

MIT

Rendered live from Totti0135/model-council's GitHub README — not stored, always reflects the source repo.

1 Install Method

NameDescriptionCategorySource
pypi packageInstall via pypi (stdio transport)mcp-servermodel-council-mcp

0 Comments

Login required
Log in to post a comment or update on this repo.

No comments yet — be the first to share an update.