Back to Discover

image-generation-mcp

connector

pvliesdonk

MCP server for AI image generation via OpenAI, Stable Diffusion (SD WebUI), or placeholders.

View on GitHub
0 starsSynced Aug 6, 2026

Install to Claude Code

/plugin marketplace add pvliesdonk/image-generation-mcp

README

Image Generation MCP

CI codecov PyPI Python License Docker Docs llms.txt Template

Multi-provider image generation MCP server built on FastMCP. Generate images from Claude Desktop, Claude Code, or any MCP client using OpenAI, Google Gemini, Stable Diffusion (SD WebUI), or a zero-cost placeholder provider.

Documentation | Config wizard | PyPI | Docker

Features

  • Multi-provider: OpenAI (gpt-image-2, gpt-image-1.5, dall-e-3), Google Gemini (gemini-3.1-flash-image, gemini-3-pro-image, gemini-3.1-flash-lite-image), SD WebUI (Stable Diffusion / Forge / reForge), and a zero-cost placeholder for testing.
  • Per-model style metadata: every model carries a style_profile (strengths, prompt grammar, lifecycle); list_providers includes a top-level warnings array for deprecated models. See Model Catalog.
  • Keyword-based auto-selection: provider="auto" routes by prompt content (text/logo → OpenAI, photoreal/anime → SD WebUI, draft → placeholder).
  • CDN-style image transforms: image://{id}/view?format=webp&width=512&crop_x=... resizes / re-encodes / crops on demand without re-generating.
  • Hybrid background tasks: long-running SD generations run with task=True (poll for status); short OpenAI calls stream progress in the foreground.
  • MCP Apps gallery + viewer: interactive UI surfaces (browse generated images, edit / crop / rotate) for clients that support app: resources.
  • Production deployment: Docker (multi-arch), .deb/.rpm with hardened systemd, OIDC + bearer auth, persistent EventStore for HTTP session resumability.

What you can do with it

With this server mounted in an MCP client, you can ask:

  • "Generate a coffee mug product photo on a worn oak table, 16:9, no text." Routes to gpt-image-1.5 for typography-aware photorealism.
  • "Create three concept-art variations of a cyberpunk alley at dusk." Composes generate_image with provider="sd_webui" and a stylised checkpoint like dreamshaperXL.
  • "Crop this image to a 1:1 square centred on the subject and resize to 512px." Uses image://{id}/view?width=512&height=512&crop_x=... resource transforms.
  • "Show me my recent generations." Browses the gallery via the image://list resource and the MCP Apps gallery viewer.
  • "Save this style as 'cyberpunk-night' so I can apply it to future requests." Uses the style library, whose markdown briefs the LLM interprets per-provider.
  • "Replace the background of my last photo with a sunset sky." Uses transform_image with the gallery image_id as a reference (image-to-image via Gemini).

Installation

From PyPI

pip install image-generation-mcp

If you add optional extras via the PROJECT-EXTRAS-START / PROJECT-EXTRAS-END sentinels in pyproject.toml, document them below:

ExtraIncludesUse when
mcpfastmcp[tasks]>=3.0,<4Background-task support (task=True), required for long SD generations.
openaiopenai>=1.0Enables the OpenAI provider.
google-genaigoogle-genai>=1.0Enables the Gemini provider.
allfastmcp[tasks] + openai + google-genaiEverything except SD WebUI (which is HTTP-only, no extra needed).

Example: pip install image-generation-mcp[all].

From source

git clone https://github.com/pvliesdonk/image-generation-mcp.git
cd image-generation-mcp
uv sync --all-extras --all-groups

Docker

docker pull ghcr.io/pvliesdonk/image-generation-mcp:latest

A compose.yml ships at the repo root as a starting point. Copy .env.example to .env, edit, and docker compose up -d.

To attach a remote Python debugger (development only; the protocol is unauthenticated), see Remote debugging.

Linux packages (.deb / .rpm)

Download .deb or .rpm packages from the GitHub Releases page. Both install a hardened systemd unit; env configuration is sourced from /etc/image-generation-mcp/env (copy from the shipped /etc/image-generation-mcp/env.example).

Claude Desktop (.mcpb bundle)

Download the .mcpb bundle from the GitHub Releases page and double-click to install, or run:

mcpb install image-generation-mcp-<version>.mcpb

Claude Desktop prompts for required env vars via a GUI wizard, with no manual JSON editing needed.

For manual Claude Desktop configuration and setup options, see Claude Desktop deployment.

Quick start

image-generation-mcp serve                                # stdio transport
image-generation-mcp serve --transport http --port 8000   # streamable HTTP

For library usage (embedding the domain logic without the MCP transport), import from the image_generation_mcp package directly. See the project's domain modules under src/image_generation_mcp/ for entry points.

Server info

The server registers a built-in get_server_info tool (via fastmcp_pvl_core.register_server_info_tool) so operators can confirm the deployed version with a single MCP call. The default response carries server_name, server_version, and core_version. Servers that talk to a remote upstream wire upstream version reporting inside the DOMAIN-UPSTREAM-START / DOMAIN-UPSTREAM-END sentinel in src/image_generation_mcp/server.py; see CLAUDE.md for the wiring pattern.

Configuration

Core environment variables shared across all fastmcp-pvl-core-based services:

VariableDefaultDescription
IMAGE_GENERATION_MCP_KV_STORE_URLfile:///data/statePersistent-state backend URL shared by every pvl-core subsystem that needs state. memory:// is in-process and lost on restart; file:///path persists on one server; redis://, dynamodb:// and mongodb:// each need their matching extra. Defaults to file:///data/state when unset.
FASTMCP_LOG_LEVELINFOLog level for FastMCP internals and app loggers (DEBUG / INFO / WARNING / ERROR / CRITICAL). The -v CLI flag overrides to DEBUG.
FASTMCP_ENABLE_RICH_LOGGINGtrueSet false for plain or structured JSON log output.

Domain-specific variables go below under Domain configuration.

Authentication

Callers authenticate via a bearer token or OIDC (mutually exclusive). See the Authentication guide for setup, mapped multi-subject tokens, OIDC, and troubleshooting.

Post-scaffold checklist

After copier copy and gh repo create --push:

  1. Fill in the DOMAIN blocks (every section marked with a DOMAIN sentinel comment) in this README and in CLAUDE.md. The GENERATED-ENV-TABLE-* regions are not DOMAIN blocks; the config generator owns them and rewrites them on every run.
  2. Configure GitHub secrets (see below).
  3. Install dev + docs tooling: uv sync --all-extras --all-groups.
  4. Install pre-commit hooks: uv run pre-commit install.
  5. Run the gate locally: uv run pytest -x -q && uv run ruff check --fix . && uv run ruff format . && uv run mypy src/ tests/.
  6. Push the first commit. CI should be green.

GitHub secrets

CI workflows reference three repository secrets. Configure them via Settings → Secrets and variables → Actions or with gh secret set:

SecretUsed byHow to generate
RELEASE_TOKENrelease.yml, copier-update.yml, renovate.yml, bootstrap.ymlFine-grained PAT at https://github.com/settings/personal-access-tokens/new with contents: write, pull_requests: write, and administration: write (bootstrap sets branch protection + auto-merge). Scoped to this repo.
CODECOV_TOKENci.ymlhttps://codecov.io: sign in with GitHub and add the repo. The upload token is on its settings page.
CLAUDE_CODE_OAUTH_TOKENclaude.yml, claude-code-review.ymlRun claude setup-token locally and paste the result.
gh secret set RELEASE_TOKEN
gh secret set CODECOV_TOKEN
gh secret set CLAUDE_CODE_OAUTH_TOKEN

Dependency updates are handled by Renovate (renovate.yml), which reuses RELEASE_TOKEN. It maintains uv.lock and auto-merges patch/minor bumps once the CI Success check is green; bootstrap.yml enables auto-merge and branch protection on first push. GitHub Actions are updated in the copier template and arrive via copier update, not per-repo.

GITHUB_TOKEN is auto-provided; no action needed.

Local development

The PR gate (matches CI):

uv run pytest -x -q                                  # tests
uv run ruff check --fix . && uv run ruff format .    # lint + format
uv run mypy src/ tests/                              # type-check

Pre-commit runs a subset of the gate on each commit; see .pre-commit-config.yaml for details, or CLAUDE.md for the full Hard PR Acceptance Gates.

Troubleshooting

Moving a scaffolded project

uv sync creates .venv/bin/* scripts with absolute shebangs pointing at the venv Python. If you move the repo after scaffolding (mv /old/path /new/path), uv run pytest fails with ModuleNotFoundError: No module named 'fastmcp' because the stale shebang resolves to a different interpreter than the venv's site-packages.

Fix:

rm -rf .venv
uv sync --all-extras --all-groups

uv run python -m pytest also works as a one-shot workaround (bypasses the stale entry-script shim).

uv.lock refresh after copier update

When copier update introduces new dependencies (such as a new extra added to pyproject.toml.jinja), CI runs uv sync --frozen which fails against a stale lockfile. Run uv lock locally and commit the refreshed uv.lock alongside accepting the copier-update PR.

Links

Domain configuration

Domain environment variables use the IMAGE_GENERATION_MCP_ prefix:

VariableDefaultRequiredDescription
IMAGE_GENERATION_MCP_A1111_HOST(none)NoDeprecated alias for IMAGE_GENERATION_MCP_SD_WEBUI_HOST; logs a warning when used.
IMAGE_GENERATION_MCP_A1111_MODEL(none)NoDeprecated alias for IMAGE_GENERATION_MCP_SD_WEBUI_MODEL; logs a warning when used.
IMAGE_GENERATION_MCP_READ_ONLYtrueNoWhen true, write-tagged tools (image generation, transforms, uploads) are hidden from clients. Set false to enable them.
IMAGE_GENERATION_MCP_SCRATCH_DIR~/.image-generation-mcp/imagesNoDirectory where generated images are saved. Created automatically on first use.
IMAGE_GENERATION_MCP_OPENAI_API_KEY(none)NoOpenAI API key. Enables the OpenAI provider (gpt-image-2, gpt-image-1.5, dall-e-3) when set.
IMAGE_GENERATION_MCP_GOOGLE_API_KEY(none)NoGoogle API key. Enables the Gemini provider (gemini-3.1-flash-image and others) when set. Get a key at https://aistudio.google.com/apikey.
IMAGE_GENERATION_MCP_SD_WEBUI_HOST(none)NoSD WebUI base URL (such as http://localhost:7860). Enables the SD WebUI provider when set. Compatible with AUTOMATIC1111, Forge, reForge, and Forge-neo.
IMAGE_GENERATION_MCP_SD_WEBUI_MODEL(none)NoSD WebUI checkpoint name, used for model-aware preset detection (SD 1.5 / SDXL / Lightning) and checkpoint override. Unset uses the instance's current model.
IMAGE_GENERATION_MCP_DEFAULT_PROVIDERautoNoProvider used when no keyword triggers auto-selection: auto, openai, gemini, sd_webui, or placeholder. auto picks the first configured provider.
IMAGE_GENERATION_MCP_TRANSFORM_CACHE_SIZE64NoMaximum number of transformed image results (resize, crop, convert) kept in memory. Set 0 to disable caching.
IMAGE_GENERATION_MCP_PAID_PROVIDERSopenaiNoComma-separated provider names that cost money; generate_image asks for confirmation (client elicitation) before using them. Set empty to disable confirmation.
IMAGE_GENERATION_MCP_STYLES_DIR~/.image-generation-mcp/stylesNoDirectory for style preset files (Markdown with YAML front matter). Created automatically if it does not exist.
IMAGE_GENERATION_MCP_ALLOW_LOCAL_FILE_INPUTfalseNoAllow reading input images from local filesystem paths. Off by default: only URLs and uploads are accepted.
IMAGE_GENERATION_MCP_MAX_INPUT_IMAGE_BYTES20971520NoMaximum accepted input image size in bytes.
IMAGE_GENERATION_MCP_TRANSFER_TTL_DEFAULT_S3600.0NoDefault lifetime in seconds of a create_download_link / create_upload_link URL when the caller omits one. HTTP transports only.
IMAGE_GENERATION_MCP_TRANSFER_TTL_MAX_S86400.0NoCeiling in seconds a caller-requested transfer-link lifetime is clamped to.
IMAGE_GENERATION_MCP_TRANSFER_GRACE_TTL_S60.0NoPost-success grace window in seconds a one-time transfer link stays reclaimable, so a stalled download can retry.
IMAGE_GENERATION_MCP_TRANSFER_LEASE_S60.0NoReclaim window in seconds for an in-flight transfer whose handler crashed.
IMAGE_GENERATION_MCP_TRANSFER_MAX_UPLOAD_BYTES104857600NoPer-upload size cap in bytes for create_upload_link bodies.
IMAGE_GENERATION_MCP_FETCH_TIMEOUT_S30.0NoHTTP timeout in seconds when fetching remote image URLs (fetch_image and URL inputs).

The create_download_link / create_upload_link tools and the /transfer/{token} route register only on an HTTP or SSE transport with BASE_URL set, and store link tokens in IMAGE_GENERATION_MCP_KV_STORE_URL; the IMAGE_GENERATION_MCP_TRANSFER_* knobs above tune link lifetime and upload limits. Security: IMAGE_GENERATION_MCP_ALLOW_LOCAL_FILE_INPUT grants callers server-filesystem read access via reference-image paths; enable it only for trusted callers or local single-user deployments.

Domain-config fields are composed inside src/image_generation_mcp/config.py between the CONFIG-FIELDS-START / CONFIG-FIELDS-END sentinels; env reads go through fastmcp_pvl_core.env(_ENV_PREFIX, "SUFFIX", default) so naming stays consistent, and field invariants go in __post_init__ between the CONFIG-VALIDATE-START / CONFIG-VALIDATE-END sentinels. Each field's metadata help and tags generate the table above directly, so keep them accurate and complete.

Key design decisions

  • Multi-provider with capability discovery, not feature flags. Each provider's discover_capabilities() reports its actual supported aspect ratios / qualities / formats / negative-prompt support at startup; routing logic asks the capability surface, not a hard-coded enum. New providers slot in by satisfying the protocol, with no router edits needed. (See docs/decisions/0001-…, 0002-…, 0007-….)
  • Per-model style_profile metadata, surfaced via list_providers. Closed-list providers (OpenAI, Gemini, placeholder) use exact-key lookup; SD WebUI uses a regex-ordered pattern table. Profiles include lifecycle flags (current / legacy / deprecated) and feed an auto-built top-level warnings array. (See docs/decisions/0009-….)
  • Hybrid background tasks. Short calls (OpenAI ~5 s) stream progress in-line; long calls (SD WebUI 30-180 s) run as background tasks with check_generation_status polling; clients pick the mode via task=True. (See docs/decisions/0005-….)
  • Image asset model: content-addressed registry + sidecar JSON metadata + on-demand transforms. Generated images keep their full-resolution original; image://{id}/view?format=webp&width=512&crop_x=… resources do format conversion / resize / crop on demand without re-generating. Transforms are cached. (See docs/decisions/0006-….)
  • Style library. User-saved markdown briefs (with YAML frontmatter for tags / aspect ratio / quality) that the LLM interprets per-provider, not copy-pasted verbatim. Distinct from per-model style_profile: style library is the brief; style_profile describes the model. (See docs/decisions/0008-… and 0009-… for disambiguation.)
  • Composes fastmcp_pvl_core.ServerConfig, never inherits. Domain config goes between CONFIG-FIELDS-START / CONFIG-FIELDS-END sentinels; env reads route through fastmcp_pvl_core.env(...) to keep prefix naming consistent.

Rendered live from pvliesdonk/image-generation-mcp's GitHub README — not stored, always reflects the source repo.

2 Install Methods

NameDescriptionCategorySource
pypi packageInstall via pypi (stdio transport)mcp-serverimage-generation-mcp
oci packageInstall via oci (streamable-http transport)mcp-serverghcr.io/pvliesdonk/image-generation-mcp:v1.12.0

0 Comments

Login required
Log in to post a comment or update on this repo.

No comments yet — be the first to share an update.