Guidance for AI coding agents working in this repository.
Xinference is a Python model-serving project for language, embedding, rerank, image, video, audio, and multimodal models. It exposes CLI commands, a Python client, REST/OpenAI-compatible APIs, an xoscar-based distributed runtime, and a React Web UI.
Primary package and entry points:
xinference/: Python package.xinference/deploy/cmdline.py: CLI entry points forxinference,xinference-local,xinference-supervisor, andxinference-worker.xinference/core/: supervisor/worker runtime and actor orchestration.xinference/model/: model families, engines, built-in model specs, and model tests.xinference/api/: API server and OpenAI-compatible routes.xinference/client/: sync and async Python clients.frontend/: Next.js Web UI.doc/source/: Sphinx documentation..github/workflows/python.yaml: main lint and test CI.
- Prefer small, focused changes that match the current module's style.
- Preserve backward compatibility. If a breaking change is unavoidable, document the reason and add deprecation behavior where practical.
- Do not edit
xinference/thirdparty/unless the task explicitly concerns vendored code. - Avoid broad refactors while fixing a localized bug.
- Keep public API behavior, request/response schemas, model registration names, and CLI flags stable unless the task requires changing them.
- Add or update tests for behavior changes. For model-runtime changes, prefer
tests close to the affected model family under
xinference/model/**/tests/. - Use type hints for new Python code when practical; this project encourages PEP 484 style annotations.
- Treat docs and examples as user-facing API. Keep command examples accurate.
Recommended local setup:
conda create --name xinf python=3.12 nodejs
conda activate xinf
pip install -e ".[dev]"Notes:
- Wheel, sdist, and editable builds (including
pip install -e .) run the Web UI build through the in-tree build backend (build_backend.py/build_web.py) unlessNO_WEB_UI=1is set. - For Python 3.12 and newer, CI installs
setuptools<82; use the same pin if packaging or editable installs fail. - The project supports Python 3.10 through 3.13 in CI.
- Optional model engines have extras in
pyproject.toml, such astransformers,vllm,mlx,embedding,rerank,image,video, andaudio.
Python formatting and checks are managed through pre-commit:
pip install pre-commit
pre-commit run --files <modified-files>For a branch-wide check against upstream main:
pre-commit run --from-ref=upstream/main --to-ref=HEAD --all-filesConfigured hooks include Black, end-of-file/trailing-whitespace checks, Ruff,
isort, mypy with missing imports ignored, and codespell. Configuration lives in
.pre-commit-config.yaml and pyproject.toml.
Run focused tests first:
pytest -vv path/to/test_file.pyThe broad CI-style non-GPU test command is approximately:
pytest --timeout=3000 -W ignore::PendingDeprecationWarning -vv \
--cov-config=pyproject.toml --cov-report=xml --cov=xinference \
--ignore xinference/core/tests/test_continuous_batching.py \
--ignore xinference/model/image/tests/test_stable_diffusion.py \
--ignore xinference/model/image/tests/test_got_ocr2.py \
--ignore xinference/model/audio/tests \
--ignore xinference/model/embedding/tests/test_integrated_embedding.py \
--ignore xinference/model/llm/transformers/tests/test_tensorizer.py \
--ignore xinference/model/llm/tests/test_llm_model.py \
--ignore xinference/model/llm/vllm \
--ignore xinference/model/llm/sglang \
--ignore xinference/client/tests/test_client.py \
--ignore xinference/client/tests/test_async_client.py \
--ignore xinference/model/llm/mlx \
xinferenceUse narrower commands for daily development. Many model tests require large dependencies, GPU, Metal, network access, or model downloads.
The Web UI is under frontend and is a Next.js app (React, TypeScript,
Tailwind CSS), built as a static export and served by the Python backend from
xinference/ui/web/dist.
Common commands:
cd frontend
npm ci
npm run dev
npm run build
npx eslint .Use npm run format only when you intentionally want Prettier writes across
the frontend tree.
Documentation source is in doc/source.
Common docs dependencies are included in the doc extra:
pip install -e ".[doc]"
cd doc
make htmlWhen changing CLI behavior, API behavior, deployment behavior, or model support,
update the relevant documentation pages in doc/source.
Built-in model documentation under doc/source/models/builtin/ is generated by
doc/source/gen_docs.py; do not edit those generated files by hand. After
changing a built-in model registry or model_spec.json, run the generator from
doc/source and commit any resulting documentation changes:
cd doc/source
python gen_docs.pyThe PR workflow in .github/workflows/pr_auto_run_gen_docs.yaml only pushes
generated documentation for eligible same-repository chore/models-sync/*
branches. For other branches and fork PRs, run the generator explicitly rather
than assuming that the workflow will update the PR.
- Keep model-family logic inside the relevant
xinference/model/<family>/package. - Keep built-in model metadata changes close to existing specs and tests.
- Be careful with lazy imports and optional dependencies. Import heavyweight model libraries only where needed so unrelated installs still work.
- Preserve platform guards for Linux-only, CUDA-only, and macOS Metal/MLX paths.
- For distributed runtime changes, consider both local mode and supervisor/worker mode.
- For OpenAI-compatible behavior, verify request/response fields and streaming behavior against existing API and client tests.
The main CI workflow:
- Runs
pre-commit run --all-files. - Runs UI
npm ci,npx eslint ., and Prettier check. - Tests Python 3.10 through 3.13 across Linux, macOS, and Windows.
- Has special GPU and macOS Metal jobs for model-specific paths.
Before marking a change done, run the smallest meaningful validation command that covers the behavior you changed, and mention any broader checks that were not run because of environment cost or missing hardware.
- Keep commits scoped to the requested change.
- Do not rewrite or revert user changes in an existing worktree unless asked.
- If the active checkout is busy or on an unrelated branch, use a separate
worktree and a semantic branch name such as
fix/...,feat/..., ordocs/.... - In PR reviews, inspect current GitHub review threads before adding duplicate comments.