Sandbox Setup
This guide walks you through setting up a local Michelangelo AI environment on your laptop. The sandbox runs a fully functional cluster — API server, controller manager, workflow engine, object storage, and supporting services — entirely on your machine, so you can explore Michelangelo AI or develop against it without any cloud infrastructure.
Who this is for: ML engineers, platform engineers, and contributors who want to try Michelangelo AI locally or develop new features against it.
What you'll have at the end: a running sandbox cluster and a successful demo pipeline run, ready for you to build your own workflows on top of.
Time estimate: 30–60 minutes on first run (image pulls for all services can take 20–40 minutes depending on your connection). Subsequent recreates with cached images typically take 5–10 minutes.
Supported platforms: macOS (Apple Silicon and Intel) and Linux. Windows is not officially supported, but WSL2 with Docker Desktop should work for most steps.
Get the code
Clone the Michelangelo AI repository to your machine:
git clone https://github.com/michelangelo-ai/michelangelo.git
cd michelangelo
Throughout this guide, <repo-root> refers to the directory you just cloned (for example, ~/michelangelo).
Prerequisites
Before you begin, make sure you have the following installed. Install commands below show macOS (Homebrew) and Linux options where they differ; on Linux, follow the linked official guide if you don't see a direct command. Run each verification command to confirm:
| Tool | Install (macOS) | Install (Linux) | Verify |
|---|---|---|---|
| Docker | Docker Desktop or Colima | Docker Engine | docker info |
| docker buildx | brew install docker-buildx (see note below) | usually bundled with Docker Engine | docker buildx version |
| kubectl | brew install kubectl | official guide | kubectl version --client |
| k3d | brew install k3d | official guide | k3d --version |
| Helm | brew install helm | official guide | helm version |
| Python 3.11 or 3.12 | python.org or brew install python@3.11 | distro package manager (e.g., apt install python3.11) | python3 --version |
| Poetry | curl -sSL https://install.python-poetry.org | python3 - | same as macOS | poetry --version |
| temporal (Temporal only) | brew install temporal | official guide | temporal --version |
Python version note: Python 3.11 or 3.12 is strongly recommended. Python 3.13+ may fail during
poetry installbecause pre-built wheels for some ML dependencies are not yet available for newer interpreter versions.
Docker daemon note:
docker info(the verify command above) requires the Docker daemon to be running — unlikedocker --version, which only checks the binary. Ifdocker infofails, start Docker Desktop or Colima before continuing.
buildx note (Homebrew):
brew install dockerinstalls the Docker CLI only — the buildx plugin is a separate formula. Without it,scripts/kuberay/build-kuberay-images.shfails withunknown flag: --platform. After installing, link the plugin so the CLI can find it:mkdir -p ~/.docker/cli-plugins
ln -sfn /opt/homebrew/opt/docker-buildx/bin/docker-buildx ~/.docker/cli-plugins/docker-buildx
Poetry install note (macOS, python.org builds): if the Poetry install command aborts with
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED], your interpreter has no CA bundle wired intossl— python.org framework builds ship one but do not install it. Run the bundled installer once, then retry:"/Applications/Python 3.11/Install Certificates.command"Alternatively,
brew install poetryavoids the issue entirely.
Colima resource requirements (macOS only)
If you are on macOS and using Colima as your Docker runtime, the default VM resources are too limited for the sandbox. Start Colima with at least:
colima start --cpu 6 --memory 12 --disk 100
| Resource | Minimum | Recommended |
|---|---|---|
| CPU cores | 4 | 6 |
| Memory (GiB) | 12 | 12 |
| Disk (GB) | 60 | 100 |
Warning: Starting Colima with the default settings (2 CPU, 2 GB RAM) will cause pods to crash or fail to schedule. Always pass explicit resource flags.
Memory is not a soft limit.
--memory 8yieldsTotal Memory: 7.737GiB(the flag is GiB), and at that size thebert_colaexample inpython/examples/fails: its data-prep task succeeds, then the training task is killed by Ray's memory monitor withTask was killed due to the node running low on memory ... 7.35GB / 7.74GB (0.950357), which exceeds the memory usage threshold of 0.95. The same pipeline succeeds unchanged at--memory 12. Because the VM needs 12 GiB and macOS needs the rest, a 16 GB host is effectively the floor for running the bundled examples.
Apple Silicon: add
--vz-rosettato yourcolima startcommand. Colima defaults torosetta: false, and some published images run amd64 binaries that fall back to QEMU emulation, where the Go runtime can die withfatal error: fault— typically seen asmichelangelo-apiserverinCrashLoopBackOff.
If Colima is already running with insufficient resources, stop it and restart with the new settings:
colima stop
colima start --cpu 6 --memory 12 --disk 100
Configure host.docker.internal
Docker containers need to communicate with services on your host machine. Verify this hostname resolves correctly:
- Open your hosts file:
sudo nano /etc/hosts - Look for this line:
127.0.0.1 host.docker.internal - If missing, add it to the end of the file and save.
Install Python dependencies
From the repository you cloned, install the Michelangelo AI Python packages:
cd <repo-root>/python
poetry install
Quick start
Once prerequisites and Python dependencies are installed:
# 1. Activate the Poetry virtual environment (from <repo-root>/python)
source .venv/bin/activate
# 2. Build local-only images required by the sandbox
cd <repo-root>
bash scripts/kuberay/build-kuberay-images.sh
# 3. Create the sandbox (30–60 min on first run; images are cached after)
ma sandbox create
# 4. Verify everything works by running the demo pipeline
ma sandbox demo pipeline
Note on step 2: the build script also imports the images it builds into your k3d clusters, but those clusters do not exist until step 3. It will report that cluster
michelangelo-compute-0was not found and skip it, which is expected. Ifhistory-serverlater sits inImagePullBackOff, re-run the script afterma sandbox create— the images are already built, so the second run only performs the import.
Tip: If you prefer not to activate the venv, prefix each
macommand withpoetry run(e.g.,poetry run ma sandbox create). If you seezsh: command not found: ma, you either skipped step 1 or need to usepoetry run. See troubleshooting below.
Note on image pull times: During
ma sandbox create, several pods will sit inContainerCreatingfor several minutes while images are pulled (some images are 500–800 MB). This is normal. The relevant question is whether pods are making progress — check withkubectl get pods -wand look for status transitions. Only act if a pod stays inImagePullBackOfforCrashLoopBackOfffor more than a few minutes.
Choosing a workflow engine
ma sandbox create defaults to Cadence, which is the recommended choice for most users — it's the most-tested path and matches the examples in this guide. Pass --workflow temporal only if you specifically want to develop or test against Temporal (for example, if your team is migrating to it). The two engines are interchangeable from a workflow-author perspective; the choice mainly affects which web UI and CLI you use.
Verifying success
When ma sandbox create completes, all Michelangelo AI services start in your k3d cluster. Verify with:
kubectl get pods
You should see roughly 10–15 pods (the exact count depends on which engine you chose and any --exclude flags). On first run, pods may spend 5–10 minutes in ContainerCreating while images are pulled — this is expected. All pods should eventually reach Running status; if any remain in ContainerCreating beyond 15 minutes, check the pod events with kubectl describe pod <pod-name>.
Then open the Michelangelo AI UI at http://localhost:8090 — if the dashboard loads, your sandbox is healthy. See Sandbox Ports and Endpoints for the full list of services and their URLs.
Sandbox commands
The ma sandbox command manages your local Kubernetes development environment.
For a complete command reference, see the CLI Reference - Sandbox Commands.
Lifecycle
The typical sandbox workflow:
create → (develop) → stop → start → (develop) → delete
↓ (if create fails partway)
sync → (develop) → ...
Create
ma sandbox create [OPTIONS]
| Flag | Description | Default |
|---|---|---|
--workflow cadence|temporal | Choose workflow engine | cadence |
--exclude [services] | Exclude services: apiserver, controllermgr, ui, worker, prometheus, grafana | none |
--create-compute-cluster | Create an additional Ray compute cluster for distributed jobs | disabled |
--compute-cluster-name <name> | Custom name for the compute cluster | auto-generated |
--include-experimental [services] | Include experimental services | none |
Examples:
# Full sandbox with all services (default: Cadence workflow engine)
ma sandbox create
# Sandbox with Temporal workflow engine
ma sandbox create --workflow temporal
# Sandbox without UI, with a Ray compute cluster
ma sandbox create --exclude ui --create-compute-cluster
Sync
ma sandbox sync
Redeploys services into an existing cluster, skipping cluster creation and image import. Use this to recover from a partially failed create — for example, if operator deployments timed out but the cluster itself was created successfully.
If ma sandbox create fails and you see Failed to create cluster ... already exists when you try again, run ma sandbox sync instead. See Recovering from a failed create below.
Stop / Start
Pause and resume your sandbox without losing state:
ma sandbox stop # preserves state
ma sandbox start # resume where you left off
Delete
Tear down the cluster and remove all resources:
ma sandbox delete
Demo
Create pre-configured demo resources for testing:
ma sandbox demo pipeline # registers and runs a sample pipeline
ma sandbox demo inference # sets up demo inference server
Recovering from a failed create
If ma sandbox create fails partway through — for example, with a Helm timeout on kuberay-operator or spark-operator — the k3d cluster may already exist even though not all services deployed successfully.
Running ma sandbox create again will fail with Failed to create cluster ... because a cluster with that name already exists.
Do not delete and recreate — that discards the cluster and costs another full image-pull cycle. Instead, use sync:
ma sandbox sync
sync redeploys services into the existing cluster without recreating it or re-importing images. In most cases this resolves transient operator timeouts without the 30–60 minute penalty of a fresh create.
If sync doesn't resolve the issue, check pod events with kubectl describe pod <pod-name> to understand what failed, then delete and recreate as a last resort:
ma sandbox delete
ma sandbox create
Smoke test: run the BERT CoLA example
After ma sandbox demo pipeline succeeds, you've already proven the sandbox works end to end. If you'd like to run a real example workflow against it before moving on, the BERT CoLA text-classification example is a quick way to confirm local execution works:
cd <repo-root>/python
poetry install --extras example
PYTHONPATH=. poetry run python ./examples/bert_cola/bert_cola.py
You should see workflow logs in your terminal and, when it finishes, a trained model artifact written to local storage.
For the full story on local vs. remote execution, building Docker images, configuring storage, and using either workflow engine end to end, see:
- Pipeline Running Modes — the four execution modes Michelangelo AI supports
Note: Local execution doesn't support caching, retries, or resource constraints. Use remote execution (covered in the ML Pipelines guides) for production-like behavior.
Optional: set a custom dev identity
The UI shows a placeholder "Local Developer" user (name, email, and avatar) in the nav bar. To swap in your own GitHub name, email, and photo, visit the sandbox URL with a ?ghUser=<your-github-username> query param, e.g. http://localhost:8090/?ghUser=<username> (optionally add &email=<your-email> to set a specific email). On that first visit, the app fetches your public GitHub profile (no auth needed) and caches your name, email, and avatar in your browser's localStorage, so no UI rebuild or backend restart is needed — a normal page reload afterward keeps showing them.
yarn set-avatar <your-github-username> [--email <your-email>] (run from javascript/) is a shortcut that prints this URL for you; editing the URL directly works the same way. Most GitHub profiles don't expose a public email, so pass --email (or add &email= to the URL yourself) if you want a real one shown — otherwise it falls back to your local git email (only works against yarn dev, not the built sandbox) or a placeholder.
To go back to the default placeholder identity, use Sign out from the user menu — in the sandbox it just clears the cached profile and strips ghUser/email from the URL (there's no real auth here), so a manual localStorage edit alone won't stick if those params are still in the URL.
Troubleshooting
command not found: ma
You either skipped the venv activation step or the Poetry environment isn't active in your current shell. Two options:
# Option 1: activate the venv (from <repo-root>/python)
source .venv/bin/activate
# Option 2: prefix each command with poetry run
poetry run ma sandbox create
If neither works, run poetry install from <repo-root>/python first to make sure the environment was created.
Failed to create cluster ... already exists
The k3d cluster was created but ma sandbox create failed before all services deployed. Do not delete and recreate — use sync to redeploy services into the existing cluster without re-importing images:
ma sandbox sync
See Recovering from a failed create for the full recovery flow.
history-server stuck in ImagePullBackOff
The kuberay-historyserver image is not available in any public registry — it must be built locally before running ma sandbox create. If you skipped the build step:
cd <repo-root>
bash scripts/kuberay/build-kuberay-images.sh
ma sandbox sync # redeploy without recreating the cluster
Progress deadline exceeded from operator installs
During ma sandbox create, Helm deploys operators (kuberay, spark) with a timeout. On slow connections or underpowered machines, the timeout can fire before the operator pod finishes pulling its image — even though the deployment will succeed on its own a few minutes later.
This is a first-run phenomenon caused by large image pulls racing against Helm's deadline. On subsequent create or sync runs with cached images, operator deployments complete well within the timeout.
If you see Progress deadline exceeded in the output, verify that the operators are still coming up:
kubectl get pods -A
Within a few minutes of the Helm error, you should see the operator deployments transition to Running. If they do, the sandbox is healthy despite the error message — run ma sandbox sync to finish deploying any remaining services. Only proceed to ma sandbox delete if pods are stuck in ImagePullBackOff or CrashLoopBackOff with no sign of progress.
Grafana pod in CrashLoopBackOff
Grafana's sandbox configuration installs dashboard plugins at startup by fetching them from an external source. On restricted, proxied, or intermittent network connections this fetch can fail, causing a crash loop.
Check if this is the cause:
kubectl logs <grafana-pod-name> | grep -i "plugin\|install\|GF_INSTALL"
If plugin installation is the failure point, you can exclude Grafana from the sandbox and proceed without it:
ma sandbox sync --exclude grafana
Grafana is used for metrics dashboards and is not required for pipeline execution or the Michelangelo AI UI.
ma CLI commands fail with connection attempt timed out
The ma CLI connects to the API server over gRPC at localhost:15566. The sandbox maps this port via k3d NodePort so it works out of the box after ma sandbox create.
If ma commands time out, verify the NodePort mapping is active:
kubectl get svc michelangelo-apiserver -o jsonpath='{.spec.type}'
# Should print "NodePort"
The default CLI configuration (~/.ma/config.toml) uses address = "127.0.0.1:15566".
ModuleNotFoundError: No module named 'grpc_reflection'
This error occurs when Python dependencies aren't fully installed. Fix it by reinstalling from the python/ directory:
cd <repo-root>/python
poetry install
If the error persists, try removing the virtual environment and reinstalling:
rm -rf .venv
poetry install
Pods stuck in ImagePullBackOff or ErrImagePull
The cluster can't pull a Docker image. Check which image is failing:
kubectl describe pod <pod-name> | grep -A 5 "Events"
Common causes:
- Network issues: Ensure Docker can reach
ghcr.io(trydocker pull ghcr.io/michelangelo-ai/worker:latest) - Local-only image not built: If
history-serveris the failing pod, seehistory-serverstuck inImagePullBackOffabove
Worker crashes with Namespace default is not found (Temporal only)
The Temporal default namespace must be registered after the sandbox starts. If the worker is in CrashLoopBackOff:
# Port-forward the Temporal frontend
kubectl port-forward svc/michelangelo-temporal-frontend 7233:7233 &
# Register the default namespace
temporal operator namespace create default
# Restart the worker to pick it up
kubectl rollout restart deployment/michelangelo-worker
Pods stuck in CrashLoopBackOff
A service is starting but immediately crashing. Check its logs:
kubectl logs <pod-name>
Before escalating, rule out the two causes that deleting and recreating will not fix:
- Architecture mismatch. If the log ends in
fatal error: faultor aruntime.sigpanicstack trace, the container is running a binary built for another architecture under emulation. Compare the binary against the node:kubectl get nodes -o jsonpath='{.items[*].status.nodeInfo.architecture}', then extract the image's entrypoint and check it withfile. On Apple Silicon, start Colima with--vz-rosetta. - Memory. If the pod is OOM-killed (
kubectl describe pod <pod-name>showsReason: OOMKilled, or Ray reportsTask was killed due to the node running low on memory), the VM is too small — see Colima resource requirements.
If neither applies, the pod may simply be wedged. The simplest recovery is to delete it and let Kubernetes recreate it:
kubectl delete pod <pod-name>
If that doesn't help, recreate the sandbox cleanly:
ma sandbox delete
ma sandbox create
Port already in use
If ma sandbox create fails because a port is already bound:
# Find what's using the port (e.g., port 9090)
lsof -i :9090
# Kill the process if it's safe to do so
kill <PID>
See Sandbox Ports and Endpoints for the full list of ports used.
Poetry install fails with build errors on macOS
If you see C++ compilation errors during poetry install:
export CC=clang
export CXX=clang++
poetry install
Add those exports to your ~/.zshrc to make them permanent.
What's next?
- Build your first pipeline -- Follow Getting Started with ML Pipelines to create a training workflow (~30 min)
- Explore example projects -- Try California Housing XGBoost (
xgb_train) or California Housing PyTorch Lightning (pytorch_train) in michelangelo-examples, BERT Text Classification, or GPT Fine-tuning - Learn the CLI -- See the CLI Reference for managing pipelines and projects