Last Updated: October 7, 2026
Before a new instance can do any work, SaladCloud allocates it to a node, the node downloads your image, and your
container starts. Each node downloads at its own speed, and a large image can take a while to download on a node that
hasn’t pulled it before. You aren’t billed while an instance is allocating or downloading (see
Billing), but you do wait, and while you are still getting a
deployment right, every change can mean another wait. The habits below keep that loop short.
Prove the plumbing with a small image first
Most problems in a first deployment are in the configuration around your application: a port or path that doesn’t match,
a probe that never passes, a job document in the wrong shape, missing registry credentials. A small image shows these as
well as your real one does, and it downloads quickly.
- Deploy a small image that answers HTTP on the same port and paths your application will use, including the path your
readiness probe checks.
- Check everything around it: the
queue connection or
Container Gateway, the
health probes, the job or request your client
sends and the response it gets back, and scaling.
- When that works, update the container group to
your real image.
Debug startup by hand
If your real image fails to start or exits early, you don’t need to rebuild and redeploy for every guess. Keep the
container alive with a command that only sleeps, connect to it, and run the startup yourself:
- Set the container group’s command to
sleep 2147483647. In
the API, that is "command": ["sleep", "2147483647"]. In the portal, enter sleep as the command and 2147483647
as its argument. The command replaces the image’s ENTRYPOINT and CMD, so the container starts and does nothing
else.
- Remove the startup and liveness probes for now. Your application isn’t running, so they would fail, and a failed
startup or liveness probe moves the instance to a new node.
- When an instance is running, connect to it with
SSH or the web terminal.
- Run the image’s original entrypoint by hand and watch what happens. Check that the expected files, binaries and
environment variables are there, that the GPU is visible (
nvidia-smi on NVIDIA), and that your application listens
on the port you configured.
- When you have found the fix, build it into the image, remove the command, restore the probes, and deploy again.
sleep infinity doesn’t work with every sleep. Some BusyBox builds reject it, and the container exits straight
away. A large number of seconds works everywhere: 2147483647 seconds is about 68 years.
If your image also runs the Job Queue Worker, the sleep
command replaces that too, so the instance takes no jobs. If you start the worker by hand, it takes jobs from the queue
as usual.
Run a few spare replicas while testing
Some nodes start your container sooner than others. While you test, run 3 replicas and use whichever instance is running
first. Instances are billed only while they are running, so the others cost nothing while they are still allocating or
downloading. When you are done, scale down or stop the container group so the spares don’t keep running.
Protect busy instances when scaling down
When you lower the replica count, or the Job Queue autoscaler does, SaladCloud chooses which instances to stop. Give an
instance that is working a higher
instance deletion cost so that idle
instances are stopped first. The deletion cost applies only to scaling down. It doesn’t protect an instance from a node
interruption.
Test first, then attach the Job Queue autoscaler
The Job Queue autoscaler sets the replica count
itself, from the queue length, so it overrides any spare replicas you set for testing. Create the container group with a
fixed replica count, get jobs running end to end with the steps above, and then add queue_autoscaler with an update.
Enable Autoscaling shows the update request.
Related pages