Build a Cloud GPU AI Workstation: Images, Video, LLMs & NSFW Limits (2026)
Rent a GPU for ComfyUI, video generation and SillyTavern. Compare costs, VRAM, setup steps, storage and NSFW restrictions before you pay.
A cloud GPU workstation lets an ordinary laptop run demanding AI workflows on a rented machine: ComfyUI for images and video, plus a language-model server for chat or SillyTavern. For occasional experiments, renting can avoid a large hardware purchase. The trade-off is setup work, ongoing storage charges and responsibility for securing the server.
What you are actually renting
Start with one NVIDIA GPU, persistent storage and one workload at a time. Your browser is the control panel; the model files and computation live on the remote machine. SillyTavern can stay on your own computer and connect to the remote language-model backend. A gaming GPU in your laptop is not required for that arrangement.
A rented GPU does not automatically include models, commercial rights, an unrestricted content policy or a desktop streaming environment. ComfyUI is a workflow interface; KoboldCpp runs the language model; SillyTavern supplies character cards, conversation history and role-play controls. Each part solves a different problem.
| Route | Best fit | Main trade-off |
|---|---|---|
| Managed image/chat service | Quick results with little maintenance | Provider models, quotas and content rules |
| Rented GPU + your software | Custom workflows, weights and repeatable experiments | You manage installation, storage and access |
| Your own GPU | Frequent use and local control | Upfront cost, power, cooling and limited VRAM |
Runpod or Vast.ai: choose an environment you can maintain
Runpod is the simpler starting point for this guide because it documents an official ComfyUI Pod template. Choose a regular on-demand Pod for interactive work, rather than assuming a Serverless endpoint behaves like a persistent workstation. The standard and Blackwell templates differ: use the matching one for the selected GPU.
Vast.ai is an alternative for readers comfortable evaluating individual machine offers. Check host reliability, location, available system RAM, storage, network rates and the rental terms together. An interruptible instance may suit recoverable batches; it is a poor default for a first interactive setup or an unsaved long render. There is no single universal Vast.ai hourly rate.
For either provider, inspect the final deployment summary before starting. A low GPU price cannot compensate for repeatedly downloading large weights, a slow disk or an incompatible CUDA environment. Start with a small balance and a short validation session.


Runpod · ComfyUI · Vast.ai · Pricing
How much VRAM and storage do you need?
These are planning ranges, not guarantees for every model. Memory demand changes with quantization, resolution, frame count, batch size, context length and offloading. Check the exact workflow before renting. System RAM is separate from VRAM: moving work to the CPU can avoid an error while making generation much slower.
For a first mixed workstation, 24 GB VRAM is a useful starting point. Reserve roughly 100–200 GB of disk space as an initial budget for selected weights, caches and outputs; a large model collection will need more. Do not load a video model and an LLM simultaneously just because they each fit separately.
| Workload | Practical starting point | What to watch |
|---|---|---|
| Basic image workflows | 12–24 GB VRAM | Extra ControlNet, editing and large batches need more |
| Wan2.2 TI2V-5B example | 24 GB with the documented offload settings | The official 720p example is not a guarantee for other video models |
| 7B–14B quantized chat model | 16–24 GB for a conservative first configuration | Leave room for context/KV cache; verify actual model size |
| Larger models or longer video | 48–80 GB may be appropriate | Match the exact model, not just its parameter count |
Costs: count the whole session, not just generation time
On September 29, 2026, Runpod’s public Pods pricing page lists RTX A5000 24 GB at US$0.27/hour, RTX 4090 24 GB at US$0.74/hour and RTX 5090 32 GB at US$0.99/hour. These are displayed reference prices, not a promise of stock or a final deployment quote. Do not substitute the separate Serverless rates.
Illustration: 20 running hours at US$0.74 plus a 100 GB standard network volume at US$0.07/GB/month gives US$21.80 before other disk charges, taxes or applicable transfer costs. Leaving the same GPU running for 720 hours would make the GPU portion alone US$532.80. Downloads, setup and idle thinking time count while the instance runs.
Budget = GPU hours × selected hourly rate + retained storage + applicable transfer + tax/payment fees. Runpod lists running volume-disk storage at US$0.10/GB/month and idle volume-disk storage at US$0.20; standard network volumes under 1 TB are listed at US$0.07. Choose the storage type deliberately.
For a purchase-versus-rental decision, divide the hardware cost you would actually pay by the rental cost for your monthly workload, then account for electricity, maintenance and resale value. Comparing an always-on cloud bill with a few hours of local use gives a misleading answer.
Runpod · Pricing · Runpod · Storage
Build the image workstation first
Use a small image job to verify the environment before installing video models and chat tools. A successful first run should produce a downloadable image, a saved workflow and a known model path. This is a setup guide, not a timed benchmark from a paid instance.
- Create a provider account, set a spending reminder and add your SSH public key if you will use a private tunnel. Choose a region reasonably close to you and confirm the GPU is available.
- Select the official Runpod ComfyUI template. For RTX 5090/B200 use its Blackwell Edition. Choose On-Demand and inspect the GPU, system RAM, storage and port settings before deployment.
- Attach persistent storage before downloading models. Confirm its mount path and put model files, workflows and outputs there. A directory called /workspace is useful only when it is actually backed by persistent storage.
- Wait for initialization and read the logs. Open ComfyUI through the configured connection. If using an HTTP proxy, check its authentication; a hard-to-guess URL is not access control. For private work, use an SSH tunnel and remove unnecessary public HTTP ports.
- Load a model-specific template from Workflow → Browse Templates, or import the workflow from official ComfyUI documentation. Install the exact checkpoint, text encoders and VAE it requires on the remote machine. A browser download may save to your laptop instead.
- Start with one image and the template’s default dimensions. Enter a simple prompt, run the graph and save both the result and workflow JSON. Verify where the files were written before stopping the Pod.
First image test
A ceramic teapot on a wooden table beside a window, soft morning light, neutral background, realistic product photograph.
Runpod · ComfyUI · ComfyUI · Text to Image · Runpod · SSH
If you choose a network volume, select it while creating the Pod; Runpod does not attach it later to an existing Pod. For ComfyUI over full SSH, the tunnel below assumes the SSH server can reach ComfyUI at 127.0.0.1:8188. Replace the connection placeholders, keep the terminal open, then visit http://127.0.0.1:8188 on this computer. The ComfyUI service must already be running; do not start a second copy on the same port.
ssh -N -L 127.0.0.1:8188:127.0.0.1:8188 -p SSH_PORT -i KEY_PATH USER@POD_IP Add video with a workflow that matches the model
For a concrete downloadable example, ComfyUI provides Wan2.2 workflows and model links. The Wan2.2 TI2V-5B official repository documents a 720p path on at least 24 GB VRAM with model offloading and the text encoder on CPU. That is a useful baseline, not a claim that every 14B video workflow fits a 24 GB card.
Open the TI2V-5B example, download every component into the folder named in the guide, and keep the graph’s supported dimensions and frame settings for the first test. Try one short clip before a long queue. For image-to-video, use an image you may upload and describe motion, not just the still scene.
Save the seed, workflow, component filenames and output together. Compare warm runs with identical settings if you evaluate GPUs; initial loading and model downloads distort comparisons. Partner/API nodes in ComfyUI can still call a separately billed hosted service. They are not automatically running on your rented GPU.
ComfyUI also documents an optimized native-offloading route for the 5B model at 8 GB VRAM. That is a different configuration from the upstream 24 GB example; lower memory use is not a promise of equal speed.
Image-to-video motion prompt
The camera slowly moves toward the teapot. Gentle steam rises while the table and background remain still. Soft, consistent daylight.

Run an LLM backend, then connect SillyTavern
Keep SillyTavern on your own computer and rent only the model backend. This separates your character-card library from a disposable GPU instance. It does not make remote inference local or invisible to the hosting infrastructure.
On a CUDA-compatible Linux instance, download the matching Linux build from the official KoboldCpp releases and a supported GGUF model from its publisher or a traceable quantization repository. Check the model license and architecture support. The example below assumes you renamed the executable to koboldcpp-linux-x64 and saved the chosen model as /workspace/models/model.gguf. Start with 4,096 context tokens and automatic GPU-layer fitting; inspect the logs to confirm GPU use.
chmod +x ./koboldcpp-linux-x64 ./koboldcpp-linux-x64 --model /workspace/models/model.gguf --usecuda --gpulayers -1 --contextsize 4096 --host 127.0.0.1 --port 5001 The server listens only on its own loopback interface. On your computer, take the provider’s full SSH/TCP connection details and adapt the tunnel command below. Replace SSH_PORT, KEY_PATH, USER and POD_IP; they are placeholders, not literal values. The SSH server and KoboldCpp must share the network namespace for this loopback destination.
ssh -N -L 127.0.0.1:5001:127.0.0.1:5001 -p SSH_PORT -i KEY_PATH USER@POD_IP - Keep the SSH terminal open. Visit http://127.0.0.1:5001 on your computer and test a short message in KoboldCpp’s interface. If that fails, fix the backend or tunnel before changing SillyTavern.
- Install SillyTavern from its official instructions. On Windows, install Node.js LTS and Git, clone the release branch into a user-writable folder and run Start.bat without administrator privileges.
- In SillyTavern, open API Connections, choose Text Completion → KoboldCpp, enter http://127.0.0.1:5001/ and click Connect. This route uses the base URL, not an OpenAI /v1 endpoint.
- Create a simple character and send a short message. Match the prompt template to the model and keep SillyTavern’s context limit within the backend setting. Back up character cards and chats locally.


KoboldCpp · Releases · KoboldCpp · SillyTavern · Windows · SillyTavern · API · Runpod · SSH
NSFW: flexibility is not permission
Self-hosting gives you more control over model selection and workflow components, but no universal NSFW switch. A model may refuse adult prompts, generate them poorly or have license restrictions. Character cards do not override the model, an API provider or your GPU host.
Runpod’s terms restrict obscene and lewd content; Vast.ai’s terms also prohibit using its services to store or transmit obscene content. Neither should be described as an unrestricted adult-content hosting recommendation. For a lawful adult use case, obtain an explicit answer from the provider before paying; model availability alone is not permission.
Evaluate four layers separately: host rules, model/license terms, any external API used by the workflow, and the rules of the platform where you publish. A fully local computer removes the cloud-host layer but not the other obligations. This guide does not recommend a bypass for provider restrictions.
| Question | Practical assessment |
|---|---|
| Can ComfyUI or SillyTavern guarantee NSFW output? | No. They are interfaces; behavior depends on the backend and model. |
| Does open-weight mean unrestricted? | No. Read the exact license and acceptable-use terms. |
| Is rented GPU inference fully private? | No. It runs on someone else’s infrastructure; protect access and avoid unnecessary sensitive uploads. |
Runpod · Terms · Vast.ai · Terms
Back up, shut down and troubleshoot
Keep a portable project bundle: workflow JSON, custom-node versions, model filenames and licenses, seeds, prompts, character cards and selected outputs. Do not include API keys in a shared workflow. Download this bundle before deleting any instance.
Stopping is different from terminating. Runpod container-disk data can be lost on stop; a volume disk persists through stop but is removed with Pod termination. A network volume has its own lifecycle and bill. Verify the provider’s current behavior, stop compute, then separately review retained storage. Restarting a stopped Pod does not guarantee the same GPU capacity remains available.
| Problem | First useful check |
|---|---|
| Out of memory | Unload other models; reduce batch, context or video frames; then consider quantization/offload or more VRAM. |
| Missing nodes/models | Use the intended workflow version and exact component paths; avoid installing arbitrary node packs. |
| SillyTavern cannot connect | Test the backend locally on the server, then the SSH tunnel, then the frontend URL. |
| Generation is unexpectedly slow | Check GPU utilization, CPU offload, initial loading and disk/network bottlenecks. |
| The bill keeps growing | Closing the browser does not stop the instance; check compute and all retained volumes. |
FAQ and verdict
Do I need a powerful computer at home?
No for this browser-and-remote-backend arrangement. You still need a stable connection, enough local storage for backups and a computer capable of running the frontend. Upload and download speed matters for large video files.
Can all four workloads run on one GPU?
They can share a machine sequentially when each fits. Running image, video and chat models together may exhaust VRAM. Unload or stop one backend before starting another.
Will two 24 GB GPUs behave like one 48 GB GPU?
Not automatically. The application must support splitting the model, and communication overhead matters. Choose one sufficiently large GPU for the simplest first setup.
Can I connect from a phone?
A browser can control a secured web interface, but the localhost tunnel shown here belongs to the computer that runs it. Phone access needs its own secure network arrangement; do not expose an unauthenticated service to make it convenient.
Can I earn money from generated images or video?
That depends on the model license, input rights, hosting terms and the use of the output. Paying for GPU time does not buy commercial rights to every model.
Is a GPU rental cheaper than a subscription?
It can be for concentrated occasional use, but include setup, storage and idle hours. If you only need a few finished images and dislike maintenance, a managed service may be better value.
For a first workstation, build one reproducible image workflow on a 24 GB GPU, confirm storage survives the intended stop/restart cycle, and measure your own cost per useful result. Add video only after that baseline works.
For a private character-chat setup, keep SillyTavern and its backups on your computer and connect it to a secured remote backend. Choose larger hardware only when the exact workload needs it—and treat NSFW hosting permission as a separate decision from technical capability.