Qwen3 0.6B
0.6B · Apache-2.0 · min. VRAM 6 GB
Tiny model for classification, routing and tests.
Recommended: RTX 3060· from $0.08/hr
Model library
Open models for text, code, vision, images, video, speech and search. Pick a model — we suggest the cheapest GPU that fits.
0.6B · Apache-2.0 · min. VRAM 6 GB
Tiny model for classification, routing and tests.
Recommended: RTX 3060· from $0.08/hr
1.7B · Apache-2.0 · min. VRAM 8 GB
Small and fast, good for simple chat and extraction.
Recommended: RTX 3060· from $0.08/hr
4B · Apache-2.0 · min. VRAM 12 GB
Best quality per gigabyte for cheap GPUs.
Recommended: RTX 3060· from $0.08/hr
3B · Apache-2.0 · min. VRAM 10 GB
Open 3B model with long context and reasoning mode.
Recommended: RTX 3060· from $0.08/hr
3.8B · MIT · min. VRAM 12 GB
Compact Microsoft model strong at math and logic.
Recommended: RTX 3060· from $0.08/hr
3B · Llama 3.2 · min. VRAM 12 GB
Meta's small instruct model (needs your HF token).
Recommended: RTX 3060· from $0.08/hr
4B · Gemma · min. VRAM 12 GB
Google's small multimodal model (needs your HF token).
Recommended: RTX 3060· from $0.08/hr
8B · Apache-2.0 · min. VRAM 24 GB
Strong all-rounder with thinking mode, 100+ languages.
Recommended: 2× RTX 3060· from $0.16/hr
21B MoE · Apache-2.0 · min. VRAM 24 GB
OpenAI open-weight reasoning model with tool use.
Recommended: 2× RTX 3060· from $0.16/hr
7B · Apache-2.0 · min. VRAM 24 GB
Proven chat model for assistants and RAG.
Recommended: 2× RTX 3060· from $0.16/hr
7B · Apache-2.0 · min. VRAM 24 GB
Code completion and generation for IDE assistants.
Recommended: 2× RTX 3060· from $0.16/hr
7B · Apache-2.0 · min. VRAM 24 GB
Understands images, documents and screenshots.
Recommended: 2× RTX 3060· from $0.16/hr
7B · MIT · min. VRAM 24 GB
R1-style step-by-step reasoning in a small model.
Recommended: 2× RTX 3060· from $0.16/hr
8B · MIT · min. VRAM 24 GB
R1 reasoning distilled into Llama 8B.
Recommended: 2× RTX 3060· from $0.16/hr
7B · Apache-2.0 · min. VRAM 24 GB
Fast European model with function calling.
Recommended: 2× RTX 3060· from $0.16/hr
8B · Llama 3.1 · min. VRAM 24 GB
Meta's popular 8B model (needs your HF token).
Recommended: 2× RTX 3060· from $0.16/hr
14B · Apache-2.0 · min. VRAM 40 GB
Noticeably smarter than 8B, still affordable.
Recommended: 4× RTX 3060· from $0.32/hr
14B · MIT · min. VRAM 40 GB
Excellent at math, code and structured answers.
Recommended: 4× RTX 3060· from $0.32/hr
14B · MIT · min. VRAM 40 GB
Deeper reasoning than the 7B distill.
Recommended: 4× RTX 3060· from $0.32/hr
12B · Apache-2.0 · min. VRAM 40 GB
128k context, strong multilingual chat.
Recommended: 4× RTX 3060· from $0.32/hr
12B · Gemma · min. VRAM 40 GB
Multimodal Gemma (needs your HF token).
Recommended: 4× RTX 3060· from $0.32/hr
32B · Apache-2.0 · min. VRAM 80 GB
Flagship dense Qwen3 for demanding assistants.
Recommended: 8× RTX 3060· from $0.64/hr
30B MoE · Apache-2.0 · min. VRAM 80 GB
MoE with only 3B active: big-model quality at high speed.
Recommended: 8× RTX 3060· from $0.64/hr
30B MoE · Apache-2.0 · min. VRAM 80 GB
Agentic coding model for Cline, Roo, Continue.
Recommended: 8× RTX 3060· from $0.64/hr
32B · Apache-2.0 · min. VRAM 80 GB
Top open dense coding model.
Recommended: 8× RTX 3060· from $0.64/hr
32B · Apache-2.0 · min. VRAM 80 GB
Strong vision-language model for documents and charts.
Recommended: 8× RTX 3060· from $0.64/hr
32B · MIT · min. VRAM 80 GB
The strongest R1 distill.
Recommended: 8× RTX 3060· from $0.64/hr
24B · Apache-2.0 · min. VRAM 80 GB
Balanced 24B with vision and function calling.
Recommended: 8× RTX 3060· from $0.64/hr
117B MoE · Apache-2.0 · min. VRAM 80 GB
OpenAI's largest open model on a single 80 GB GPU.
Recommended: 8× RTX 3060· from $0.64/hr
27B · Gemma · min. VRAM 80 GB
Largest Gemma 3 (needs your HF token).
Recommended: 8× RTX 3060· from $0.64/hr
70B · Llama 3.3 · min. VRAM 160 GB
GPT-4-class open model; 2×80 GB (needs your HF token).
Recommended: 4× A40· from $1.60/hr
106B MoE · MIT · min. VRAM 320 GB
Agent-oriented MoE; 4×80 GB.
Recommended: 8× A40· from $3.20/hr
235B MoE · Apache-2.0 · min. VRAM 320 GB
Frontier-level open model; 4×80 GB (Hopper, FP8).
Recommended: 8× A40· from $3.20/hr
480B MoE · Apache-2.0 · min. VRAM 640 GB
Best open agentic coder; 8×80 GB (Hopper, FP8).
Recommended: 8× A100 80GB· from $8.80/hr
671B MoE · MIT · min. VRAM 1000 GB
Full DeepSeek-R1; 8×H200.
Recommended: 8× H200 141GB· from $20.80/hr
671B MoE · MIT · min. VRAM 1000 GB
Full DeepSeek-V3.1 hybrid thinking model; 8×H200.
Recommended: 8× H200 141GB· from $20.80/hr
0.6B · Apache-2.0 · min. VRAM 6 GB
Fast multilingual embeddings for search and RAG.
Recommended: RTX 3060· from $0.08/hr
8B · Apache-2.0 · min. VRAM 24 GB
Top-quality embeddings (MTEB leader class).
Recommended: 2× RTX 3060· from $0.16/hr
568M · MIT · min. VRAM 6 GB
Dense + sparse multilingual embeddings.
Recommended: RTX 3060· from $0.08/hr
560M · MIT · min. VRAM 6 GB
Classic multilingual retrieval embeddings.
Recommended: RTX 3060· from $0.08/hr
568M · Apache-2.0 · min. VRAM 6 GB
Reranker for better RAG results (/v1/rerank).
Recommended: RTX 3060· from $0.08/hr
3.5B · OpenRAIL++ · min. VRAM 8 GB
Classic SDXL: huge ecosystem of LoRAs, runs on 8 GB.
Recommended: RTX 3060· from $0.08/hr
3.5B · CreativeML OpenRAIL-M · min. VRAM 8 GB
Photorealistic SDXL fine-tune, great for portraits.
Recommended: RTX 3060· from $0.08/hr
2.6B · Apache-2.0 · min. VRAM 12 GB
Efficient text-to-image with good prompt following.
Recommended: RTX 3060· from $0.08/hr
12B · Apache-2.0 · min. VRAM 16 GB
Top-quality images in 4 steps, commercial use allowed.
Recommended: RTX 4060 Ti 16GB· from $0.12/hr
17B · MIT · min. VRAM 24 GB
State-of-the-art open image model, MIT license.
Recommended: 2× RTX 3060· from $0.16/hr
20B · Apache-2.0 · min. VRAM 40 GB
Best-in-class text rendering in images (posters, UI, Chinese/English).
Recommended: 4× RTX 3060· from $0.32/hr
1.3B · Apache-2.0 · min. VRAM 12 GB
Video generation even on 12 GB cards (480p).
Recommended: RTX 3060· from $0.08/hr
5B · Apache-2.0 · min. VRAM 24 GB
720p text- and image-to-video on a single 24 GB GPU.
Recommended: 2× RTX 3060· from $0.16/hr
14B MoE · Apache-2.0 · min. VRAM 40 GB
Cinematic 720p video; 4-step LoRAs included for speed.
Recommended: 4× RTX 3060· from $0.32/hr
14B MoE · Apache-2.0 · min. VRAM 40 GB
Animate any image into a 720p clip.
Recommended: 4× RTX 3060· from $0.32/hr
10B · Apache-2.0 · min. VRAM 24 GB
Smooth-motion text-to-video from Genmo.
Recommended: 2× RTX 3060· from $0.16/hr
min. VRAM 6 GB
faster-whisper transcription and text-to-speech behind an OpenAI-compatible audio API.
min. VRAM 8 GB
Speech recognition service (faster-whisper, large-v3) with a Swagger UI.
min. VRAM 4 GB
Fast, natural text-to-speech (82M) with an OpenAI-compatible /v1/audio/speech endpoint.
No models match these filters.
Rent 2, 4 or 8 GPUs in one server and use them as one:
vLLM and SGLang split one model across all GPUs of the server, so their memory adds up: Llama 3.3 70B on 2×A100 80GB, Qwen3 235B on 4×H100, DeepSeek-R1 671B on 8×H200.
PyTorch templates see every GPU: use torchrun with DDP, FSDP or DeepSpeed ZeRO to train faster or fit larger models.
GPUs are combined inside one machine over NVLink/PCIe. Splitting a model across computers over the internet would be tens of times slower, so we never do it.
Templates
Every server starts from a template. Templates with a model slot load the model automatically; others give you a clean environment over SSH.
vLLM · OpenAI API
High-throughput LLM server with an OpenAI-compatible API. Call it from your code with your Gysga API key.
Multi-GPU: tensor parallel
SGLang · OpenAI API
Fast LLM serving with RadixAttention and an OpenAI-compatible API.
Multi-GPU: tensor parallel
Open WebUI + Ollama
ChatGPT-like interface with Ollama inside. Pull any model from the Ollama library in one click.
Ollama API
Ollama server: run Llama, Qwen, Gemma, DeepSeek and hundreds of other models with a simple API.
Text Generation WebUI
oobabooga web UI for chatting with and testing local models (GGUF, EXL2, Transformers).
ComfyUI
Node-based image and video generation: SDXL, FLUX, Wan, LTX-Video. ComfyUI-Manager included.
Speech-to-text & TTS · OpenAI API
faster-whisper transcription and text-to-speech behind an OpenAI-compatible audio API.
Whisper large-v3 API
Speech recognition service (faster-whisper, large-v3) with a Swagger UI.
Kokoro TTS · OpenAI API
Fast, natural text-to-speech (82M) with an OpenAI-compatible /v1/audio/speech endpoint.
LLaMA-Factory
Fine-tune 100+ LLMs with LoRA/QLoRA or full training from a web UI.
Multi-GPU: torchrun / DeepSpeed
Axolotl
YAML-driven fine-tuning toolkit (LoRA, QLoRA, full, DPO). Connect over SSH and run axolotl train.
Multi-GPU: torchrun / DeepSpeed
Unsloth
2× faster fine-tuning with less memory. Connect over SSH and run your Unsloth scripts.
Multi-GPU: torchrun / DeepSpeed
PyTorch 2.8
PyTorch 2.8 with CUDA 12.8 and cuDNN 9. Connect over SSH and work in /workspace.
Multi-GPU: torchrun / DeepSpeed
Jupyter Lab + PyTorch
Jupyter Lab on top of PyTorch 2.8. Open it through an SSH tunnel with the token from the connection panel.
Multi-GPU: torchrun / DeepSpeed
TensorFlow + Jupyter
TensorFlow with GPU support and Jupyter.
Multi-GPU: torchrun / DeepSpeed
Ubuntu 24.04 + CUDA 12.8
Clean Ubuntu 24.04 with the CUDA 12.8 toolkit and cuDNN. Install anything you need.
Multi-GPU: torchrun / DeepSpeed
Custom Docker image
Run any public Docker image with GPU access. Its default command is used unless you override it.
Multi-GPU: torchrun / DeepSpeed
Embeddings · OpenAI API
Serve embedding models (BGE-M3, E5) for search and RAG through /v1/embeddings.