gysga

Model library

Deploy AI models in one click

Open models for text, code, vision, images, video, speech and search. Pick a model — we suggest the cheapest GPU that fits.

Type
Size
Instant start

Qwen3 0.6B

0.6B · Apache-2.0 · min. VRAM 6 GB

Tiny model for classification, routing and tests.

tinyfast

Recommended: RTX 3060· from $0.08/hr

Instant start

Qwen3 1.7B

1.7B · Apache-2.0 · min. VRAM 8 GB

Small and fast, good for simple chat and extraction.

fast

Recommended: RTX 3060· from $0.08/hr

Instant start

Qwen3 4B Instruct 2507

4B · Apache-2.0 · min. VRAM 12 GB

Best quality per gigabyte for cheap GPUs.

multilingual

Recommended: RTX 3060· from $0.08/hr

Instant start

SmolLM3 3B

3B · Apache-2.0 · min. VRAM 10 GB

Open 3B model with long context and reasoning mode.

Recommended: RTX 3060· from $0.08/hr

Instant start

Phi-4 mini

3.8B · MIT · min. VRAM 12 GB

Compact Microsoft model strong at math and logic.

reasoning

Recommended: RTX 3060· from $0.08/hr

Instant startNeeds your HF token

Llama 3.2 3B

3B · Llama 3.2 · min. VRAM 12 GB

Meta's small instruct model (needs your HF token).

Recommended: RTX 3060· from $0.08/hr

Instant startNeeds your HF token

Gemma 3 4B

4B · Gemma · min. VRAM 12 GB

Google's small multimodal model (needs your HF token).

vision

Recommended: RTX 3060· from $0.08/hr

Instant start

Qwen3 8B

8B · Apache-2.0 · min. VRAM 24 GB

Strong all-rounder with thinking mode, 100+ languages.

reasoningmultilingual

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

gpt-oss 20B

21B MoE · Apache-2.0 · min. VRAM 24 GB

OpenAI open-weight reasoning model with tool use.

reasoningmoetools

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Qwen2.5 7B Instruct

7B · Apache-2.0 · min. VRAM 24 GB

Proven chat model for assistants and RAG.

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Qwen2.5 Coder 7B

7B · Apache-2.0 · min. VRAM 24 GB

Code completion and generation for IDE assistants.

code

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Qwen2.5 VL 7B

7B · Apache-2.0 · min. VRAM 24 GB

Understands images, documents and screenshots.

visionocr

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

DeepSeek-R1 Distill Qwen 7B

7B · MIT · min. VRAM 24 GB

R1-style step-by-step reasoning in a small model.

reasoning

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

DeepSeek-R1 Distill Llama 8B

8B · MIT · min. VRAM 24 GB

R1 reasoning distilled into Llama 8B.

reasoning

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Mistral 7B Instruct v0.3

7B · Apache-2.0 · min. VRAM 24 GB

Fast European model with function calling.

tools

Recommended: 2× RTX 3060· from $0.16/hr

Instant startNeeds your HF token

Llama 3.1 8B

8B · Llama 3.1 · min. VRAM 24 GB

Meta's popular 8B model (needs your HF token).

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Qwen3 14B

14B · Apache-2.0 · min. VRAM 40 GB

Noticeably smarter than 8B, still affordable.

reasoningmultilingual

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Phi-4 14B

14B · MIT · min. VRAM 40 GB

Excellent at math, code and structured answers.

reasoning

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

DeepSeek-R1 Distill Qwen 14B

14B · MIT · min. VRAM 40 GB

Deeper reasoning than the 7B distill.

reasoning

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Mistral NeMo 12B

12B · Apache-2.0 · min. VRAM 40 GB

128k context, strong multilingual chat.

multilingual

Recommended: 4× RTX 3060· from $0.32/hr

Instant startNeeds your HF token

Gemma 3 12B

12B · Gemma · min. VRAM 40 GB

Multimodal Gemma (needs your HF token).

vision

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Qwen3 32B

32B · Apache-2.0 · min. VRAM 80 GB

Flagship dense Qwen3 for demanding assistants.

reasoningmultilingual

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

Qwen3 30B-A3B Instruct 2507

30B MoE · Apache-2.0 · min. VRAM 80 GB

MoE with only 3B active: big-model quality at high speed.

moefast

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

Qwen3 Coder 30B-A3B

30B MoE · Apache-2.0 · min. VRAM 80 GB

Agentic coding model for Cline, Roo, Continue.

codeagents

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

Qwen2.5 Coder 32B

32B · Apache-2.0 · min. VRAM 80 GB

Top open dense coding model.

code

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

Qwen2.5 VL 32B

32B · Apache-2.0 · min. VRAM 80 GB

Strong vision-language model for documents and charts.

visionocr

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

DeepSeek-R1 Distill Qwen 32B

32B · MIT · min. VRAM 80 GB

The strongest R1 distill.

reasoning

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

Mistral Small 3.2 24B

24B · Apache-2.0 · min. VRAM 80 GB

Balanced 24B with vision and function calling.

visiontools

Recommended: 8× RTX 3060· from $0.64/hr

Instant start

gpt-oss 120B

117B MoE · Apache-2.0 · min. VRAM 80 GB

OpenAI's largest open model on a single 80 GB GPU.

reasoningmoetools

Recommended: 8× RTX 3060· from $0.64/hr

Instant startNeeds your HF token

Gemma 3 27B

27B · Gemma · min. VRAM 80 GB

Largest Gemma 3 (needs your HF token).

vision

Recommended: 8× RTX 3060· from $0.64/hr

Instant startNeeds your HF token

Llama 3.3 70B

70B · Llama 3.3 · min. VRAM 160 GB

GPT-4-class open model; 2×80 GB (needs your HF token).

Recommended: 4× A40· from $1.60/hr

Instant start

GLM-4.5 Air

106B MoE · MIT · min. VRAM 320 GB

Agent-oriented MoE; 4×80 GB.

agentscodemoe

Recommended: 8× A40· from $3.20/hr

Instant start

Qwen3 235B-A22B FP8

235B MoE · Apache-2.0 · min. VRAM 320 GB

Frontier-level open model; 4×80 GB (Hopper, FP8).

moefrontier

Recommended: 8× A40· from $3.20/hr

Instant start

Qwen3 Coder 480B FP8

480B MoE · Apache-2.0 · min. VRAM 640 GB

Best open agentic coder; 8×80 GB (Hopper, FP8).

codeagentsfrontier

Recommended: 8× A100 80GB· from $8.80/hr

Instant start

DeepSeek-R1 0528 (671B)

671B MoE · MIT · min. VRAM 1000 GB

Full DeepSeek-R1; 8×H200.

reasoningfrontier

Recommended: 8× H200 141GB· from $20.80/hr

Instant start

DeepSeek-V3.1 (671B)

671B MoE · MIT · min. VRAM 1000 GB

Full DeepSeek-V3.1 hybrid thinking model; 8×H200.

frontiertools

Recommended: 8× H200 141GB· from $20.80/hr

Instant start

Qwen3 Embedding 0.6B

0.6B · Apache-2.0 · min. VRAM 6 GB

Fast multilingual embeddings for search and RAG.

multilingual

Recommended: RTX 3060· from $0.08/hr

Instant start

Qwen3 Embedding 8B

8B · Apache-2.0 · min. VRAM 24 GB

Top-quality embeddings (MTEB leader class).

multilingual

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

BGE-M3

568M · MIT · min. VRAM 6 GB

Dense + sparse multilingual embeddings.

multilingual

Recommended: RTX 3060· from $0.08/hr

Instant start

Multilingual E5 Large

560M · MIT · min. VRAM 6 GB

Classic multilingual retrieval embeddings.

multilingual

Recommended: RTX 3060· from $0.08/hr

Instant start

BGE Reranker v2 M3

568M · Apache-2.0 · min. VRAM 6 GB

Reranker for better RAG results (/v1/rerank).

rerank

Recommended: RTX 3060· from $0.08/hr

Instant start

Stable Diffusion XL 1.0

3.5B · OpenRAIL++ · min. VRAM 8 GB

Classic SDXL: huge ecosystem of LoRAs, runs on 8 GB.

photoart

Recommended: RTX 3060· from $0.08/hr

Instant start

Juggernaut XL v9

3.5B · CreativeML OpenRAIL-M · min. VRAM 8 GB

Photorealistic SDXL fine-tune, great for portraits.

photorealism

Recommended: RTX 3060· from $0.08/hr

Instant start

Lumina Image 2.0

2.6B · Apache-2.0 · min. VRAM 12 GB

Efficient text-to-image with good prompt following.

art

Recommended: RTX 3060· from $0.08/hr

Instant start

FLUX.1 schnell (FP8)

12B · Apache-2.0 · min. VRAM 16 GB

Top-quality images in 4 steps, commercial use allowed.

phototext-in-imagefast

Recommended: RTX 4060 Ti 16GB· from $0.12/hr

Instant start

HiDream-I1 Fast (FP8)

17B · MIT · min. VRAM 24 GB

State-of-the-art open image model, MIT license.

photoart

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Qwen-Image (FP8)

20B · Apache-2.0 · min. VRAM 40 GB

Best-in-class text rendering in images (posters, UI, Chinese/English).

text-in-imageposters

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Wan 2.1 Text-to-Video 1.3B

1.3B · Apache-2.0 · min. VRAM 12 GB

Video generation even on 12 GB cards (480p).

text-to-videolight

Recommended: RTX 3060· from $0.08/hr

Instant start

Wan 2.2 TI2V 5B

5B · Apache-2.0 · min. VRAM 24 GB

720p text- and image-to-video on a single 24 GB GPU.

text-to-videoimage-to-video

Recommended: 2× RTX 3060· from $0.16/hr

Instant start

Wan 2.2 Text-to-Video 14B (FP8)

14B MoE · Apache-2.0 · min. VRAM 40 GB

Cinematic 720p video; 4-step LoRAs included for speed.

text-to-videocinematic

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Wan 2.2 Image-to-Video 14B (FP8)

14B MoE · Apache-2.0 · min. VRAM 40 GB

Animate any image into a 720p clip.

image-to-video

Recommended: 4× RTX 3060· from $0.32/hr

Instant start

Mochi 1 Preview (FP8)

10B · Apache-2.0 · min. VRAM 24 GB

Smooth-motion text-to-video from Genmo.

text-to-video

Recommended: 2× RTX 3060· from $0.16/hr

Speech-to-text & TTS · OpenAI API

min. VRAM 6 GB

faster-whisper transcription and text-to-speech behind an OpenAI-compatible audio API.

Whisper large-v3 API

min. VRAM 8 GB

Speech recognition service (faster-whisper, large-v3) with a Swagger UI.

Kokoro TTS · OpenAI API

min. VRAM 4 GB

Fast, natural text-to-speech (82M) with an OpenAI-compatible /v1/audio/speech endpoint.

No models match these filters.

Combine several GPUs for bigger models

Rent 2, 4 or 8 GPUs in one server and use them as one:

Tensor parallelism for inference

vLLM and SGLang split one model across all GPUs of the server, so their memory adds up: Llama 3.3 70B on 2×A100 80GB, Qwen3 235B on 4×H100, DeepSeek-R1 671B on 8×H200.

Distributed training

PyTorch templates see every GPU: use torchrun with DDP, FSDP or DeepSpeed ZeRO to train faster or fit larger models.

Why one server

GPUs are combined inside one machine over NVLink/PCIe. Splitting a model across computers over the internet would be tens of times slower, so we never do it.

Templates

Ready environments

Every server starts from a template. Templates with a model slot load the model automatically; others give you a clean environment over SSH.

LLM serving

vLLM · OpenAI API

High-throughput LLM server with an OpenAI-compatible API. Call it from your code with your Gysga API key.

Multi-GPU: tensor parallel

SGLang · OpenAI API

Fast LLM serving with RadixAttention and an OpenAI-compatible API.

Multi-GPU: tensor parallel

Open WebUI + Ollama

ChatGPT-like interface with Ollama inside. Pull any model from the Ollama library in one click.

Ollama API

Ollama server: run Llama, Qwen, Gemma, DeepSeek and hundreds of other models with a simple API.

Text Generation WebUI

oobabooga web UI for chatting with and testing local models (GGUF, EXL2, Transformers).

Images & video

ComfyUI

Node-based image and video generation: SDXL, FLUX, Wan, LTX-Video. ComfyUI-Manager included.

Speech

Speech-to-text & TTS · OpenAI API

faster-whisper transcription and text-to-speech behind an OpenAI-compatible audio API.

Whisper large-v3 API

Speech recognition service (faster-whisper, large-v3) with a Swagger UI.

Kokoro TTS · OpenAI API

Fast, natural text-to-speech (82M) with an OpenAI-compatible /v1/audio/speech endpoint.

Fine-tuning

LLaMA-Factory

Fine-tune 100+ LLMs with LoRA/QLoRA or full training from a web UI.

Multi-GPU: torchrun / DeepSpeed

Axolotl

YAML-driven fine-tuning toolkit (LoRA, QLoRA, full, DPO). Connect over SSH and run axolotl train.

Multi-GPU: torchrun / DeepSpeed

Unsloth

2× faster fine-tuning with less memory. Connect over SSH and run your Unsloth scripts.

Multi-GPU: torchrun / DeepSpeed

Development

PyTorch 2.8

PyTorch 2.8 with CUDA 12.8 and cuDNN 9. Connect over SSH and work in /workspace.

Multi-GPU: torchrun / DeepSpeed

Jupyter Lab + PyTorch

Jupyter Lab on top of PyTorch 2.8. Open it through an SSH tunnel with the token from the connection panel.

Multi-GPU: torchrun / DeepSpeed

TensorFlow + Jupyter

TensorFlow with GPU support and Jupyter.

Multi-GPU: torchrun / DeepSpeed

Ubuntu 24.04 + CUDA 12.8

Clean Ubuntu 24.04 with the CUDA 12.8 toolkit and cuDNN. Install anything you need.

Multi-GPU: torchrun / DeepSpeed

Custom Docker image

Run any public Docker image with GPU access. Its default command is used unless you override it.

Multi-GPU: torchrun / DeepSpeed

Search & RAG

Embeddings · OpenAI API

Serve embedding models (BGE-M3, E5) for search and RAG through /v1/embeddings.