GLM-5.3-Flash-Uncensored-FP8 - 4× H200 141GB
glm-5-3-flash-uncensored-fp8
glm-5-3-flash-uncensored-fp8
About
GLM-5.3-Flash-Uncensored-FP8 is in the OpenLLM catalog on a dedicated 4× H200 141GB. Deploy it with a flat time pack and call your instance through an OpenAI-compatible API.
Compare
Model Cost Across Durations
Live pack pricing vs typical API estimates from 11 hours through 1 month.
Live pack pricing for this model — API competitor estimates coming soon.
Time pack
GLM-5.3-Flash-Uncensored-FP8 on 4× H200 141GB
24 hours cost
$530 USD
Lowest
Models in chart
- GLM-5.3-Flash-Uncensored-FP8 on 4× H200 141GB
At a glance
GPU
4× H200 141GB
GPUs
4
Memory
141GB
Apps & integrations
Choose an app below. Each guide shows how to point the app at your OpenAI-compatible endpoint.
n8n
Automate workflows and call your model as a node.
Open
OpenClaw
Build AI agents and tools on an OpenAI-compatible endpoint.
Open
Hermes
Connect agent runners to your chat completions endpoint.
Open
OpenCode
Power developer tools with your OpenAI-compatible model.
Open
Cursor
Override OpenAI Base URL in Cursor Settings and use your model with BYOK.
Open
VS Code
Use the Cline extension in VS Code to connect your OpenAI-compatible endpoint.
Open
Codex
Run OpenAI Codex CLI against your Chat Completions endpoint via config.toml.
Open
Raspberry Pi
Full Pi OS guide: SSH, API keys, curl, Python venv, systemd, and troubleshooting.
Open