Taiwan LLM Inference Benchmarks
Compare H100, H200, B200, L40S, A100 and MI300X on real workloads — inference, fine-tuning, and RAG — with transparent methodology and NTD pricing.
The recipe landscape
Every cell is a recipe × chip pairing across all three suites. Dot color names the suite; an empty cell is a real gap you can target with a submission.
| recipe | H100 80GB | H200 141GB | B200 | H100 PCIe 80GB | GB200 NVL72 | L4 24GB | A100 PCIe 40GB | RTX 6000 Ada 48GB | MI250X 128GB |
|---|---|---|---|---|---|---|---|---|---|
| SGLang BF16 | |||||||||
| SGLang FP16 | |||||||||
| SGLang FP8 | |||||||||
| TensorRT-LLM BF16 | |||||||||
| TensorRT-LLM FP16 | |||||||||
| TensorRT-LLM FP8 | |||||||||
| TensorRT-LLM INT8 | |||||||||
| TGI FP16 | |||||||||
| vLLM BF16 | |||||||||
| vLLM FP16 | |||||||||
| vLLM FP8 | |||||||||
| vLLM INT8 |
Three test suites
Pick a workload to see top performers and full rankings.
LLM Inference
Top runEnd-to-end throughput on Llama-3.1-8B / 70B at batch=8, covering both online and offline serving.
LLM Fine-tuning
Top runLoRA fine-tuning throughput on Qwen2.5-7B, measuring training efficiency and GPU utilization.
RAG End-to-End
Top runRetrieval + generation throughput on a 100-document bilingual corpus.
Hardware coverage
NVIDIA lineup
- NVIDIA A100 80GB
- NVIDIA A100 PCIe 40GB
- NVIDIA B200
- NVIDIA GB200 NVL72
- NVIDIA H100 80GB
- NVIDIA H100 PCIe 80GB
- NVIDIA H200 141GB
- NVIDIA L4 24GB
- NVIDIA L40S 48GB
- NVIDIA RTX 3090 24GB
- NVIDIA RTX 4080 SUPER 16GB
- NVIDIA RTX 4090 24GB
- NVIDIA RTX 5090 32GB
- NVIDIA RTX 6000 Ada 48GB
AMD lineup
- AMD MI250X 128GB
- AMD MI300X 192GB
- AMD RX 7900 XTX 24GB
Data sovereignty note: Learn about sovereign AI factory
Recent submissions
View all →| Chip | Model | Framework | Suite | Primary metric | Date |
|---|---|---|---|---|---|
| NVIDIA GB200 NVL72 ×36 | Llama-3.1-8B-Instruct | SGLang | RAG End-to-End | 920 queries/sec | 2026-09-21 |
| NVIDIA GB200 NVL72 ×36 | Qwen2.5-7B | TensorRT-LLM | LLM Fine-tuning | 2,120 samples/hour | 2026-09-20 |
| NVIDIA GB200 NVL72 ×36 | Llama-3.1-70B-Instruct | TensorRT-LLM | LLM Inference | 268,000 tokens/sec | 2026-09-19 |
| NVIDIA GB200 NVL72 ×72 | Qwen2.5-7B | vLLM | LLM Fine-tuning | 4,800 samples/hour | 2026-09-18 |
| NVIDIA GB200 NVL72 ×72 | Llama-3.1-8B-Instruct | vLLM | RAG End-to-End | 1,480 queries/sec | 2026-09-15 |
Want to submit your results?
We accept submissions from any Taiwan-based lab, university, or partner. Reproducible scripts encouraged.