DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026

Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026

Comments
2 min read
Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism

Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism

Comments
4 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

1
Comments
2 min read
Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming

Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming

2
Comments
4 min read
How We Cut Inference Cold Starts from Minutes to Seconds

How We Cut Inference Cold Starts from Minutes to Seconds

1
Comments
6 min read
Renting GPUs for AI? Start with VRAM, Not the GPU

Renting GPUs for AI? Start with VRAM, Not the GPU

Comments
2 min read
Why memory bandwidth matters more than TFLOPS for LLM inference

Why memory bandwidth matters more than TFLOPS for LLM inference

Comments
3 min read
Deploying DeepSeek V3 (LLM) Using SGLang

Deploying DeepSeek V3 (LLM) Using SGLang

6
Comments 1
2 min read
One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract

One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract

1
Comments 1
8 min read
KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM

KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM

1
Comments 1
3 min read
Running a 122B Parameter Model and Agent Locally on AMD MI300X GPU — What I Learned

Running a 122B Parameter Model and Agent Locally on AMD MI300X GPU — What I Learned

1
Comments
7 min read
GPU Monitoring & Metrics for MLOps

GPU Monitoring & Metrics for MLOps

Comments
1 min read
The 60% idle GPU that turned out to be a network policy

The 60% idle GPU that turned out to be a network policy

Comments
3 min read
Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower

Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower

1
Comments
6 min read
Building CI/CD Pipelines for GPU Validation

Building CI/CD Pipelines for GPU Validation

1
Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.