Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
gpu
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026
Dmitry Noranovich
Dmitry Noranovich
Dmitry Noranovich
Follow
Aug 13
Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026
#
nvidia
#
gpu
#
ai
#
deeplearning
Comments
Add Comment
2 min read
Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism
Charles
Charles
Charles
Follow
Aug 13
Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism
#
rust
#
gpu
#
programming
#
performance
Comments
Add Comment
4 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
Aarush Karak
Aarush Karak
Aarush Karak
Follow
Aug 13
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
#
gpu
#
cuda
#
vram
#
memory
1
 reaction
Comments
Add Comment
2 min read
Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming
Charles
Charles
Charles
Follow
Aug 11
Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming
#
rust
#
gpu
#
programming
#
performance
2
 reactions
Comments
Add Comment
4 min read
How We Cut Inference Cold Starts from Minutes to Seconds
Adi Ziv
Adi Ziv
Adi Ziv
Follow
for
AWS Community Builders
Aug 12
How We Cut Inference Cold Starts from Minutes to Seconds
#
aws
#
kubernetes
#
ai
#
gpu
1
 reaction
Comments
Add Comment
6 min read
Renting GPUs for AI? Start with VRAM, Not the GPU
Kavya
Kavya
Kavya
Follow
Aug 5
Renting GPUs for AI? Start with VRAM, Not the GPU
#
machinelearning
#
llm
#
gpu
#
cloud
Comments
Add Comment
2 min read
Why memory bandwidth matters more than TFLOPS for LLM inference
Kavya
Kavya
Kavya
Follow
Aug 3
Why memory bandwidth matters more than TFLOPS for LLM inference
#
llm
#
gpu
#
nvidia
#
machinelearning
Comments
Add Comment
3 min read
Deploying DeepSeek V3 (LLM) Using SGLang
Sanskriti Harmukh
Sanskriti Harmukh
Sanskriti Harmukh
Follow
for
Vultr
Aug 12
Deploying DeepSeek V3 (LLM) Using SGLang
#
ai
#
llm
#
gpu
#
docker
6
 reactions
Comments
1
 comment
2 min read
One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract
Christopher Maher
Christopher Maher
Christopher Maher
Follow
Aug 10
One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract
#
kubernetes
#
gpu
#
ai
#
devops
1
 reaction
Comments
1
 comment
8 min read
KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM
Ken Imoto
Ken Imoto
Ken Imoto
Follow
Jul 28
KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM
#
llm
#
ai
#
gpu
#
performance
1
 reaction
Comments
1
 comment
3 min read
Running a 122B Parameter Model and Agent Locally on AMD MI300X GPU — What I Learned
Yogi
Yogi
Yogi
Follow
Aug 10
Running a 122B Parameter Model and Agent Locally on AMD MI300X GPU — What I Learned
#
gpu
#
llm
#
agents
#
roc
1
 reaction
Comments
Add Comment
7 min read
GPU Monitoring & Metrics for MLOps
Harsha Kumarasingha (TMHKThennakoon)
Harsha Kumarasingha (TMHKThennakoon)
Harsha Kumarasingha (TMHKThennakoon)
Follow
Jul 25
GPU Monitoring & Metrics for MLOps
#
ai
#
webdev
#
mlops
#
gpu
Comments
Add Comment
1 min read
The 60% idle GPU that turned out to be a network policy
Leo
Leo
Leo
Follow
Jul 23
The 60% idle GPU that turned out to be a network policy
#
kubernetes
#
kubeflow
#
cilium
#
gpu
Comments
Add Comment
3 min read
Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower
Yehor Cherednichenko
Yehor Cherednichenko
Yehor Cherednichenko
Follow
Jul 22
Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower
#
machinelearning
#
python
#
performance
#
gpu
1
 reaction
Comments
Add Comment
6 min read
Building CI/CD Pipelines for GPU Validation
TANYA SRIVASTAVA
TANYA SRIVASTAVA
TANYA SRIVASTAVA
Follow
Jul 21
Building CI/CD Pipelines for GPU Validation
#
gpu
#
cicd
#
productivity
#
infrastructure
1
 reaction
Comments
Add Comment
10 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account