Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
cuda
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 13
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
#
aws
#
vllm
#
cuda
#
machinelearning
5
 reactions
Comments
Add Comment
9 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
xbill
xbill
xbill
Follow
for
AWS Community Builders
Aug 13
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
#
aws
#
vllm
#
cuda
#
machinelearning
Comments
Add Comment
9 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
Aarush Karak
Aarush Karak
Aarush Karak
Follow
Aug 13
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit
#
gpu
#
cuda
#
vram
#
memory
1
 reaction
Comments
Add Comment
2 min read
I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware
Zero_planck
Zero_planck
Zero_planck
Follow
Jul 28
I Built a CUDA Engine That Streams 744B Parameter AI Models on Consumer Hardware
#
ai
#
cuda
#
opensource
#
machinelearning
2
 reactions
Comments
Add Comment
4 min read
Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization
Tamiz Uddin
Tamiz Uddin
Tamiz Uddin
Follow
Jul 14
Breaking CUDA's Chains: Open-Source Alternatives for Cross-GPU Performance Optimization
#
ai
#
breaking
#
cuda
#
chains
Comments
Add Comment
2 min read
Porting a 1,200-line persistent CUDA megakernel to Qwen3-TTS: ~25 ms to first audio chunk
Pratham Sharma
Pratham Sharma
Pratham Sharma
Follow
Jul 6
Porting a 1,200-line persistent CUDA megakernel to Qwen3-TTS: ~25 ms to first audio chunk
#
cuda
#
ai
#
machinelearning
#
performance
1
 reaction
Comments
Add Comment
3 min read
RTX 5090 survival guide: sm_120, CUDA 12 and 13 side by side, and the xformers trap
Will Kline
Will Kline
Will Kline
Follow
Jul 28
RTX 5090 survival guide: sm_120, CUDA 12 and 13 side by side, and the xformers trap
#
nvidia
#
cuda
#
hardware
#
ai
1
 reaction
Comments
Add Comment
7 min read
GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension
ranjithvutnoor
ranjithvutnoor
ranjithvutnoor
Follow
Aug 5
GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension
#
pytorch
#
cuda
#
machinelearning
#
performance
4
 reactions
Comments
2
 comments
11 min read
Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)
skyne
skyne
skyne
Follow
Jun 27
Resurrecting Kepler: Getting Modern LLMs Running on a GTX 770 (Kernel 7.x)
#
cuda
#
linux
#
llm
#
gpu
1
 reaction
Comments
Add Comment
4 min read
I Built Flash Attention From Scratch — Here's What Nobody Tells You About It
JITENDRA KUMAR SINGH
JITENDRA KUMAR SINGH
JITENDRA KUMAR SINGH
Follow
Jun 20
I Built Flash Attention From Scratch — Here's What Nobody Tells You About It
#
showdev
#
machinelearning
#
deeplearning
#
cuda
Comments
Add Comment
2 min read
Bypassing the OS to Run LLMs: What I Learned Building a Firmware-Centric Runtime
RoTSL
RoTSL
RoTSL
Follow
Jun 21
Bypassing the OS to Run LLMs: What I Learned Building a Firmware-Centric Runtime
#
llm
#
nvidia
#
cuda
#
firmware
Comments
Add Comment
5 min read
One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof
soy
soy
soy
Follow
Jul 18
One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof
#
gpu
#
benchmark
#
machinelearning
#
cuda
Comments
Add Comment
4 min read
Learn CUDA and GPU programming without owning a GPU
I Want To Learn Programming
I Want To Learn Programming
I Want To Learn Programming
Follow
Jun 9
Learn CUDA and GPU programming without owning a GPU
#
cuda
#
gpu
#
parallel
Comments
Add Comment
2 min read
Adding GPU backends to a pure-C TTS engine: Metal, CUDA, and the rented-Mac trick
Gabriele Mastrapasqua
Gabriele Mastrapasqua
Gabriele Mastrapasqua
Follow
Jul 7
Adding GPU backends to a pure-C TTS engine: Metal, CUDA, and the rented-Mac trick
#
c
#
cuda
#
metal
#
machinelearning
7
 reactions
Comments
4
 comments
6 min read
Notes on CUDA Tensor Core GEMM (WMMA)
member_2e5ba30f
member_2e5ba30f
member_2e5ba30f
Follow
May 31
Notes on CUDA Tensor Core GEMM (WMMA)
#
cuda
#
gpu
#
cpp
#
performance
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account