gpu-optimization.
4 writings found
Latest Archives
Transferring GPU Expertise Across Hardware with AI-Driven Kernel Search
How evolutionary search and CUDA-to-MLX translation layers enable automatic GPU kernel optimization for Apple Silicon, bridging decades of NVIDIA expertise.
Cross-Platform GPU Kernels: When AI Learns to Translate Optimization
How evolutionary kernel search with structured translation layers can port decades of CUDA expertise to Apple Silicon without rebuilding from scratch.
Transferring GPU Expertise Across Hardware with AI-Driven Kernel Search
How evolutionary search and structured translation layers enable CUDA optimization knowledge to transfer to Apple Silicon, unlocking near-expert performance without rebuilding from scratch.
Why FlashAttention Breaks the Profiler (And Why That's Good)
FlashAttention shows low GPU occupancy yet outperforms all other attention backends. Here's what the profiler isn't telling you about modern kernel design.
Prev
Page 1 of 1 Next