AICL-Lab
Popular repositories Loading
-
cursor-rules
cursor-rules PublicArchive-grade .mdc rule library for Cursor AI — 26 production-ready rules for teams seeking stable, reusable AI coding guidance
-
the-book-of-secret-knowledge-zh
the-book-of-secret-knowledge-zh Public📚 秘密知识之书中文版 - A curated collection of tools, manuals, cheatsheets, and resources for SysAdmins, DevOps, Pentesters and Security Researchers. Chinese translation of the-book-of-secret-knowledge.
-
sgemm-optimization
sgemm-optimization PublicBilingual CUDA SGEMM optimization tutorial and reference implementation, from naive kernels to Tensor Core WMMA | 双语 CUDA SGEMM 优化教程与参考实现,从朴素内核到 Tensor Core WMMA
-
bookmarks-cleaner
bookmarks-cleaner PublicOffline-first bookmark cleaner: rules-first, ML-assisted, LLM-optional | 智能书签清理与分类:规则+ML+LLM(可选)
Python 4
-
the-art-of-hpc-zh
the-art-of-hpc-zh PublicThe Art of HPC Series Chinese Translation - HPC textbooks covering MPI, OpenMP, CUDA, and Scientific Computing (CC-BY 4.0) | 《高性能计算艺术》系列中文翻译 - 涵盖 MPI、OpenMP、CUDA 与科学计算的 HPC 教材(CC-BY 4.0)
Repositories
- compress-kit Public
Classic lossless compression algorithms in C++17 with round-trip verification | 经典无损压缩算法,C++17 实现,round-trip 验证
- cuflash-attn Public
CUDA C++ FlashAttention reference implementation - O(N) memory, FP32/FP16, forward/backward
- tiny-llm Public
CUDA-native C++ Transformer inference engine with W8A16 quantization, KV cache management, and optimized CUDA kernels
- ai-system-optimization-series Public
从编译器优化到 GPU 内核开发 — AI 基础设施工程师综合学习资源 | TVM, ONNX Runtime, CUTLASS, Triton
- modern-ai-kernels Public
TensorCraft-HPC: A header-only C++/CUDA kernel library for learning high-performance AI operators with progressive optimization paths
- gpu-spmv Public
High-Performance CUDA Sparse Matrix-Vector Multiplication Library • 70%+ bandwidth utilization • 4 optimized kernels • Spec-Driven Development
- mini-opencv Public
CUDA-accelerated GPU image processing library. 30-50x faster than CPU OpenCV. High-performance operators for computer vision: convolution, morphology, filters, geometric transforms, and more.
- sgemm-optimization Public
Bilingual CUDA SGEMM optimization tutorial and reference implementation, from naive kernels to Tensor Core WMMA | 双语 CUDA SGEMM 优化教程与参考实现,从朴素内核到 Tensor Core WMMA
- hetero-paged-infer Public
High-Performance LLM Inference Engine with PagedAttention & Continuous Batching in Rust
- n-body Public
High-performance N-body particle simulation with Barnes-Hut algorithm, GPU acceleration, and real-time visualization
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…