Change the repository type filter
All
Repositories list
44 repositories
compress-kit
PublicClassic lossless compression algorithms in C++17 with round-trip verification | 经典无损压缩算法,C++17 实现,round-trip 验证hetero-paged-infer
PublicHigh-Performance LLM Inference Engine with PagedAttention & Continuous Batching in Rustbuild-your-own-tools
Publicmini-image-pipe
PublicGPU-accelerated image processing pipeline with DAG scheduling, CUDA operators, and multi-stream executionn-body
PublicHigh-performance N-body particle simulation with Barnes-Hut algorithm, GPU acceleration, and real-time visualizationmini-opencv
PublicCUDA-accelerated GPU image processing library. 30-50x faster than CPU OpenCV. High-performance operators for computer vision: convolution, morphology, filters, …graph-viewer
Publicmind-gym
Public🧠 Browser-based memory training PWA with FSRS-4.5 spaced repetition, N-back training, and adaptive difficulty | 基于 FSRS-4.5 间隔重复、N-back 训练和自适应难度的浏览器记忆力训练 PWAbookmarks-cleaner
PublicOffline-first bookmark cleaner: rules-first, ML-assisted, LLM-optional | 智能书签清理与分类:规则+ML+LLM(可选)sgemm-optimization
PublicBilingual CUDA SGEMM optimization tutorial and reference implementation, from naive kernels to Tensor Core WMMA | 双语 CUDA SGEMM 优化教程与参考实现,从朴素内核到 Tensor Core WMM…webgpu-sorting
PublicHigh-performance GPU sorting library using WebGPU compute shaders (Bitonic Sort, Radix Sort) with TypeScript API, live demo, and comprehensive documentationmodern-ai-kernels
PublicTensorCraft-HPC: A header-only C++/CUDA kernel library for learning high-performance AI operators with progressive optimization pathsaurora-signal
Publicwebrtc
PublicLightweight WebRTC Demo: Go Signaling Server + Vanilla JavaScript Client, OpenSpec-Driven Development | 轻量级 WebRTC 演示:Go 信令服务 + 原生 JavaScript 客户端,OpenSpec 驱动开发yolo-toys
PublicMulti-Model Real-Time Visual Recognition System - Supports YOLO, DETR, OWL-ViT, Grounding DINO, BLIP with REST API and WebSocket Streaming Inference | 多模型实时视觉识别…fq-compressor-rust
PublicHigh-performance FASTQ compressor with block-indexed archive format, random access support, and multiple compression modescursor-rules
PublicArchive-grade .mdc rule library for Cursor AI — 26 production-ready rules for teams seeking stable, reusable AI coding guidancecuflash-attn
PublicCUDA C++ FlashAttention reference implementation - O(N) memory, FP32/FP16, forward/backwardmini-inference-engine
PublicCUDA GEMM optimization tutorial and mini inference engine with progressive kernels, benchmarks, and OpenSpec docstiny-llm
PublicCUDA-native C++ Transformer inference engine with W8A16 quantization, KV cache management, and optimized CUDA kernelsgpu-spmv
PublicHigh-Performance CUDA Sparse Matrix-Vector Multiplication Library • 70%+ bandwidth utilization • 4 optimized kernels • Spec-Driven Developmentray-tracer
PublicGPU-accelerated ray tracer with BVH acceleration and path tracing support- 从编译器优化到 GPU 内核开发 — AI 基础设施工程师综合学习资源 | TVM, ONNX Runtime, CUTLASS, Triton
gpu-fft
PublicHigh-performance GPU-accelerated FFT library for JavaScript/TypeScript using WebGPU compute shaders. Zero runtime dependencies, dual GPU/CPU paths, TypeScript-f…particle-fluid-sim
PublicHigh-performance WebGPU particle fluid simulation with compute shaderstiny-dl-inference
PublicZero-dependency WebGPU deep learning inference engine (~50KB vs TensorFlow.js ~2MB)llm-speed
PublicCUDA kernels for LLM inference: FlashAttention forward, Tensor Core GEMM, and PyTorch bindingshpc-ai-optimization-lab
PublicCUDA kernel optimization lab: GEMM, FlashAttention, quantization, and GPU performance learning.- C++17 DAG scheduler for heterogeneous CPU/GPU workloads - production-ready with CPU-only validation path
cuda-kernel-academy
Public
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.