Tags¶
Following is a list of relevant tags:
0/1 Adam¶
1-bit Adam¶
1-bit LAMB¶
3FS¶
4+1 View Model¶
5G¶
64bits vs 32bits¶
A3¶
A5¶
ADR¶
ADXL¶
AGENTX¶
AGRS¶
AI¶
- AI Coding Harness
- AI Compiler
- AI Image
- Agile Governance: Balancing IPD and AI Innovation
- Code Migration And Alignment
- Codex Model Speed Benchmark
- IPD Q&A
- Introduction to AI and Machine Learning Basics
- Subagent Control Plane
- 成功的软件商业解决方案
AI AGENT¶
AI Agent¶
AI CODING¶
AI Chip¶
AI Cluster¶
AI Engineering¶
AI INFRA¶
AI Infrastructure¶
AI STARTUP¶
AI Systems¶
AIInfra¶
AISBench¶
AIV¶
- AIV Direct Drive
- AIV Direct Drive Programming
- AIV MoE All-to-All
- Ascend Communication Stack
- Ascend DeepEP MoE Dispatch Optimization
- Ascend DeepEP Ref Dispatch SIMT
- DualPath KV Loading
- Mooncake Ascend Transports
- SHMEM Symmetric Memory
- TileXR-SHMEM Abstractions
ALL-TO-ALL¶
ALST¶
AMAT¶
AMD¶
APR¶
ASCEND¶
ASPLOS¶
ASTRA-Sim¶
ASYNC I/O¶
ATI¶
ATTENTION RESIDUALS¶
AVX¶
Acceptance Criteria¶
Activation Offload¶
Agent¶
Agent Harness¶
Agent Loop¶
Agentic RL¶
Algorithm¶
All-to-All¶
AllGather¶
Alpha¶
Analysis¶
Apt¶
Architecture Analysis¶
Architecture Governance¶
Archon¶
Ascend¶
- BSND TND Operator Layout
- CloudMatrix384 LLM Serving
- ClusterHealthDetect A3 Performance
- Echo Training Simulator
- Memory-Semantic TransferQueue
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- NPU Training Operators - RoPE MRoPE
- Next of My Ascend Career
- SimAI Architecture
- TileXR-SHMEM Abstractions
Ascend 910C¶
Ascend 950¶
Ascend A5¶
Ascend C¶
Ascend NPU¶
AscendCL¶
Assembly¶
Async¶
Async RL¶
Attention¶
AudioVideo¶
AutoResearch¶
AutoTP¶
Autonomous Driving¶
Autotuning¶
B200¶
BAM¶
BATCH¶
BFS¶
BHive¶
BLOCK DEVICE¶
BPMN¶
BSHD¶
BSND¶
BT¶
BTL¶
Baka Mitai¶
Bash¶
Benchmark¶
Benchmarking¶
Big-Endian¶
Blackwell¶
C¶
- Naming
- [C++ Basic] Exploring Useful Built-in Functions
- [C++ Basic] Grammar
- [C++ Basic] Types
- [C++] Destructor Order
C++¶
- Gperftools
- Hash map
- Naming
- [C++ Basic] Exploring Useful Built-in Functions
- [C++ Basic] Grammar
- [C++ Basic] STL Data Structure
- [C++ Basic] Types
- [C++ Basic] User-Defined Types: Class
- [C++] Destructor Order
CACHE OFFLOADING¶
CANN¶
CANN VMM¶
- Ascend Communication API Evolution
- Ascend Communication Runtime Evolution
- Ascend Communication Stack
- Mooncake Ascend Transports
CCD¶
CCX¶
CI¶
CISC¶
CODE WALKTHROUGH¶
COMPENSATION¶
CONTINUOUS BATCHING¶
CP¶
CPP¶
CPU¶
CPU Thread Socket¶
CUDA¶
CUDA Allocator¶
CUFILE¶
- Mooncake Classic NVMeoF Transport
- Mooncake File vs Block Device
- Mooncake TENT GDS
- Mooncake TENT Request Path
CV¶
Calling Conventions¶
Capability Map¶
Career¶
- 1.2 Career
- 1.2 Career:1 秋招
- Career Transferable skill / Durable skills / Core capabilities
- DFX: Design for X
- Multi-Objective Decision Making
- QCC:Quality Control Circle
Career Strategy¶
Checkpoint¶
Chunk Layer¶
Class¶
Claude Code¶
CloudMatrix384¶
Cloudflare¶
ClusterHealthDetect¶
Co-design¶
Codeforces¶
Codex¶
Cohousing¶
Collective Communication¶
Communication¶
Communication Logging¶
Communication-Compute Fusion¶
Compatibility¶
Compiler¶
Computer Architecture¶
Conference¶
Context Parallelism¶
Crawler¶
DDR¶
DEEPEP¶
DFD¶
DFX¶
DFlash¶
DIRECT STORAGE¶
DISAGGREGATION¶
DISTRIBUTED STORAGE¶
DLP¶
DMA¶
DNS¶
DP¶
DPO¶
DRAM¶
DSA¶
DanceGRPO¶
DanceOPD¶
Data Lineage¶
Data Placement¶
DataStates¶
Database¶
Databases¶
Dataset¶
Debug¶
Decision Analysis¶
Decode¶
Deep Learning¶
DeepEP¶
DeepNVMe¶
DeepSeek-R1¶
DeepSpeed¶
- DeepSpeed Communication Compression and Hiding
- DeepSpeed I/O, Offload, and Asynchrony
- DeepSpeed Memory and Parallelism
- DeepSpeed MoE and Model Compression
- DeepSpeed Observability and Autotuning
Deepfake¶
Design Patterns¶
DevLog¶
- [DevLog] PLAN
- [DevLog]24Q3P1 - Optimize PTA With Thread Affinity
- [DevLog]24Q4P2 - Lazy Initialize during old dispatcher way
DiT¶
Diffusion¶
Diffusion Language Model¶
Diffusion Model¶
Digital Worker¶
- AI Documentation Workflow
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
- Personal Advantage Workflow
Disk¶
Distributed Systems¶
Distributed Training¶
Domain Model¶
Domino¶
Domino EP¶
DualPath¶
Dynamic Programming¶
ECC¶
EDP¶
EP¶
ER Model¶
ESOP¶
ETP¶
EXTENSIBILITY¶
Echart¶
Echo¶
Ed2k¶
Evals¶
Evaluation¶
Evidence¶
Executable file¶
Expert Parallelism¶
- CloudMatrix384 LLM Serving
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- XTuner Domino EP
Explain¶
FABRIC MEMORY¶
FAST¶
FBS¶
FLA¶
FLOPs Profiler¶
FMA¶
FP8¶
FPDT¶
FSDP¶
FSDP2¶
FUSELINK¶
Fabric Memory¶
- Ascend Communication API Evolution
- Ascend Communication Runtime Evolution
- Ascend Communication Stack
FeatureMatrix¶
FlashAttention¶
FlashInfer¶
Flow-Factory¶
FlowServe¶
Flowchart¶
Foundation Model Training¶
Frontier Models¶
Function Tree¶
Functional Decomposition¶
FurtherStudy¶
GCN¶
GDDR6x¶
GDN¶
- Attention Architecture Evolution
- Attention Cache and Sequence Parallelism
- BSND TND Operator Layout
- NPU Training Operators - GDN
- Training Performance Model
- vLLM Inference Profiling
GDS¶
- GPU-Initiated I/O
- Mooncake Classic NVMeoF Transport
- Mooncake File vs Block Device
- Mooncake NDS Integration
- Mooncake TENT GDS
- Mooncake TENT Request Path
GIL¶
GLIBC¶
GLM-5¶
GLM-5.2¶
GMM¶
GMT¶
GNN¶
GNU¶
GPT¶
GPT-5.6¶
GPTQ¶
GPU¶
GPU Benchmark¶
GPU COMMUNICATION¶
GPU DIRECT STORAGE¶
GPU OFFLOADING¶
GPU-Aware MPI¶
GPU-Centric Communication¶
GPU-INITIATED I/O¶
GPU-INITIATED IO¶
GPUDIRECT STORAGE¶
GPUDirect¶
GQA¶
GRPO¶
- BSND TND Operator Layout
- Diffusion LLM Post-Training
- NPU Training Operators - GDN
- RL Algorithms: PPO-RLHF & GRPO-family
Game¶
Goal¶
H100¶
H200¶
H2D¶
HARDWARE-SOFTWARE CO-DESIGN¶
HCCL¶
- AIV Direct Drive
- AIV Direct Drive Programming
- Ascend Communication API Evolution
- Ascend Communication Stack
- ClusterHealthDetect A3 Performance
- DualPath KV Loading
- Echo Training Simulator
- Mooncake Ascend Transports
- NPU Training Operators - MC2
- SHMEM Symmetric Memory
- SimAI Architecture
- TileXR-SHMEM Abstractions
HCCS¶
HCOMM¶
HDMI¶
HIGH AVAILABILITY¶
HIXL¶
HPCAI¶
HPL¶
HPL-PL¶
HTML¶
Hardware Software Co-design¶
HiXL¶
Hopper¶
Housing¶
Huawei¶
Hugo¶
HybridEP¶
HyperParallel-MoE¶
ILP¶
INFERENCE¶
INFERENCEX¶
IO_URING¶
IP¶
IPC¶
IPCC¶
IPD¶
IPO¶
Image¶
Inference¶
Inference Quantization¶
Inference Serving¶
Jellyfin¶
Jupyter¶
KDA¶
KIMI¶
KV CACHE¶
- Long-Context KV Cache Systems
- Mooncake
- Mooncake Codebase Architecture
- Mooncake KV Metadata
- Mooncake NDS Integration
- Mooncake Store Design
- PagedAttention
- Tutti SSD-Backed KV Cache
- vLLM KV Offloading and GDS
KV CONNECTOR¶
KV Cache¶
- Attention Cache and Sequence Parallelism
- CloudMatrix384 LLM Serving
- Distributed KV Cache Management
- DualPath KV Loading
- KV Cache Fundamentals
KVCACHE¶
Kavita¶
Kernel¶
Kernel Launch¶
Kimi K3¶
Knowledge Distillation¶
Komga¶
Kunpeng¶
LATENTMOE¶
LDA¶
LINUX¶
LLM¶
LLM Application¶
LLM INFERENCE¶
LLM Inference¶
LLM SERVING¶
LLM SYSTEMS¶
LLM Serving¶
LLM Wiki¶
LLaDA¶
LONG CONTEXT¶
LeetCode¶
Leetcode¶
LegacyBugs¶
Linking¶
Lock¶
Long Context¶
MAC¶
MATERIAL FOR MKDOCS¶
MC2¶
MCA¶
MCDA¶
MESH NOC¶
METADATA¶
MFU¶
MHA¶
MILAN¶
MINDIE¶
MIPS¶
MKDOCS¶
MLA¶
MLP¶
MOE¶
- AIV MoE All-to-All
- Ascend DeepEP MoE Dispatch Optimization
- Ascend DeepEP Ref Dispatch SIMT
- Kimi K3 Report
MOONCAKE¶
- Mooncake
- Mooncake Ascend Transports
- Mooncake Classic NVMeoF Transport
- Mooncake Classic vs TENT Engine
- Mooncake Codebase Architecture
- Mooncake File vs Block Device
- Mooncake KV Metadata
- Mooncake NDS Integration
- Mooncake Store Design
- Mooncake TENT GDS
- Mooncake TENT Request Path
- Mooncake vLLM Ascend
- vLLM KV Offloading and GDS
MOPD¶
MPI¶
- Dynamic pool dispatch
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- MPI
- Memory Semantics vs RDMA
- Python MPI
- Why MPI_Init is slow
MPI option¶
MPI_Init¶
MPK¶
MPS¶
MQA¶
MRoPE¶
MTE¶
- Ascend Data Movement Evolution
- DMA Communication Terminology
- Memory Semantics vs RDMA
- Memory-Semantic TransferQueue
MTP¶
MULTI-NIC¶
MULTIMODAL¶
MXFP8¶
Mac¶
Machine Learning¶
Marriage¶
Matchmaking¶
MegaKernel¶
MegaMoE¶
Megatron¶
Megatron Bridge¶
Megatron Core¶
MemFabric¶
Memory Hierarchy¶
Memory Optimization¶
Memory Semantics¶
- Ascend Communication Runtime Evolution
- Ascend Communication Stack
- Ascend Data Movement Evolution
- Memory Semantics vs RDMA
- Memory-Semantic TransferQueue
Mesh Interconnect Architecture¶
Metrics¶
Micro¶
Micro-Fusion¶
Micro-architecture¶
Microarchitecture¶
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Out-Of-Order execution(OoOE/OOE) & Register Renaming
- Microarchitecture: Pipeline of Intel Core CPUs
- Microarchitecture: Zero (one) idioms & Mov Elimination
MindSpeed¶
MindSpeed-MM¶
- AI Training Parallelism
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- VeRL Backend Parallelism
MixZ++¶
MoE¶
- AI Training Parallelism
- AIV Direct Drive
- CloudMatrix384 LLM Serving
- DeepSpeed MoE and Model Compression
- Inference MegaKernel
- Kimi K3 NPU Training
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- Training Performance Model
- VeRL Router Replay
- XTuner Domino EP
- XTuner Memory Optimization
- xDeepServe on CloudMatrix384
MoQ¶
Model Compression¶
Module¶
Money¶
Monitor¶
MoonEP¶
Mooncake¶
Mount¶
Mov-Elimination¶
Multi-Agent¶
Multimodal¶
Multimodal Generation¶
Multimodal Model¶
Muon¶
NAS¶
NCCL¶
- Echo Training Simulator
- FuseLink Multi-NIC Communication
- GPU-Centric Communication
- SimAI Architecture
NDS¶
NFT¶
NIXL¶
NLP¶
NPU¶
- AI Chip System Barriers
- AI Infra Daily Radar
- BSND TND Operator Layout
- Kimi K3 NPU Training
- Mooncake NDS Integration
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- NPU Training Operators - MC2
- NPU Training Operators - RoPE MRoPE
- VeRL Async Policy
- VeRL Performance Optimization
NUMA¶
NV¶
NVIDIA¶
NVIDIA DYNAMO¶
NVLINK¶
NVLink¶
NVME¶
NVMEOF¶
NVSHMEM¶
Nas¶
NoC¶
Nvidia¶
O2¶
O3¶
OPD¶
OS¶
OSI¶
OSS¶
Observability¶
Omni¶
OmniNFT¶
Open Source¶
OpenCode¶
OpenLDAP¶
OpenMP¶
OpenSHMEM¶
OpenWRT¶
Openssl¶
Options¶
- AMD Epyc Compiler Options
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Intel Compile Options
OutOfOrder¶
PAGEDATTENTION¶
PAPER¶
PCIE¶
PCIe¶
PD Disaggregation¶
PGAS¶
PIM¶
PKI¶
PLMR¶
PML¶
PP¶
PPO¶
PPT¶
PREFILL-DECODE¶
PT¶
PTA¶
- Debug/Profile/Devlop Tools of PTA
- [DevLog] PLAN
- [DevLog]24Q3P1 - Optimize PTA With Thread Affinity
- [DevLog]24Q4P2 - Lazy Initialize during old dispatcher way
PTX¶
PVE¶
Pagerank¶
Parallel¶
Parallelism¶
Perf¶
Performance¶
Performance Engineering¶
Performance Model¶
Performance Modeling¶
Performance Simulation¶
Pi¶
PicBed¶
PicGo¶
Pin¶
Pipeline¶
Post Training¶
Post-Training¶
PostTraining¶
Powershell¶
Prefill¶
Presentation¶
Priority¶
Privacy¶
Probability Theory¶
Product Discovery¶
Professional Skills¶
Prompt Engineering¶
PyG¶
PyPTO¶
PyTorch¶
PyTorch Profiler¶
Pytorch¶
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 2.5 :Dataset & Dataloader
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
- Pytorch 7 :Memory Optimization(Freeing GPU/NPU Memory Early)
- Pytorch 8 :Hyperparameter
QAT¶
QCC¶
QuantitativeFinance¶
Quantization¶
Qwen3-Omni¶
Qwen3-VL¶
Qwen3.5¶
- BSND TND Operator Layout
- Inference Quantization Formats
- NPU Training Operators - GDN
- NPU Training Operators - GMM
- Training Performance Model
RAG¶
RAID¶
RAW¶
RDMA¶
- AIV Direct Drive
- Ascend Communication Axes
- Ascend Communication Stack
- Ascend Data Movement Evolution
- Distributed KV Cache Management
- DualPath KV Loading
- FuseLink Multi-NIC Communication
- Memory Semantics vs RDMA
- Mooncake
- Mooncake Store Design
- Mooncake vLLM Ascend
RGB Lab¶
RHLF¶
RISC¶
RISC-V¶
RL¶
- AI Post Traning: DPO + MPO
- AI Post Traning: DanceGRPO
- Agent & Agentic RL
- DiffusionNFT
- Fast Debug: VeRL example
- Frontier Model RL
- Kimi K3 Report
- Memory-Semantic TransferQueue
- Multimodal Generation Evaluation
- Multimodal RL
- RL Algorithms: PPO-RLHF & GRPO-family
- RL DFX Metrics
- RL Data Flow
- RL Infra Series
- RL Next: Meta-Learning
- RL: Training Inference Mismatch
- RL: xPU Mismatch - metrics
- The Mechanics of RL: How Inference Sampling Shapes the Probability Landscape
- Train Stages: Pretrain, Mid-Train(CT), SFT, RL
- VLM RL Evaluation Datasets
- VeRL
- VeRL Async
- VeRL Backend Parallelism
- VeRL Checkpoint
- VeRL Feature Matrix
- VeRL Feature Survey
- VeRL Performance Optimization
- VeRL Rollout Inference
- VeRL Router Replay
- VeRL Speculative Decoding
- VeRL Training Flow
RLHF¶
RLInfra¶
RMA¶
ROMA¶
RTX 3090¶
Ramulator¶
Register-Renaming¶
Reinforcement Learning¶
Relationship Design¶
Reliability Engineering¶
Rental¶
Reporting¶
Requirements Engineering¶
Ring Attention¶
Risk Management¶
RoCE¶
RoPE¶
RouterReplay¶
Rust¶
SASS¶
SCADA¶
SCALE-UP¶
SDMA¶
SE¶
SFT¶
SHMEM¶
- AIV Direct Drive Programming
- AIV MoE All-to-All
- Ascend Communication API Evolution
- Ascend Communication Runtime Evolution
- Ascend Communication Stack
- DMA Communication Terminology
- Memory Semantics vs RDMA
- Memory-Semantic TransferQueue
- SHMEM Symmetric Memory
- TileXR-SHMEM Abstractions
SIMD¶
SIMT¶
SLIC¶
- Hybrid Multithreaded/OpenMP + MPI parallel Programs
- IPCC Preliminary SLIC Analysis
- IPCC Preliminary SLIC Analysis part2 : Run process
- IPCC Preliminary SLIC Analysis part3 : Hot spot analysis
- IPCC Preliminary SLIC Analysis part4 : cluster environment
- IPCC Preliminary SLIC Case1/2/3
- IPCC Preliminary SLIC Optimization 2
- IPCC Preliminary SLIC Optimization 3
- IPCC Preliminary SLIC Optimization 4: EnforceLabelConnectivity
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- IPCC Preliminary SLIC algorithm
- IPCC Preliminary SLIC test
SMA¶
SNAPSHOT¶
SP¶
SPDK¶
SRAM¶
SSD¶
SSD OFFLOADING¶
SSE¶
STEPFUN¶
STORAGE-NEXT¶
SWAP¶
Scale-Out¶
Scale-Up¶
Scaling Law¶
Scratchpad¶
Sequence Diagram¶
Server¶
Services¶
SimAI¶
Simulation¶
Skill¶
Skylake¶
Slurm¶
Software Architecture¶
SpeculativeDecoding¶
Staleness¶
Stanford¶
State Machine¶
Streaming¶
Subagent¶
SuperPixel¶
Survey¶
Swimlane Diagram¶
System Optimization¶
Systemclt¶
Systemd¶
TCP/IP¶
TENSORRT-LLM¶
TENT¶
- Mooncake Classic vs TENT Engine
- Mooncake Codebase Architecture
- Mooncake File vs Block Device
- Mooncake NDS Integration
- Mooncake TENT GDS
- Mooncake TENT Request Path
TILEXR¶
TLB¶
TLP¶
TND¶
TP¶
TRANSFER ENGINE¶
- Mooncake Classic NVMeoF Transport
- Mooncake Classic vs TENT Engine
- Mooncake Codebase Architecture
- Mooncake File vs Block Device
TRANSPORT¶
TURBOBUS¶
TUTTI¶
Team Architecture¶
Technical Due Diligence¶
Technical Feasibility¶
Technical Leadership¶
Technology Radar¶
Technology Strategy¶
Terminal¶
Threat Modeling¶
Thunderbolt¶
TileXR¶
Topcoder¶
Training¶
Training Systems¶
TransferQueue¶
Transformer¶
Transformerless¶
Triton¶
Type-C¶
UB-Mesh¶
UBSHMEM¶
UCX¶
UDMA¶
- Ascend Communication Stack
- Ascend Data Movement Evolution
- Ascend DeepEP MoE Dispatch Optimization
- DMA Communication Terminology
- TileXR-SHMEM Abstractions
UPI¶
URMA¶
USP¶
Ubuntu¶
UltraEP¶
Ulysses¶
Ulysses-Offload¶
UniGRPO¶
UnifiedBus¶
- Ascend Communication Axes
- Ascend Communication Stack
- Ascend Data Movement Evolution
- Ascend Interconnect Evolution
- DMA Communication Terminology
- Memory Semantics vs RDMA
- UB-Mesh Architecture
Unreal¶
Useful¶
VBench¶
VEFX-Bench¶
VERL¶
VLA¶
VLLM¶
- InferenceX Data Pipeline
- Mooncake Store Design
- Mooncake vLLM Ascend
- PagedAttention
- vLLM KV Offloading and GDS
VLM¶
- Frontier Model RL
- Ideas around Vision-Language Models (VLMs) / Reasoning Models
- Multimodal Generation Evaluation
- VLA VLM + DiT
- VLM RL Evaluation Datasets
VNC¶
VP¶
VPN¶
VR¶
Value Capture¶
Value of Information¶
VeOmni¶
VeRL¶
- AI Infra Daily Radar
- BSND TND Operator Layout
- Memory-Semantic TransferQueue
- NPU Training Operators - GDN
- Training Performance Model
- VLM RL Evaluation Datasets
- VeRL Async
- VeRL Async Policy
- VeRL Backend Parallelism
- VeRL Checkpoint
- VeRL Feature Matrix
- VeRL Feature Survey
- VeRL Performance Optimization
- VeRL Rollout Inference
- VeRL Router Replay
- VeRL Speculative Decoding
- VeRL Training Flow
VeRL-Omni¶
Vectorization¶
Video¶
Virtual memory¶
Visualization¶
Vscode¶
Vue¶
W8A4¶
W8A8¶
WAFER-SCALE¶
WAFERLLM¶
WAR¶
WAW¶
WBS¶
WLAN¶
WQE¶
Work Management¶
X86¶
XCCL¶
XLA¶
XTuner¶
Xeon¶
ZeRO++¶
ZeRO-Offload¶
ZenFlow¶
Zero-Idiom¶
Zsim¶
abi¶
agent¶
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
ai¶
- AI Documentation Workflow
- AI Model Visualization
- Building Large-Scale AI Systems on Ascend: Training, Inference, and Multimodal Optimization
- Business Trip: 2601-2602 verl + DanceGRPO
- Deploy OpenLLM to one A100
- HPCAI
- My Digital Worker
- My Digital Worker : AutoMoneyMaker - AutoTrader
- My Digital Worker : Model / Software Usage
- My Digital Worker : New Coding Way
- My Digital Worker : New Coding Way Part0 —— Building AI-Coding Env
- My Digital Worker : Work with AI
- Personal Advantage Workflow
- Probability Theory
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 2.5 :Dataset & Dataloader
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
- Pytorch 7 :Memory Optimization(Freeing GPU/NPU Memory Early)
- Pytorch 8 :Hyperparameter
algorithm¶
amd¶
ampere¶
anaconda¶
anime¶
- Anime Super Resolution to 4K & Interpolation to 120 fps
- Diary 230827: 上海二次元之旅
- UnimportantView: Anime Recommendation
- UnimportantView: Film & TV(Anime) Works Rating
aocc¶
apache¶
app¶
apple¶
apt¶
architecture¶
- GPU
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Out-Of-Order execution(OoOE/OOE) & Register Renaming
- Microarchitecture: Pipeline of Intel Core CPUs
arm¶
array¶
ascend¶
assembly¶
auto¶
avx¶
avx256¶
bank¶
batch¶
benchmark¶
bilibili¶
blog¶
broadcast¶
bug¶
bugs¶
business¶
c¶
c++¶
cache¶
calloc¶
cat¶
centos¶
chatgpt¶
chivier¶
chsh¶
clang¶
clash¶
- Clash Config 4 yourself
- Clash on LAN/linux/Dockers
- ClashX Pro and Wireguard on Macbook In School Net
- OpenWRT on router
- Wireguard
class¶
cloud¶
clustering¶
cmake¶
code¶
color¶
comic¶
command¶
commands¶
compile¶
compile options¶
compression¶
conda¶
context switch¶
cpp¶
cpu¶
cpu flags¶
crontab¶
cs¶
css¶
ctags¶
cuda¶
- Cuda Optimize
- Cuda Optimize : Stencil
- Cuda Optimize : Vectorized Memory Access
- Cuda Program Basic
- Nvidia Nsight
- Nvprof
- The CUDA Execution Model
cuobjdump¶
dLLM¶
ddns¶
debug¶
deepseek¶
diagram¶
disassembly¶
disk¶
- Disk
- Disk C: make room for installation
- Migrate From Synology DS220J to UGREEN DX4600
- Mount Network Disk
- Nas Disk Speed Test
- Ubuntu server reInstall
dit¶
div¶
dnat¶
dnf¶
docker¶
docuwiki¶
dokuwiki¶
domain¶
domestic¶
dram¶
echart¶
email¶
english¶
enter¶
entertainment¶
epic¶
epyc¶
ethernet¶
family¶
filesystem¶
firewall¶
firewalld¶
flags¶
flood fill¶
focusk¶
fork¶
forward¶
fpic¶
fstab¶
ftp¶
fun¶
- (Research) Team/Lab Organization
- 0 Overview
- 1.1 Living Needs & Meaning
- 1: Target2chase
- 2.2 Social Part
- 2: Courage to move on
- 3 EfficientJumpingRunning
- 3.2 taskPriority
- 3.3 EfficientWorkLearning
- 6 FPS
- AI Hardware & Accelerators
- AI Infra: 10k-GPU cluster
- AI Training Optimization
- Anime Auto add Chinese Subtitle
- AntiCheat
- Audio
- Balance (Efficient) work & life time
- Benchmark
- Blind Date 1st
- Blind Date 1st(2)
- Blind Date Tips
- Blog writing
- Burnout Monitor : Healthy Body Model + hair/heart-aware exercise
- CProgramReading
- CV Model
- Chrome://tracing
- Classical AI Models
- Colorful Life (TOP)
- Cuda Driver Runtime
- Data Link
- Data Structure Summary
- Datastruture: Tree
- DeviceExpansion
- Disease And Prevention
- Disordered Ideas
- Excel
- Experiments For PIM Motivation
- FPGA
- Financial Values
- Future Plans: House & Car
- GUIAgents
- Go Templates
- HTML
- Host-Core With PIM-Core In 3D-stacked Mem
- Huawei Ascend Domain-Specific Architectures : DaVinci
- Important Date
- Inference Basic
- Inference Optimization
- Japanese
- Keyboard
- LLM Model
- LLM Model Basic
- Lab homepage Template & Website builders choice
- LinuxFolderInstall
- Mathematical Logic & Algebraic structure
- Mkdocs
- Muon Optimizer
- Muon Optimizer + FSDP
- OOTD: outfit of the day
- Open &Free Multimodel AI Tools
- OpenCL Basic
- OpenWRTNetworkManage
- Parallel_sort
- Personal Image Management
- Php
- Piano: Transcribe Piano Sheet Music from Video using AI model
- Postgraduate dormitory
- Predictor
- RL Weekly News
- Research logic
- Salary & Tax & Insurance
- Scientifically Concocting Glasses
- Search, Ads, and Recommendation
- Security
- Social Science
- Synology terminal
- TODO
- Team Cooperation / Relationship
- Tmux
- Topology
- Training Data Usage
- Turing Machine & P versus NP problem
- URLs
- UnimportantView: Film & TV(Anime) Works Rating
- UnimportantView: Game
- User Kernel Mode
- Wake On Lan(Wol)
- Weekly
- When & How 4 team presentation page & knowledge database pool
- 宛如泥潭的大型项目开发困境
function call¶
game¶
gateway¶
gcc¶
- C program compile&run process
- GCC Compile Error
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Inline Assembly
gdb¶
gdbgui¶
gef¶
gem5¶
gif¶
git¶
- Git Lfs
- Git Push 2 Homepage
- Git Standardization
- Git Submodule: Data & Code Repository Separate
- Introduction to Git Commands
- Network Basic
github¶
github action¶
glibc¶
gnu¶
go¶
- Go Install and Command
- Go mod
- Golang Syntax
- Web Design 2 : Content Organization & Link Content using go template
golang¶
gprof¶
gpt¶
gpu¶
graph¶
group¶
gzip¶
h265¶
hash¶
header¶
health¶
heap¶
hexo¶
home¶
homepage¶
- Git Push 2 Homepage
- Homepage Template Conflict
- How SSG Get Work? & Hugo theme creation
- Miscellaneous
- Web Server: Nginx V.S. Apache2
hpc¶
html¶
htop¶
http¶
huawei¶
- Forecasting Housing Prices Around Huawei Xicen Base: A Home Buying Plan for Qingpu District
- Huawei Kunpeng workload
- Travel and Business Trip Checklist
hugo¶
hw¶
icc¶
icecream¶
icpc¶
incomplete¶
- AOCC
- Cache
- GCC Compiler Option 1 : Optimization Options
- GCC Compiler Option 2 : Preprocessor Options
- Memalloc
inference¶
- SGLang
- Speculative Decoding & eagle3
- Vllm Basic
- vLLM Inference Profiling
- vllm-omni & DiT Inference Accelerate
initd¶
intel¶
- Intel Compile Options
- Intel Pin
- Intel SDM(Software Developer's Manual)
- Microarchitecture: Micro-Fusion & Macro-Fusion
- Microarchitecture: Zero (one) idioms & Mov Elimination
- Old Pintool Upgrade with newest pin
interval model¶
ip¶
- IP Forward
- IPV4 && IPV6
- Linux Auto Run : crontab
- Linux Network Command Guide
- ServerLogin
- USTC Network Information Center
ipcc¶
- Hybrid Multithreaded/OpenMP + MPI parallel Programs
- IPCC Preliminary SLIC Analysis
- IPCC Preliminary SLIC Analysis part2 : Run process
- IPCC Preliminary SLIC Analysis part3 : Hot spot analysis
- IPCC Preliminary SLIC Analysis part4 : cluster environment
- IPCC Preliminary SLIC Case1/2/3
- IPCC Preliminary SLIC Optimization 1
- IPCC Preliminary SLIC Optimization 2
- IPCC Preliminary SLIC Optimization 3
- IPCC Preliminary SLIC Optimization 4: EnforceLabelConnectivity
- IPCC Preliminary SLIC Optimization 5: MPI + OpenMP
- IPCC Preliminary SLIC Optimization 6: Non-blocking MPI
- IPCC Preliminary SLIC algorithm
- IPCC Preliminary SLIC test
- Training course - IPCC 5 Optimize common tools
- VtuneOptimize
iptable¶
iptables¶
ipv4¶
ipv6¶
ispc¶
ive¶
izone¶
j4125¶
java¶
jekyll¶
job¶
jpg¶
k-fold cross validation¶
- Pytorch 1 :Basic Components
- Pytorch 2 :more conceptions about training and inference
- Pytorch 3 :Model & Training
- Pytorch 4 :Save & Load & Pretrain
- Pytorch 5 : Distributed Training & Parallelism (Mem Re-allocated)
- Pytorch 6 :Visualization
kill¶
kpop¶
kunpeng 920¶
latex¶
linux¶
live2d¶
llm¶
llvm¶
- Kunpeng
- LLVM Mca : huawei HiSilicon's TSV110 work
- LLVM Mca :with BHive (2019)
- LLVM-MCA: Install&RunTests
- LLVM-MCA: docs
- llvm
- llvm Backend
- llvm Pass
llvm-mca¶
- LLVM-MCA: docs
- Static Code Analysis
- uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures (2019)
local-dev¶
localhost¶
logic core¶
loop¶
lscpu¶
lsof¶
macbook¶
make¶
malloc¶
manga¶
map¶
master¶
mca¶
- Kunpeng
- LLVM Mca : huawei HiSilicon's TSV110 work
- LLVM Mca :with BHive (2019)
- LLVM-MCA: Install&RunTests
- LLVM-MCA: docs
memory¶
micro-op fusions¶
minicoda¶
mkdocs¶
moonlight¶
mount¶
mpi¶
mpicc¶
mpiicc¶
mtr¶
nD-FullMesh¶
name¶
nas¶
nasm¶
neon¶
network¶
networkmanager¶
newline¶
nginx¶
npm¶
nsight¶
nvidia¶
nvprof¶
omni¶
oneApi¶
oneapi¶
openmp¶
openmpi¶
openvpn¶
opkg¶
optimization¶
optimize¶
optimizer¶
overleaf¶
page table¶
pam¶
paper¶
- DualPath KV Loading
- FuseLink Multi-NIC Communication
- GPU-Centric Communication
- GPU-Initiated I/O
- LLVM Mca :with BHive (2019)
- Latent Dirichlet Allocation (2003)
- Micro2023: Utopia
- NVIDIA SCADA
- TurboBus PCIe Bandwidth Pooling
- uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures (2019)
parallel¶
password¶
pdf¶
perf¶
piano¶
ping¶
pip¶
png¶
podman¶
port¶
ppt¶
prefetch¶
price¶
process¶
profiling¶
proxy¶
- Clash on LAN/linux/Dockers
- Cloudflare warp proxy
- Introduction to Git Commands
- Network Basic
- SSHForward
ps¶
pull request¶
pyc¶
pyo¶
pypy¶
python¶
- Debug/Profile/Devlop Tools of PTA
- Latent Dirichlet Allocation (2003)
- Pip Package
- Python
- Python Class
- Python Graph & Visualization
- Python MPI
- Python: DataStructure
- PythonRegex
- WebCrawler first try
pytorch¶
qt¶
ram¶
ranking¶
rar¶
rating¶
readelf¶
reboot¶
reduce¶
regex¶
register¶
registers¶
report¶
rl¶
route¶
router¶
rpm¶
safety¶
scons¶
scss¶
security¶
segmentation fault¶
server¶
sh¶
simd¶
skill¶
slime¶
smanga¶
smb¶
snap¶
snat¶
sniper¶
socket¶
speaking¶
sql¶
sram¶
ssh¶
ssl¶
stack¶
- CSAPP: Machine Programming III: Procedures
- Linux Executable file: Structure & Running
- [C++ Basic] STL Data Structure
stencil¶
step-video¶
stream¶
switch¶
synology¶
systemctl¶
systemd¶
t2v¶
- 250217 Step-Video-T2V Reading & Porting
- 260117 Step-3-VL 10B
- AI Model Memory
- Ideas around T2I2V models
- Ideas around Vision-Language Models (VLMs) / Reasoning Models
- World Model/UFMs/Omni-Modal: AR vs DiT
tar¶
tcp¶
tcpdump¶
terminal¶
thread¶
tick¶
tips¶
tlb¶
tmm¶
tmpfs¶
tools¶
top¶
topology¶
torch¶
torch.dist¶
torch_npu¶
torchrun¶
traceroute¶
uTorrent¶
udp¶
ufw¶
ugreen¶
unix¶
unravel¶
usb¶
user¶
useradd¶
usermod¶
vLLM¶
vLLM-Ascend¶
vLLM-Omni¶
vector¶
vectorization¶
verl¶
video¶
vim¶
virtual machine¶
visualization¶
vllm¶
- Speculative Decoding & eagle3
- Vllm Basic
- vLLM Inference Profiling
- vllm-omni & DiT Inference Accelerate
vlm¶
vpn¶
vscode¶
vscode debug¶
vtune¶
wake¶
warp¶
webdav¶
website¶
- Web Design 1 : Layout Overview
- Web Design 2 : Content Organization & Link Content using go template
- Web Design 3 : Future Features
- Web Design 4 : Customize Markdown Grammar In SSG
wget¶
wifi¶
windows¶
wireguard¶
- ClashX Pro and Wireguard on Macbook In School Net
- Github Access
- OpenWRT on router
- Wireguard
- Wireguard Server 2 Server in OpenWRT