DEV Community

Mingxin Technology profile picture

Mingxin Technology

Storage acceleration for LLM inference. KV-cache tiering on NVMe-oF: +29-40% throughput (signed benchmarks, reproducible). mingxinstorage.xyz

Joined Joined on 
A Comparative Framework for KV Cache Performance on Domestic Accelerators: Measured Insights

A Comparative Framework for KV Cache Performance on Domestic Accelerators: Measured Insights

Comments
5 min read
KV Cache Storage Product Comparison: Methodology and Measured Benchmarks

KV Cache Storage Product Comparison: Methodology and Measured Benchmarks

Comments
5 min read
KV Cache Memory vs. Hit Rate: Why Larger Caches Yield Diminishing Returns

KV Cache Memory vs. Hit Rate: Why Larger Caches Yield Diminishing Returns

Comments
4 min read
The Boundary Between External Memory and VRAM for KV Cache in Local LLM Deployment

The Boundary Between External Memory and VRAM for KV Cache in Local LLM Deployment

Comments
5 min read
Optimizing Compute Rental Costs: Dynamic Scaling and On-Demand Allocation Strategies

Optimizing Compute Rental Costs: Dynamic Scaling and On-Demand Allocation Strategies

Comments
5 min read
Measured Application Analysis of Domestic KV Cache Products in AI Inference

Measured Application Analysis of Domestic KV Cache Products in AI Inference

Comments
5 min read
Three Clause Types Often Missed in Compute Rental Contracts

Three Clause Types Often Missed in Compute Rental Contracts

Comments
5 min read
Energy Consumption Assessment and Optimization Paths for Data Center-Scale KV Cache Deployment

Energy Consumption Assessment and Optimization Paths for Data Center-Scale KV Cache Deployment

Comments
7 min read
Choosing Compute Rental: Focus on Three SLA Metrics, Not Unit Price

Choosing Compute Rental: Focus on Three SLA Metrics, Not Unit Price

Comments
5 min read
Bandwidth and Latency Optimization Strategies for KV Cache in Edge Computing

Bandwidth and Latency Optimization Strategies for KV Cache in Edge Computing

Comments
5 min read
Four Key Dimensions for Evaluating the Reliability of Domestic AI Storage

Four Key Dimensions for Evaluating the Reliability of Domestic AI Storage

Comments
6 min read
Three-Layer Adaptation of the Ascend Inference Stack: Driver, Operator, and Framework Are All Indispensable

Three-Layer Adaptation of the Ascend Inference Stack: Driver, Operator, and Framework Are All Indispensable

Comments
6 min read
Can Compressed Sensing Be Used for Inference Storage Data Compression?

Can Compressed Sensing Be Used for Inference Storage Data Compression?

Comments
4 min read
NVMe-oF vs. RDMA: Performance Comparison in Inference Storage

NVMe-oF vs. RDMA: Performance Comparison in Inference Storage

Comments
5 min read
Storage Selection Strategy for Inference in Domestic AI Computing Centers

Storage Selection Strategy for Inference in Domestic AI Computing Centers

Comments
5 min read
How KV Cache Prefetch Cuts Storage Latency

How KV Cache Prefetch Cuts Storage Latency

Comments
6 min read
KV Cache Reuse in Multi-Turn Dialogue: 29% Throughput Gain Measured

KV Cache Reuse in Multi-Turn Dialogue: 29% Throughput Gain Measured

Comments
5 min read
Evaluating Long-Context Inference Performance of Domestic Accelerator Cards

Evaluating Long-Context Inference Performance of Domestic Accelerator Cards

Comments
4 min read
Storage Bottlenecks in Real-Time Video Inference and a Tiered Acceleration Approach

Storage Bottlenecks in Real-Time Video Inference and a Tiered Acceleration Approach

Comments
5 min read
KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

KV Cache Reuse in Multi-Turn Dialogue: A Deployment Case Study

Comments
5 min read
How KV Cache Pooling and Sharing Improves Inference Resource Utilization

How KV Cache Pooling and Sharing Improves Inference Resource Utilization

Comments
5 min read
Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Distributed KV Cache Load Balancing: Strategies and Measured Trade-offs

Comments
6 min read
Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Measured Evaluation of KV Cache Acceleration for Inference in Online Education Scenarios

Comments
4 min read
Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Deployment Path and Measured Performance of Domestic AI Inference Acceleration Cards in Real-Time Database Queries

Comments
5 min read
Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Case Analysis of Huawei OceanStor UCM Inference Acceleration Solution

Comments
5 min read
Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Why KV Cache Tiering Is the Key TCO Optimization Point in a Three-Tier Storage Architecture for Compute Centers

Comments
6 min read
The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

The Hidden Cost of Model Switching: A Measured Path from 46.7% to 62.8% Effective Compute Utilization

Comments
5 min read
Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Cross-Platform Methodology Porting: Validating KV Cache Memory Efficiency from Muxi N260 to MI308X

Comments
3 min read
Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Inference Deployment in Xinchuang Scenarios: Architecture Essentials for Data Not Leaving the Domain

Comments
5 min read
Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Clos Network Architecture: A Cost-Effectiveness and Selection Framework for Thousand-Card Inference Clusters

Comments
7 min read
How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

How to Determine the Storage-to-Compute Ratio for Inference Clusters: Measured Basis for One Array Serving 8 Nodes

Comments
6 min read
Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Engineering Practice Essentials of NVMe-oF + RoCEv2 in Inference Storage Scenarios: A Case Study with FX100 Measurements

Comments
5 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

Comments
5 min read
loading...