DEV Community

# llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
KVQuant: Run 70B LLMs on 8GB RAM with Real-Time KV Cache Compression

KVQuant: Run 70B LLMs on 8GB RAM with Real-Time KV Cache Compression

1
Comments
1 min read
I Built a Knowledge Base That Thinks — Inspired by Karpathy’s LLM Wiki

I Built a Knowledge Base That Thinks — Inspired by Karpathy’s LLM Wiki

5
Comments
6 min read
Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications

Cencori: A Serverless Infrastructure Layer for Secure and Scalable AI Applications

2
Comments
5 min read
KVQuant: Run 70B LLMs on 8GB RAM with 4-bit KV Cache Quantization

KVQuant: Run 70B LLMs on 8GB RAM with 4-bit KV Cache Quantization

Comments
1 min read
Securing Agentic Workflows: A Deterministic 'Human-in-the-Loop' Pattern for LLMs

Securing Agentic Workflows: A Deterministic 'Human-in-the-Loop' Pattern for LLMs

Comments
5 min read
software engineers are becoming reliability engineers for generated output

software engineers are becoming reliability engineers for generated output

Comments
5 min read
I just wanted to chat with my Raspberry Pi.

I just wanted to chat with my Raspberry Pi.

Comments
9 min read
Fix Your Prompt Structure Before You Touch Your Infrastructure

Fix Your Prompt Structure Before You Touch Your Infrastructure

Comments
4 min read
Why File-to-Markdown Conversion Is Becoming an AI Input Layer

Why File-to-Markdown Conversion Is Becoming an AI Input Layer

Comments 1
7 min read
I Compressed GPT-2 to Run on an Arduino

I Compressed GPT-2 to Run on an Arduino

Comments
1 min read
How LLMs Memorize Phone Numbers (and How Labs Stop It)

How LLMs Memorize Phone Numbers (and How Labs Stop It)

Comments
7 min read
I Let My AI Agent Run Overnight. It Cost $437.

I Let My AI Agent Run Overnight. It Cost $437.

1
Comments
5 min read
TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

TurboQuant on a MacBook Pro, part 2: perplexity, KL divergence, and asymmetric K/V on M5 Max

Comments
8 min read
Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Why I'm Building a Local-First AI Coding Workspace (And How Behavioral Routing Makes It Work)

Comments
6 min read
Prompt Caching Works. Your Prompt Assembly Code Does Not.

Prompt Caching Works. Your Prompt Assembly Code Does Not.

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.