Archive
Browse posts by year and month
2026
56 postsAugust 19 posts
- AI Agent Memory Lessons From LinkedIn's Hiring Assistant
- Why Speculative Decoding Pays Nearly 4x on CPUs
- AI Agent Token Usage Overtook Humans on OpenRouter
- How WhatsApp Scam Alert Detects Scams It Cannot Read
- AI-Assisted Product Launch Teardown of Stampli's 68% Claim
- Build an AI Text Detector, Then Test If It Can Ship
- The Real Cost of Gradient Accumulation on T4 and L4
- LLM Context Window Management With a Token Budget
- The Real OpenAI Ultrafast Mode Speedup, Workload by Workload
- LangChain vs LangGraph for Stateful AI Agent Orchestration
- ChatGPT Business Premium Pricing Decodes Agent Token Math
- Muse Glimmer Local Review Tests 30B Agents Under 20GB
- Structured Output Local LLM Tactics That Survive Production
- AI Agent Cyber Security Evaluation After the Astra Slowdown
- Agentic Loop Token Costs Are an Architecture Problem
- Megakernels in LLM Inference When Fusion Actually Wins
- Stacked Pull Requests Relocate AI Mega-PR Review Cost
- Disaggregated GPU Inference Hits the KV Cache Wall
- How The Copilot Prompt Injection Worm Spreads
July 28 posts
- LLM-Native Recommendation Architecture After Netflix GenRec
- How MCP Servers for AI Agents Bridge Fragmented Data
- Control Reasoning Effort LLM APIs in Production
- BM25 vs Dense Retrieval and SPLADE for Production RAG
- Shipping LLM Browser Agents Without Breaking Production
- Claude Opus 5 Prompt Injection Hits 0% Across 129 Tests
- AI Coding Agents Cut Per-Task Cost Up to 4x With Indexing
- ChatGPT Shared Link Vulnerability Plants Rogue Agents
- Grok vs Copilot in Excel for Builders
- Run AI Models Locally on Mac With MLX and Nativ
- Hugging Face Agentic Attack Redefines AI Security Response
- AI Agent Cost Per Resolution Decides If It Ships
- Self-Host LLM Inference, the Netflix Decision Framework
- Where AI in Media Production Workflows Actually Pays Off
- Orchestrate Open Source LLMs vs Frontier Models
- OpenAI Codex Subagent Encryption Breaks Agent Observability
- ChatGPT Sales Workflows Fail on Real CRM Data
- Real-Time AI Dental Image Verification Cuts Claim Denials
- Fine-Tuning vs RAG vs Prompt Engineering Decision Framework
- How LLM Agent Scaffolding Fixes Failing Code Review Agents
- Building Proactive AI Agents With Context Graphs
- AI Coding Benchmarks Are Gameable by Design
- AI Video Editor Limitations Break Iterative Workflows
- LLM Vendor Data Risk Has a Break-Even Price
- Stop RAG Hallucination with Typed Schema Contracts
- Domain-Specific LLM Evaluation Demands Leakage-Free Data
- Why an Agent-to-Agent Gateway Beats Point-to-Point Links
- LLM API Token Inflation Hides Claude Sonnet's Real Cost
June 9 posts
- Production LLM Application Security Beyond Prompt Injection
- Secure AI Coding Agents From Supply Chain Attacks
- Build Expert-in-the-Loop AI for Reliable Automation
- Build an LLM Fallback Strategy Before Access Drops
- Human-in-the-Loop Systems: Balancing AI Speed and Operational Safety
- Custom AI Inference Chips Are Eating the GPU Market
- The MLOps Playbook: Choosing Cloud GPU Providers for LLM Inference
- The End of AI Reasoning Transparency: When Chain of Thought Becomes a Summary
- Why AI Agents Leak Sensitive Data (and How to Stop Them)