Cloud Native Deep Dive
  • Home
  • About The Author
  • Main Website
Sign in Subscribe
Distributed LLM Inference on Kubernetes: Configuring KServe and llm-d

Distributed LLM Inference on Kubernetes: Configuring KServe and llm-d

18 Jul 2026 6 min read AI Inference
Open Weight Models (Qwen, DeepSeek, Kimi) need a place to run that gives the ability to have shareable GPUs, a first-class orchestration/scheduler, and traffic routing capabilities that engineers are comfortable with.
Implementing Observability For Agent Substrate Actors

Implementing Observability For Agent Substrate Actors

18 Jul 2026 8 min read Agent Substrate
Sandboxed Agents means we're going a level deeper in terms of where AI runs. Originally, it could be an Agent Harness like opencode or Codex running on your local terminal. Now,
Live Actor Migration in Agent Substrate: Moving State Across Workers

Live Actor Migration in Agent Substrate: Moving State Across Workers

11 Jul 2026 5 min read Agent Substrate
High Availability (HA) is the cornerstone of an efficient, scalable, performant system. Without it, zero downtime and migrating/scaling state for systems would be impossible. In this case, the "system" is
Virtual Keys in Agentgateway: Per-User Token Budgets for Your LLM Gateway

Virtual Keys in Agentgateway: Per-User Token Budgets for Your LLM Gateway

04 Jul 2026 5 min read agentgateway
Cost in AI will vary per user and department. If an engineer is refactoring a codebase or observing an environment to fix anomalies, the token spend may differ from that of someone in
How Agentgateway Makes LLM & MCP Token Spend Visible

How Agentgateway Makes LLM & MCP Token Spend Visible

28 Jun 2026 5 min read AI Cost
Tokens turn into dollars, and the tokens are expanded upon in many different ways. Input tokens, output tokens, agent context racking up tokens due to ingesting MCP Server tools, and token output increasing
Previous
Page 2 of 24
Next
Cloud Native Deep Dive © 2026
  • Sign up
  • LinkedIn
Powered by Ghost