Open Weight Models (Qwen, DeepSeek, Kimi) need a place to run that gives the ability to have shareable GPUs, a first-class orchestration/scheduler, and traffic routing capabilities that engineers are comfortable with.
Cost is everything. In just about every agentic conversation, the three things that come up for enterprises implementing AI workloads are:
1. Cost
2. Observability
3. Security
and as AI continues to throw
There are many Agentic creation frameworks ranging from CrewAI to kagent to langchain and several others which are typically written in Python or JS. If you're an engineer working on Kubernetes,
There are two types of Models/LLMs you see in today's Agentic world:
1. "SaaS-based Models", which are Models that are managed for you (Claude, Gemini, GPT, etc.