AI Stack Radar

August 19, 2026

29 hosted AI products worth knowing about this week, sorted by the job they do — this is the first run of the series, so nothing here has appeared before.

🤖 Agents & Orchestration

Letta

Most agent frameworks reset to zero at the end of a session; Letta's whole pitch is that yours doesn't have to. It runs a tiered memory system — core, archival, recall — that agents write to and pull from across conversations, and ships an Agent Development Environment where you can watch the context window and memory blocks update live instead of guessing what the model saw. Replaces: the vector-DB-plus-prompt-stuffing hack most teams build for agent memory by hand. Pricing: free tier for 3 agents (bring your own key), Pro from $20/mo.

Gumloop

Gumloop's drag-and-drop canvas connects to 300+ business tools and keeps a synced "company brain" the agents can query instead of starting from a blank prompt every run. Built-in eval tracking lets an agent tune its own workflow over time rather than staying static after launch. Replaces: Zapier-plus-a-manual-LLM-call setups for teams that want agents making decisions, not just following triggers. Pricing: 14-day free trial, Pro from $37/mo plus a small orchestration fee on usage.

CrewAI AMP

CrewAI AMP is the hosted management layer bolted onto the open-source CrewAI framework: a visual studio, one-click deploys, tracing, and guardrails the free framework doesn't include. Teams that can't send data outside their network can run the same tooling in self-hosted "Factory" containers instead. Replaces: the observability stack teams currently duct-tape around CrewAI's open-source core. Pricing: free Basic tier (50 workflow runs/month), Enterprise by quote.

🔌 MCP & Tool Plumbing

Composio

Composio hands an agent OAuth-managed access to 1,000+ toolkits without the agent ever holding a raw credential — tokens get refreshed and scoped behind the scenes, and tool calls execute in sandboxed environments with results written to a queryable filesystem. Replaces: the pile of hand-rolled OAuth integrations most agent stacks accumulate one SaaS API at a time. Pricing: free for 100K tool calls/month, Pro from $29/mo plus per-call overage.

Arcade.dev

Arcade.dev cares less about connecting tools and more about knowing who's using them: agents act under scoped, per-user permissions pulled from your existing identity provider, and every action lands in a central audit log. It recently absorbed Smithery, the MCP registry, folding that catalog into its own. Replaces: the shared-static-token setups most MCP integrations default to, and the audit gap that comes with them. Pricing: free for 2,000 auth events and 2,000 tool calls/month, Team from $25/mo plus usage.

Pipedream

Pipedream exposes its entire integration catalog — over 3,000 apps, 10,000+ tools — through a single MCP server with managed OAuth, or through an SDK if you'd rather embed the connectors directly. Queues and private networking come bundled in, which is more infrastructure than most tool-plumbing platforms bother shipping. Replaces: wiring together single-purpose API wrappers one at a time; a broader but less audit-focused alternative to Composio or Arcade. Pricing: free for 300 credits/month, Basic from $29/mo.

💻 Coding & Dev

Morph

Morph doesn't sell one model — it sells a handful of small, purpose-built ones. Fast Apply merges an LLM's suggested edits into your actual files at over 10,000 tokens per second, and a separate context-compression model keeps token costs down when an agent's working set balloons. It ships as an MCP server for Cursor, Claude Code, and Windsurf. Replaces: the slow "rewrite the whole file" step between a frontier model's diff output and disk. Pricing: 200 free requests/month, then usage-based per million tokens.

Amp

Amp spun out of Sourcegraph into its own company at the end of 2025, carrying that codebase-search heritage into a terminal-and-web agentic coding tool. You pick a reasoning mode — low through ultra — depending on how much the task is worth, and heavier jobs can run remotely in containerized "orbs" instead of tying up your machine. Replaces: a terminal-first coding agent in the same slot as Claude Code or Codex CLI. Pricing: free ad-supported tier, paid tiers scale with reasoning mode.

Greptile

Greptile builds a graph of your whole codebase — files, functions, the dependencies between them — before it reviews a single pull request, which is how it catches bugs that span multiple files instead of just the diff in front of it. It also picks up your team's style preferences from the comments engineers leave on its suggestions. Replaces: the human first-pass reviewer for pattern-level bugs, sitting next to CodeRabbit and Qodo in the AI-review space. Pricing: free Starter tier (1 seat, 50 credits/month), Pro from $30/seat/month.

CodeRabbit

CodeRabbit's angle is triage, not just commentary: a "Change Stack" layer generates blast-radius visualizations and architecture diagrams, then ranks incoming PRs by risk before a human even opens them. It also runs continuous security scanning specifically aimed at catching problems in AI-generated code. Replaces: manual PR triage and a chunk of what a separate SAST tool would otherwise catch. Pricing: free trial, paid and enterprise tiers behind a quote.

🧠 Text & Reasoning APIs

Cerebras Inference

Cerebras runs inference on its own wafer-scale chip and claims north of 2,000 tokens per second on models like Llama 4 Scout — fast enough that it also sells flat-rate coding subscriptions ($50/mo for 24M tokens a day) instead of pure metered pricing. Swapping in is meant to take two lines of code since the API matches OpenAI's. Replaces: Groq or Together as the low-latency inference backend when an agent's response time matters more than which specific model answers. Pricing: $5 free trial credit, pay-per-token from there.

DeepInfra

DeepInfra hosts over 100 models — text, image, speech, video — behind one pay-per-token API, including open-weight frontier-adjacent releases from DeepSeek, Qwen, and Kimi at prices well under what the labs themselves charge. No contract, no minimum. Replaces: calling model providers directly when the priority is cost, not brand; sits next to Novita and Hyperbolic in the budget-hosting tier. Pricing: pay-as-you-go, no published free tier.

Portkey

Portkey puts one API in front of 1,600+ models and layers observability, PII-redaction guardrails, and spend governance on top — useful when the actual problem isn't picking a model, it's knowing what your agents did with it. Palo Alto Networks acquired the company in May and has started rebranding it as Prisma AIRS AI Gateway, though the self-serve signup still works under the original name. Replaces: the same slot as OpenRouter, but built for teams that need audit trails more than raw model choice. Pricing: free for 10,000 logs/month, Production from $49/mo.

OpenRouter

OpenRouter sits in front of more than 500 models across 80-plus providers with automatic failover if one goes down mid-request — buy credits once, spend them against whichever model fits the task. No subscription tier to pick, no per-provider account sprawl. Replaces: managing separate API keys for every model provider you touch. Pricing: pay-as-you-go credits from $10, no subscription.

🎨 Image & Design

Caimera

Caimera turns a flat-lay product photo into a full set of on-model catalog shots — sketch-to-product renders, print patterns, packshots — for fashion and e-commerce brands that would otherwise book a studio and a model for every SKU. H&M and Puma are both listed as customers. Replaces: a chunk of the traditional product photoshoot for online catalogs. Pricing: free trial, no card required; paid plans scale with volume.

Recraft

Recraft generates actual SVG vector output, not a raster image dressed up as one, and its brand kits keep colors and style consistent across a whole batch of generated icons or logos. The raw API prices per image rather than per subscription seat. Replaces: hand-building icon sets and vector logos in Illustrator, or paying for stock-vector libraries. Pricing: free tier is non-commercial and public-only; commercial plans start around $10-12/mo.

Flair AI

Flair generates bulk product photography for categories where a real photoshoot is expensive relative to the product price — beauty, jewelry, food — and its free tier actually includes a custom trained model plus five images a month, not just a demo. Replaces: freelance product photographers for e-commerce listings and ad creative at small-to-mid volume. Pricing: free tier available, Pro from $8/mo.

🎙️ Voice & Audio

Cartesia

Sonic, Cartesia's flagship model, streams text-to-speech at under 90 milliseconds to first audio across 40-plus languages, with 10-second voice cloning and inline tags for laughter or emotional inflection dropped right into the transcript. It's built for conversational latency, not narration. Replaces: ElevenLabs' API in voice-agent builds — call centers, IVR — where every millisecond of lag is audible to the caller. Pricing: free tier (20K credits/mo), Pro from $5/mo.

Fish Audio

Fish Audio just closed a $50M seed and claims 8 million users across its TTS, STT, and voice-cloning stack, with clones ready in 10-15 seconds and support for 30-plus languages. Unusually, it open-sources several of the underlying speech models even while selling the hosted version. Replaces: ElevenLabs or PlayHT for character-voice and audiobook narration workflows. Pricing: free tier (8,000 credits/mo, no card), Plus from $5.50/mo.

Hume AI

Hume splits into two developer products: EVI for real-time speech-to-speech conversation, and Octave, an LLM-built TTS engine, both sitting on top of an Expression Measurement API that scores 48-plus emotion categories across 50-plus languages. The company pitches itself these days more as an evaluation layer for voice AI than a straight TTS vendor. Replaces: stitching together ElevenLabs, Whisper, and an LLM by hand for voice agents that need to read tone, not just transcribe words. Pricing: free plan (10,000 characters + 5 EVI minutes/mo), paid from $3/mo.

Beatoven.ai

Beatoven turns a mood prompt into a royalty-free backing track through a model it calls Maestro, plus a separate sound-effects generator it bills as a personal AI foley artist. It's Fairly Trained certified, meaning the training data was licensed — worth knowing if a brand is nervous about music-copyright exposure. Replaces: stock-music libraries like Epidemic Sound for YouTube, podcast, and video scoring; narrower than a full song generator like Suno. Pricing: free previews, paid plans from about $2.50/mo for 15 minutes of downloads.

🔍 Research & Retrieval

Exa

Exa is a search API built for agents rather than browsers — sub-180ms responses, and a "highlights" endpoint that returns only the relevant snippet of a page instead of the whole document, cutting token spend by a claimed 90%. Cognition uses it inside Devin. Replaces: wiring an agent to Google or Bing and hand-rolling the scraping and relevance filtering yourself. Pricing: $20 signup credit plus $10/mo ongoing free tier, then per-request pricing.

Reducto

Reducto parses tables, charts, handwriting, and scans across 30-plus file types and returns structured JSON with bounding boxes attached, so citations point back to the exact spot on the page. A June feature called Deep Split handles multi-thousand-page documents across 150-plus category taxonomies. Harvey and Scale AI both run it in production. Replaces: AWS Textract or an open-source parser like unstructured.io as the ingestion layer feeding a RAG pipeline. Pricing: 15,000 free credits on signup, then $0.015/credit.

Ragie

Ragie handles the whole RAG pipeline — ingestion, chunking, hybrid vector-plus-keyword indexing, retrieval — and syncs directly from Google Drive, Notion, and Confluence instead of requiring a custom ETL job first. Agentic OCR pulls tables and forms out of scanned documents along the way. Replaces: assembling Pinecone, a chunker, an embedding pipeline, and retrieval logic by hand. Pricing: free Developer tier, Starter from $100/mo.

Vectorize

Hindsight, Vectorize's core product, gives agents memory that's meant to improve rather than just accumulate — remember, recall, and reflect endpoints spread across four separate memory networks, positioned against systems that just store facts without learning from them. It ships both as a hosted cloud service and an MIT-licensed open-source version. Replaces: the one-shot chat-history hacks agents use when they need to remember a specific user across sessions, not just retrieve documents. Pricing: free to self-host; Cloud runs $10 per million tokens stored.

🛡️ Evals, Observability & Guardrails

Galileo

Galileo turns offline evals into runtime guardrails using Luna, a family of small distilled judge models built to score 20-plus metrics at under 200ms and roughly $0.02 per million tokens — a fraction of what LLM-as-judge scoring usually costs. It deploys as SaaS, VPC, or fully on-prem. Replaces: hand-scripted LLM-as-judge harnesses plus a separate content-filter layer, folded into one product. Pricing: free for 5,000 traces/mo, Pro from $100/mo.

Helicone

Helicone is open-source and doubles as both a gateway and an observability layer — it routes requests across OpenAI, Anthropic, Azure, Together, and Groq while logging every session into a dashboard you can actually read. Prompt management and a built-in playground come with it. Replaces: the custom logging middleware most teams wrap around raw provider SDK calls. Pricing: free for 10,000 requests/mo, Pro from $79/mo.

LangWatch

LangWatch is OpenTelemetry-native and framework-agnostic, with integrations for LangGraph, DSPy, and the Vercel AI SDK, plus an agent-simulation feature that runs synthetic user scenarios — including voice — against your agent before a real user ever touches it. It also does automatic prompt optimization through Stanford's DSPy. Replaces: a bespoke eval-script-plus-tracing-tool combo; the open-source, self-hostable answer to LangSmith or Braintrust. Pricing: free forever up to 50,000 events/mo, Growth from €29/seat/mo.

Lakera Guard

Lakera Guard sits inline in front of any model call and blocks prompt injection, jailbreaks, data leakage, and toxic content before it reaches, or leaves, the model, claiming sub-50ms latency and a 0.01% false-positive rate in production. Replaces: a homegrown regex filter, or leaning entirely on a provider's built-in moderation endpoint. Pricing: free Community tier (10,000 requests/mo), Pro by quote.