Superpowers is an MIT-licensed agentic skills framework and software-development methodology that turns an approved design/spec into an implementation plan, then drives subagent execution, TDD, code review and branch completion across many coding-agent hosts.
The AI developer tooling compass
Explore the AI tooling ecosystem
Discover AI coding agents, frameworks and integrations. Explore capabilities, implementation details and licensing.
- 441 tools
- 17 categories
- 354 open-source entries
- Snapshot
Interactive search could not be loaded. You can still browse the tools and visit their websites.Reload this page to try again.
No tools match your search and filters
Try removing a filter or clearing the current search.
A configurable collection of agent skills, rules, hooks, memory and checks for coding assistants.
Small, composable engineering skills intended to extend an existing coding-agent workflow without taking ownership of the entire development process.
Portable coding guidelines focused on explicit assumptions, simple implementations, narrow changes and verifiable goals.
OpenCode is an MIT-licensed TypeScript coding agent with terminal, desktop, web, and IDE clients. Its provider-agnostic client/server architecture exposes sessions and agent operations through an OpenAPI-based HTTP service.
Web search, scraping and crawling tools that return Markdown or structured data for AI applications.
A local model runner and API server with CLI chat, model management and Python and JavaScript clients; it supplies inference to many chat frontends.
Anthropic’s example skill collection, including reusable technical and business workflows. Open-source examples and source-available document skills have different licensing scopes.
Hugging Face model-definition library for pretrained text, vision, audio and multimodal models, used across training frameworks and inference engines.
A self-hosted web chat and agent platform for Ollama and OpenAI-compatible providers, with multi-user controls and extensions; its current license limits rebranding.
LangChain is an MIT-licensed Python framework for building agents and LLM applications from interoperable models, tools, retrieval components and integrations, with LangGraph providing lower-level orchestration.
A coding-agent skill that encourages simpler solutions, reuse of existing capabilities and less unnecessary code.
Claude Code is Anthropic's proprietary coding agent for terminal, IDE, desktop, browser, and automation workflows. It reads and edits repositories, runs commands, works with Git, and supports MCP, project instructions, hooks, skills, parallel agents, and permission-controlled tools; its public repository is not the complete runtime source.
A community-maintained collection of reported system prompts and tool definitions from AI products.
GitHub Spec Kit is an MIT-licensed Python toolkit for Spec-Driven Development that gives coding agents a structured Specify → Plan → Tasks → Implement → Converge workflow and stores each stage as durable Markdown artifacts.
A desktop manager for coding-agent provider settings, MCP servers, skills and prompts.
A node-based local AI media engine for building, saving and executing image, video, audio and other generative workflows.
A suite of coding-agent skills and tools for product planning, engineering review, browser QA and release work.
A design skill and supporting search tools for selecting UI styles, design systems and implementation guidance.
A C/C++ language- and vision-model inference project with quantization, multiple hardware backends, command-line tools and an OpenAI-compatible serving interface.
Codex CLI is OpenAI's Apache-2.0, Rust-based coding agent for inspecting, editing, running, and reviewing local code. It combines configurable sandboxing and approvals with interactive sessions, non-interactive automation, subagents, MCP connections, and optional handoff to hosted Codex environments.
Builds a queryable knowledge graph connecting source code with project documentation and other files.
A concise-response skill plus local proxy and middleware for reducing agent output and input tokens.
Gemini CLI is Google's Apache-2.0 terminal agent for codebase exploration, editing, shell and web operations, and scripted automation. It supports multimodal prompts, Google Search grounding, MCP extensions, project context files, checkpoints, and sandbox controls.
Pi Coding Agent is an MIT-licensed TypeScript terminal harness with interactive, print/JSON, RPC, and Node.js SDK interfaces. Its small core provides read, edit, write, Bash, models, and tree-structured sessions, while extensions, skills, prompts, themes, and packages supply workflow-specific behavior.
A tensor computation and deep-learning framework with automatic differentiation, neural-network components and accelerator support.
A composable skill collection for production engineering workflows, spanning definition, planning, implementation, verification, review and shipping.
A community-maintained directory of Model Context Protocol servers and related resources, useful for discovering integrations rather than executing them itself.
Persistent agent memory that captures tool activity, summarizes observations and retrieves relevant context in later sessions.
A multilingual OCR and document AI toolkit that extracts text, layouts, tables and structured Markdown or JSON from images and PDFs.
Portable design skills for improving the layout, typography, spacing and motion of AI-built interfaces.
OpenHands Agent Canvas is an MIT-licensed TypeScript browser client and self-hosted control center for coding-agent conversations and automations. It connects to local, Docker, VM, cloud, or enterprise Agent Server backends and can drive OpenHands or ACP-compatible agents; the separate Python SDK supplies execution.
A document parsing toolkit that converts PDFs and office documents into structured Markdown or JSON with layout-aware extraction.
A Rust command-output proxy that filters and summarizes verbose developer-tool output before it enters a coding agent’s context.
Progressive Python lessons for building a coding-agent harness, from a tool loop to context management and coordinated agents.
A local desktop LLM application for private chat and document conversations, with Python bindings and a local API option.
A local model training and inference toolkit with a desktop interface, optimized fine-tuning, dataset preparation and model export workflows.
A curated directory of skills, plugins and examples for Claude and other compatible coding assistants.
A unified fine-tuning toolkit for language and vision-language models, offering command-line workflows and the Gradio-based LLaMA Board interface.
Microsoft’s educational repository for learning agent concepts through lessons, examples and notebooks rather than a production agent framework.
Local context compression for tool outputs, logs, repository files and retrieval results.
The discontinued public Daytona repository describes sandbox infrastructure for AI-generated code. Its README says core development moved to a private codebase in June 2026.
A local semantic code graph and context service for coding agents, exposing repository structure through CLI, MCP, library and browser interfaces.
MetaGPT is an MIT-licensed Python multi-agent software-engineering framework that models a software company as specialized roles passing structured artifacts through a development workflow.
OpenSpec is an MIT-licensed TypeScript SDD framework for AI coding assistants, optimized for existing codebases through persistent behavioral specs, change-local deltas and an artifact-guided proposal → specs → design → tasks → implementation workflow.
Open Interpreter is an Apache-2.0 Rust coding agent forked from Codex and optimized for low-cost models. It can emulate multiple agent harnesses, run with native sandboxing, expose ACP, and remain compatible with the Codex exec protocol.
Cline is an Apache-2.0 TypeScript coding agent available through IDE extensions, a CLI, desktop applications, and an SDK. Its workflow emphasizes reviewable diffs, command and edit approvals, plan-versus-act separation, checkpoints, broad model-provider choice, and extensibility through plugins and MCP.
A model gateway that presents compatible API surfaces, routing and fallbacks across configured providers, with usage and quota visibility.
A Python document-processing toolkit that converts varied source formats into structured data for extraction, search, retrieval and agent context.
Community documentation and examples for Claude Code features, development workflows and agent orchestration.
A local-first chat and agent workspace for conversations over documents, with desktop and self-hosted editions and a developer API.
Get Shit Done is the archived MIT-licensed predecessor of GSD Core: a meta-prompting, context-engineering and spec-driven development system for Claude Code, OpenCode, Gemini and Codex. Active development moved to open-gsd/gsd-core.
A high-level deep-learning framework with interchangeable JAX, TensorFlow and PyTorch backends, plus an inference-only OpenVINO backend.
An MCP server and CLI for retrieving current library documentation and code examples for coding assistants.
A Python and CLI toolkit for training, validating, exporting and running YOLO models for detection and other vision tasks.
Microsoft’s multi-agent application framework, now in maintenance mode. Code is MIT licensed; documentation uses CC BY 4.0.
A Python framework that combines role-oriented agent crews with event-driven flows for coordinating multi-step AI applications.
goose is an Apache-2.0 Rust agent available through desktop, CLI, and API interfaces. It supports multiple model providers, MCP extensions, reusable workflow recipes, subagents, and ACP connections for editor integration or external agent providers.
A C/C++ Whisper inference engine for local speech transcription across desktop, mobile, embedded and browser environments.
BMad Method is an MIT-licensed AI-native software-development framework that installs role-specific agents and workflow skills into coding tools such as Claude Code and Cursor, carrying decisions from product discovery and requirements through architecture, stories, implementation and review.
LlamaIndex is an MIT-licensed Python framework for data-connected LLM applications, combining ingestion, indexing, retrieval, RAG, agents and workflows.
A cross-platform desktop chat client for cloud models and local backends such as Ollama and LM Studio, with assistants, documents and MCP support.
A self-hostable workspace for assigning issues to coding agents and reviewing their progress and results.
Reusable Python utilities for computer vision predictions, annotations, tracking, datasets and video processing across model providers.
Aider is an Apache-2.0 Python terminal pair programmer that edits existing repositories or starts new projects across hosted and local model providers. Its repository map supplies selective codebase context, while automatic Git commits, linting, tests, and undo-oriented workflows make changes reviewable.
A vector database for similarity search and retrieval workloads, commonly used as a data component in AI applications rather than an agent itself.
An offline-capable desktop AI chat app with bundled local model execution, optional cloud providers, MCP and an OpenAI-compatible local API.
Open-source AI code-review CLI that combines deterministic file selection and rules with an LLM agent for contextual, line-level findings.
A distributed deep-learning optimization library for training and inference, including ZeRO memory partitioning and offload techniques.
A terminal-native runtime for keeping many coding-agent sessions persistent, visible and controllable across local and SSH-connected machines.
LangGraph is an MIT-licensed Python framework for stateful, durable agent workflows with explicit graphs, persistence, human checkpoints and multi-agent control.
A distributed training framework with data, tensor, pipeline and sequence parallelism plus memory-management and optimization components.
CodeWhale is an MIT-licensed Rust coding agent with terminal and graphical clients over a shared runtime. It supports hosted or local models, plan/work modes, permission levels, and larger jobs delegated across agents with different roles and models.
A document converter that produces Markdown, JSON, HTML and chunks while preserving tables, equations, images and other document structure.
A multi-host collection of specialist agents, skills, commands and workflow plugins maintained in wshobson/agents.
A Claude Code orchestration plugin for specialist agents, coordinated teams and persistent task workflows.
Design-engineering and animation skills for reviewing and improving interfaces, including web and React Native motion guidance.
DSPy is an MIT-licensed Python framework for programming and automatically optimizing modular LM systems, including RAG pipelines and agent loops.
A context database for agents that organizes resources, memories and skills through a virtual filesystem with hierarchical summaries and retrieval.
Microsoft’s MCP server for giving agents structured browser automation through Playwright.
A local model gateway for managing providers, routing rules and requests from coding agents.
Anthropic’s official directory of Claude Code plugins, including first-party plugins and external integrations.
A Python numerical-computing library with composable automatic differentiation, compilation, vectorization and accelerator execution.
A client-neutral collection of reverse-engineering skills and routing playbooks that organizes reusable analysis workflows and tool selection.
Continue is an Apache-2.0 TypeScript coding agent for CLI, VS Code, and JetBrains, with shared YAML configuration for models, rules, MCP servers, and tools. It remains useful for historical comparison, but its maintainers issued a final 2.0.0 release and marked the repository read-only and no longer actively maintained.
Meta's modular PyTorch library for object detection, segmentation and visual recognition research with reusable models and deployment utilities.
A modular Hugging Face library for diffusion-model inference and training, with pretrained pipelines, schedulers and reusable model components.
An OpenAI-maintained Claude Code plugin for Codex reviews, task delegation and session handoff.
Cursor is a proprietary coding agent spanning its editor, terminal CLI, and ACP integrations. It offers agent, plan, and read-only ask modes, explicit tool permissions, resumable sessions, and local-to-cloud handoff without publishing its core implementation.
OpenMMLab's modular detection toolbox with model implementations, training configurations and evaluation workflows for visual recognition.
GitHub’s MCP server exposes repository, issue, pull-request and workflow operations to compatible AI clients through authenticated tools.
A framework that organizes PyTorch model code and automates training infrastructure such as mixed precision, checkpointing and distributed execution.
A fork of Pi that integrates coding tools, language-server and debugger operations into a terminal agent, with an extensible multi-provider environment.
A ready-to-use multilingual OCR library that returns detected text, bounding boxes and confidence values from images.
Semantic code retrieval and editing tools exposed to agents through MCP.
OpenAI Agents SDK is an MIT-licensed Python framework with lightweight primitives for agents, tools, handoffs, guardrails, sessions and tracing.
Persistent memory for coding agents with hybrid retrieval, knowledge relationships and cross-session context.
Semantic Kernel is Microsoft's MIT-licensed .NET/Python SDK for model, plugin and agent orchestration; Microsoft now directs new agent development toward Microsoft Agent Framework.
Packs a repository into an AI-friendly document with configurable file selection and output formatting.
Firecrawl’s example application for generating React websites with AI, combining web content acquisition, model calls and an execution environment.
An open-source annotation application for text, images, audio, video and time series with customizable labeling interfaces and model integrations.
A generative image application with a web canvas, model management and reusable visual workflows for iterative creative work.
An open-source AI lifecycle platform for experiment tracking, model management, LLM tracing, evaluation and prompt workflows.
A PyTorch-based deep-learning library combining high-level training APIs with reusable lower-level components for customized workflows.
A kanban workspace for planning coding-agent tasks, executing work and reviewing changes.
A Claude Code status-line plugin showing context usage, active tools, agent activity and task progress.
A TypeScript framework for AI agents and workflows. Most repository code is Apache 2.0, while directories named ee contain separately licensed enterprise features.
Qwen Code is an Apache-2.0 TypeScript coding agent centered on terminal workflows while supporting multiple model protocols and local inference. It combines code actions with planning, approvals, sandboxing, MCP, skills, subagents, memory, and IDE integrations.
A native Ghostty-based macOS terminal that organizes many coding-agent sessions with workspaces, attention notifications, session restore and a scriptable control API.
Kilo is an MIT-licensed coding agent spanning VS Code, JetBrains, CLI, and cloud workflows. It emphasizes model choice, task-specific agent modes, MCP extensibility, and the ability to switch models during an active task.
A Dolt-backed, dependency-aware issue tracker designed to preserve structured task state and project memory across coding-agent sessions.
Persistent file-based planning that keeps task plans, findings and progress available across agent sessions.
Grok Build is an Apache-2.0 Rust coding agent with an interactive TUI, headless automation, and ACP integration. It edits files, runs commands, searches the web, and supports extensibility through MCP servers, skills, plugins, and hooks.
Haystack is an Apache-2.0 Python orchestration framework for modular LLM pipelines, RAG and agents with explicit retrieval, routing, memory and generation stages.
A Python Whisper implementation using CTranslate2 for memory-efficient transcription, quantized inference and batched processing.
The portable Agent Skills format and reference tooling: version-controlled SKILL.md folders that compatible agents discover and load progressively.
An open Markdown convention for repository instructions consumed by coding agents, with a reference website and examples rather than an agent runtime.
A transcription pipeline combining Whisper recognition, forced alignment and optional speaker diarization to produce word-level timestamps.
Roo Code is an archived Apache-2.0 TypeScript coding agent for VS Code. It remains useful as an architectural reference for mode-based agents, provider choice, tool approvals, MCP integration and browser-assisted workflows, but it is no longer an active upstream.
Anthropic’s role-oriented plugin collection bundles skills, commands and connectors for business workflows in Claude Cowork and compatible Claude Code setups.
A reinforcement-learning post-training framework for language models using modular training and generation backends with flexible GPU placement.
A workflow engine for defining repeatable AI coding processes in YAML and running them across projects.
A terminal application combining shell sessions, file previews, editing and an integrated AI assistant, with local and remote development workflows.
A Python library for loading, streaming and transforming multimodal datasets from local files and the Hugging Face Hub.
Hugging Face library for adapting pretrained models by training a small set of additional or selected parameters, including LoRA and other adapter methods.
Google ADK is an Apache-2.0, code-first agent framework for building, evaluating and deploying multi-agent systems, optimized for Gemini but designed to be model- and deployment-agnostic.
A document OCR and layout toolkit with reading-order and table recognition, combining open-source code with separately licensed models.
Prime Agent is an MIT-licensed TypeScript autonomous agent built around persistent execution rather than disposable chat sessions. Its daemon, goals and scheduled work make it a useful reference for long-lived coding-agent state.
SWE-agent is an MIT-licensed Python software-engineering agent from Princeton and collaborators that turns GitHub issues into attempted repository fixes. It is also a research platform for studying agent-computer interfaces and SWE-bench behavior.
Pydantic AI is an MIT-licensed Python agent framework built around typed models, validated structured data, dependency injection, tools and testable application code.
jcode is an MIT-licensed Rust coding agent focused on low runtime overhead and parallel agent work. Its swarm design includes worker coordination, messaging, conflict awareness and optional worktree isolation.
Devika is an MIT-licensed Python autonomous software-engineering agent with a web-oriented interface. It decomposes high-level development requests, researches context, plans work and writes project code.
Hugging Face trainers and command-line workflows for foundation-model post-training, including supervised fine-tuning, preference optimization and reinforcement learning.
A browser inference library that uses WebGPU to run supported language models locally, with an OpenAI-style application interface.
NVIDIA's PyTorch framework for developing, customizing and deploying speech recognition, speech synthesis and speech language models.
A Neovim assistant for conversational coding and proposed edits, integrating model providers and agent protocols within the editor.
NVIDIA’s transformer training reference implementation and composable Megatron Core library for large distributed GPU workloads.
A reusable skill collection focused on context engineering: selecting, organizing and evaluating the information and execution structures surrounding AI agents.
A prose-editing skill with reference material for identifying and removing repetitive AI-writing patterns, rather than a text-generation model or detector.
DeepCode is an MIT-licensed Python multi-agent software factory that turns research papers, specifications and higher-level requirements into implementation artifacts through staged analysis, planning, retrieval and generation.
A Go MCP server for defining and exposing database tools to agents and applications, formerly maintained under the genai-toolbox name.
A Python framework for streaming voice and multimodal agents that connects audio, video, AI providers and transport pipelines.
Git-oriented data and model versioning with reproducible pipelines, local experiment tracking and remote artifact storage.
A Python agent harness with tool execution, skills, memory and multi-agent coordination, plus the built-in ohmo personal-assistant application.
Plandex is an MIT-licensed Go coding agent built around explicit planning for large, multi-file tasks. It keeps plans and changes reviewable before application and supports broad codebase context.
An offline speech recognition toolkit with streaming transcription, configurable vocabularies and bindings for desktop, mobile and embedded applications.
An ONNX Runtime speech toolkit for offline recognition, synthesis, diarization, keyword spotting and other audio tasks on many device platforms.
An agentic IDE for running many CLI coding agents in isolated worktrees with integrated terminals, review, browser previews, remote hosts and automation.
A framework for real-time voice and multimodal agents with provider integrations, turn handling and LiveKit media transport.
Microsoft’s framework for building agents and coordinated workflows in Python and .NET, with provider integrations and explicit application orchestration.
Community-maintained AI pull-request reviewer with configurable prompts, self-hosted deployment and multiple Git and model providers.
MiMo Code is Xiaomi MiMo's MIT-licensed TypeScript coding agent. It provides terminal and editor-oriented agent workflows and serves as the reference harness developed alongside MiMo coding models.
A configuration-driven framework for language and multimodal model post-training, with full fine-tuning, adapters, preference optimization and distributed execution.
Trae Agent is ByteDance's MIT-licensed Python software-engineering agent and research harness for repository-level tasks. It is suitable for autonomous issue solving, benchmark experiments and custom agent workflows.
Freebuff is the Apache-2.0 TypeScript coding agent formerly published as Codebuff. It uses multiple specialized agents and repository context from a terminal workflow rather than relying on one monolithic prompt.
A PyTorch toolkit for speech recognition, synthesis, enhancement, separation and other conversational audio research and applications.
A data-quality library that uses model predictions and dataset statistics to detect likely label errors, outliers, duplicates and other issues.
A differentiable computer vision library for PyTorch covering geometry, image transformations, augmentation and learned vision components.
Kimi CLI is Moonshot AI's Apache-2.0 Python terminal agent with direct shell mode, autonomous planning, MCP tools, and ACP-based editor integration. Its official repository is being gradually wound down in favor of Kimi Code CLI.
The Weights & Biases Python SDK for tracking training runs, hyperparameters, metrics and artifacts in the W&B model-development platform.
GitHub Copilot CLI brings Copilot's coding agent to interactive terminals, scripts, and ACP-compatible clients, with planning, tool approvals, MCP extensibility, and GitHub operations. Its public repository uses a restrictive proprietary license and does not make the broader Copilot service open source.
A Python toolkit and visual application for exploring multimodal datasets, finding data-quality issues and evaluating computer-vision models.
A command-line tool for installing and exposing reusable skills to coding agents, including synchronization of available-skill metadata into AGENTS.md.
Open SWE is LangChain's MIT-licensed Python asynchronous coding agent for turning software tasks into isolated implementation runs and pull requests. It is built for internal, event-driven engineering workflows rather than only interactive terminal sessions.
A Python speaker diarization toolkit with neural segmentation, speaker embeddings, pretrained pipelines and fine-tuning support.
A Python speech-to-text library with voice activity detection, wake words, incremental updates and optional streaming-server components.
A distributed reinforcement-learning framework for LLM post-training, with Ray orchestration, inference-engine integration and extensible agent interactions.
An end-to-end speech processing toolkit with reproducible recipes for recognition, synthesis, translation, enhancement and speaker tasks.
OpenMMLab's semantic segmentation toolbox with model implementations, configurable training recipes and benchmark evaluation.
A Hugging Face library and launcher that handles device placement, distributed execution and mixed precision around existing PyTorch training loops.
GSD Core is an MIT-licensed JavaScript context-engineering and Spec-Driven Development framework that drives existing coding agents through a repeatable Discuss → Plan → Execute → Verify → Ship loop while preserving durable project state between sessions.
A Go terminal application for supervising multiple coding-agent sessions in isolated Git worktrees with diff review and explicit checkout or push controls.
PyTorch quantization components for lower-memory inference and training, including low-bit linear layers and 8-bit optimizers.
mini-SWE-agent is an MIT-licensed Python minimal software-engineering agent from the SWE-agent project. It intentionally strips the architecture down to a small, readable issue-solving loop.
Kimi Code is Moonshot AI's MIT-licensed TypeScript successor to Kimi CLI. It is the current coding-agent harness for Kimi models and is designed for terminal use and editor integration.
A configurable web application and server for AI image and video generation, captioning, upscaling and media processing.
A neural-network library for JAX with flexible model construction, state management and transformation-compatible training workflows.
Swarms is an Apache-2.0 Python framework for composing and running multi-agent systems with prebuilt sequential, concurrent and hierarchical orchestration patterns.
A configurable data-processing framework for cleaning, synthesizing and analyzing text and multimodal datasets for foundation models.
A Python SDK for ClearML experiment tracking, dataset management and reproducible execution within its ML operations ecosystem.
An open, structured database of AI model specifications, capabilities and pricing metadata, maintained as a reference rather than an inference service.
Gentle-AI is an MIT-licensed Go configuration framework that adds persistent memory, Spec-Driven Development, skills, MCP servers, personas and bounded review to existing coding agents without locking a project to one agent.
A self-hosted experiment tracker with a Python SDK and interface for querying, visualizing and comparing training runs and AI metadata.
Ouroboros is an MIT-licensed Python Agent OS for replayable AI coding workflows that crystallizes a vague request into a locked Seed specification, records execution in a ledger, evaluates results through staged gates and can evolve subsequent attempts without changing the approved contract.
An open-source desktop ADE for running CLI coding agents in isolated worktrees, scheduling repeat work, reviewing diffs and managing agent resources.
A PyTorch-native LLM post-training library with editable recipes for fine-tuning and reinforcement learning; the project states it is no longer actively maintained.
A local neural text-to-speech engine maintained in the Open Home Foundation's GPL repository, with CLI, Python and native integration.
An interface-design skill that guides consistent application UI work and preserves design decisions for reuse across agent sessions.
Kode CLI is an Apache-2.0 TypeScript coding-agent runtime and SDK from shareAI-lab. It exposes terminal workflows together with programmable agent primitives suitable for subagents and custom orchestration.
A collaborative data-curation application for collecting human feedback, reviewing examples and improving datasets used by AI models.
Mistral Vibe is an open-source Python coding agent for interactive and programmatic terminal workflows. It combines project-aware context, configurable tool approvals, specialized agent profiles, subagent delegation, and MCP extensibility.
Kiro is a proprietary agentic development platform for IDE, CLI, web, and mobile workflows, centered on executable specifications and parallel task execution. It turns requirements into design and task artifacts, integrates hooks, steering, skills, MCP, permissions, and checkpoints, and exposes ACP from its CLI.
A Python streaming text-to-speech library that converts strings, generators and LLM output into audio through multiple synthesis engines.
fast-agent is an Apache-2.0 Python agent framework that also works as a terminal coding agent. It combines broad model-provider support with MCP, ACP, A2A, skills, plugins, evaluation tooling, and declarative multi-agent workflows.
cc-sdd is an MIT-licensed TypeScript SDD harness that installs the same discovery, requirements, design, task and autonomous implementation skills across multiple coding agents, with human approval gates, fresh implementers and independent per-task review.
A lightweight MCP database gateway that exposes schema inspection and SQL tools for supported relational databases to AI clients.
A framework for synthetic-data generation and AI feedback pipelines, with reusable generation and judging components across model providers.
A pipeline library for processing, filtering and deduplicating large text corpora used in model training.
A TypeScript ACP adapter that makes Anthropic's Claude Agent SDK available to compatible clients, including permission-aware tools, edit review, MCP, terminals, goals, and subagent transcripts. The adapter is Apache-2.0; Claude's underlying SDK and runtime remain commercially governed.
A JAX optimization library with composable gradient transformations, optimizer implementations and loss functions.
The maintained Idiap Coqui TTS fork provides speech synthesis, model training, fine-tuning and dataset utilities.
Google Antigravity is a proprietary agentic development platform with IDE and terminal interfaces. Its CLI shares the platform's agent engine and permissions while emphasizing keyboard-driven operation, remote sessions, parallel subagents, and reviewable artifacts.
A practical guide to authoring, evaluating and maintaining agent skills, with emphasis on discoverability, progressive disclosure and reliable supporting scripts.
Kimchi is an Apache-2.0 terminal coding agent built on the pi-mono SDK and connected to Kimchi's model infrastructure. Its central workflow routes planning, implementation, exploration, research, and independent review among role-specific models.
Amazon Q Developer CLI is the dual-licensed MIT OR Apache-2.0 Rust predecessor of Kiro CLI. Its source remains useful for studying a production Rust coding-agent runtime, but AWS now limits maintenance to critical security fixes.
AI code review for GitHub and GitLab, with pull-request summaries, suggested fixes, editor integrations and repository security scanning.
Stakpak is an Apache-2.0 Rust agent specialized for DevOps and infrastructure operations rather than general application coding. It combines interactive and scheduled execution with infrastructure context, credential protection, network guardrails, MCP, ACP, and provider flexibility.
A third-party KV-cache quantization implementation for language-model inference, with Python and GPU-kernel integrations rather than a standalone agent.
An OpenCode plugin for persistent project memory and session context, using embedded Turso/libSQL vector search and a management interface.
A GNOME desktop client for chatting with and managing models through Ollama, with optional cloud-provider connections.
Spec Kitty is an MIT-licensed Python SDD CLI and workflow system for governed multi-agent software delivery. It keeps specs, plans, work packages, review state and merge decisions in Git, isolates parallel agents with worktrees, and exposes progress through an optional kanban dashboard.
A desktop AI chat app with bundled local inference and an OpenAI-compatible server for agents and developer tools. It began as a fork of Jan and also connects to cloud providers.
DeepAgents is LangChain's MIT-licensed TypeScript agent harness, not a standalone proprietary coding product. It provides a LangGraph runtime with pluggable filesystems, context management, subagents, tools, streaming, and optional planning, sandbox, memory, and human-approval capabilities.
Dirac is an Apache-2.0 coding agent focused on precise, context-efficient changes through hash-anchored edits, AST operations, batched tools, and concurrent subagents. It runs in VS Code, terminals, and ACP clients with user-selected providers.
AI code-review platform with self-hosted and cloud editions, bring-your-own-model support and repository-specific review rules.
A GPU-oriented language-model inference engine and EXL3 quantization toolkit for running supported models on local hardware.
A Python, OpenAI-compatible local inference server built around ExLlamaV3. The maintainers describe it as a hobby project rather than a production service.
A self-contained TUI for running and restoring multiple coding-agent sessions across Git worktrees and projects without depending on tmux.
A local-first coding agent with project planning, verification, persistent work state and worktree-aware Git automation, distinct from the separate GSD Core repository.
Shopify’s Ruby toolkit for composing structured AI workflows from reusable steps, including model conversations, agent calls, code and command execution.
MoAI-ADK is an Apache-2.0 Go harness around Claude Code that combines a SPEC lifecycle with verification-backed completion claims, plan/run/sync execution, isolated worktrees, quality gates and optional multi-session Kanban or factory orchestration.
A deployment and service-management layer for AI agents, providing runtime integration, communication and supporting application infrastructure.
OpenHands Software Agent SDK is the MIT-licensed Python runtime/framework beneath the newer OpenHands architecture. It supplies agents, tools, conversations, workspaces and execution backends independently from the separately cataloged Agent Canvas UI.
The source repository for web-platform guidance skills, including authoring, calibration and evaluation tools plus a CLI for targeted guidance retrieval.
A self-hostable, multi-provider agent system that combines assistant interfaces, tools and multi-agent workflows rather than focusing only on repository editing.
A Go TUI for managing coding-agent sessions across projects with status detection, groups, worktrees, session forking, search and a remote conductor.
VT Code is a Rust coding harness for interactive, headless, and long-running terminal work. It emphasizes sandboxed execution, durable and replayable sessions, explicit approvals, verification, multi-provider support, and protocol-based extension.
Shotgun is an MIT-licensed Python SDD tool that indexes an existing codebase, researches the problem, builds a code-aware specification and staged implementation plan, then exports focused instructions for coding agents rather than asking them to infer architecture from a single prompt.
pi ACP is an MIT-licensed TypeScript adapter that exposes the separate Pi Coding Agent to ACP clients. It translates ACP JSON-RPC messages over stdio to Pi's RPC mode rather than implementing an independent agent runtime.
A TypeScript collection of shadcn/ui-based components and blocks for building interfaces. It is a general UI resource, not an AI model or coding agent.
Poolside's `pool` CLI is a proprietary coding agent and ACP client/server for interactive terminal work, editor integration, and automation. It combines planning, permissions, sandboxes, subagents, hooks, skills, MCP resources, and configurable model backends.
An MCP server exposing local image-processing operations such as cropping, resizing, detection, OCR and background removal to compatible agents.
LeanSpec is an MIT-licensed, tool-agnostic SDD framework that gives coding agents a unified spec model, CLI, MCP interface and visual management surface while keeping specifications small, living and independent from a particular issue tracker or AI agent.
Auggie CLI is Augment's proprietary terminal agent, built around repository-wide context for understanding and changing existing codebases. It supports interactive development and non-interactive automation, but its public repository distributes the client and examples rather than an openly licensed agent core.
An OpenCode plugin that reports subscription and quota status from supported provider accounts, without acting as a coding agent or model gateway.
A local-first software-development orchestrator that runs parallel coding agents in isolated worktrees and drives features through specs, CI fixes and draft PRs.
A tmux-native agent dashboard that discovers existing coding sessions, tracks live states, manages worktrees and supports direct handoff between agents.
Autohand Code is a TypeScript terminal agent with planning, editing, testing, memory, skills, subagents, provider selection, and Git automation. Its repository presents Apache-2.0 licensing together with an additional commercial-license requirement for organizations above a stated revenue threshold.
A chat-driven meta-agent that orchestrates multiple CLI coding agents through tmux with routing, persistent memory, auto-continuation and independent review loops.
Amp is a proprietary multi-model coding agent spanning terminal, web, and managed execution environments. This catalog entry pairs that product with a community Apache-2.0 ACP adapter that invokes an installed Amp CLI and translates its sessions, tools, permissions, images, and MCP configuration.
SpecKit Companion is an MIT-licensed TypeScript VS Code workspace for managing an SDD lifecycle: visual specs, inline review comments, live pipeline state, living capability specs, drift detection and configurable Spec Kit-compatible workflows over plain repository files.
Commercial AI code-review and governance platform that applies organizational rules and cross-repository context throughout developer and coding-agent workflows.
A repository-based architecture workflow for AI assistants, using documented decisions, multi-perspective reviews, implementation guidance and recalibration rather than an autonomous runtime.
Okto Pulse is a local-first Python SDD workbench that exposes a governed Stories → Ideation → Refinement → Spec → Sprint → Tasks/Tests/Bugs lifecycle to coding agents over MCP while humans inspect and steer the same artifacts in a web UI.
Codebase-aware AI reviews in GitHub, GitLab, Bitbucket, IDEs and the command line, with custom guidelines and suggested fixes.
Go client libraries for loading and calling tools exposed by MCP Toolbox, including integrations for supported Go agent frameworks.
crow-cli is a Python, ACP-native coding agent centered on durable, searchable session state. It combines terminal and editor use with SQL-backed shared memory, MCP tools, resumable conversations, and cross-agent delegation.
GLM Agent is a community-built Apache-2.0 TypeScript agent for running Z.AI GLM Coding Plan models through ACP-compatible clients. It provides persistent sessions, configurable permissions, model switching, and local filesystem and command tools.
siGit Code is an Apache-2.0 Rust coding agent designed primarily for on-device inference. It offers an interactive terminal UI and an ACP stdio server for editors, with a distinct optional hosted service.
Minion Code is an AGPL-3.0 Python coding assistant built on the separate Minion framework. It can run as an interactive CLI, an embeddable agent, or an ACP server with configurable models, permissions, and MCP tools.
Agoragentic is a governance, deployment, discovery, and transaction layer for autonomous agents rather than a conventional coding agent. Its MIT-licensed integration repository includes protocol adapters and catalogs, while the hosted Triptych OS control plane and live execution paths are separate.
Corust Agent is an ACP-packaged coding companion positioned specifically for Rust development. Public materials verify cross-platform binary releases and an ACP server, but do not expose or document the implementation architecture.
A local-first browser control room for supervising parallel Claude Code and Codex sessions with unified approvals, diffs, replay, memory and history.
DimCode is a proprietary, multi-provider terminal coding agent with interactive and headless workflows, local session persistence, MCP integration, plan and agent modes, and tool approvals. Its public GitHub repository is only an issue tracker, not implementation source.
A local-first Rust personal agent using a Markdown vault as memory and grammar-constrained tool calls through llama.cpp, with optional external integrations.
Harn is a pre-1.0 programming language and Rust runtime for explicitly defined agent pipelines, rather than a conventional chat-based coding assistant. It provides typed tools, durable execution, replay, capability checks, and native agent protocols.
Factory Droid is a proprietary development agent spanning terminal, IDE, web, desktop, and collaboration workflows. Its persistent sessions, headless execution, extensible harness, custom subagents, and orchestrated Missions connect interactive coding with longer structured projects.
A Go daemon, CLI and local web dashboard for running multiple coding agents in managed Git workspaces backed by tmux.
Nova is Compass Platform's proprietary terminal coding agent, with ACP editor integration and optional session synchronization across CLI, desktop, and web clients. Its public MIT repository documents and distributes the product but does not contain the core implementation.
AI code reviewer that indexes repository structure and dependencies to analyze pull requests with broader codebase context.
Devin CLI is Cognition's local terminal coding agent with resumable sessions, permission modes, sandboxed autonomous execution, and handoff to cloud Devin sessions. The CLI and proprietary cloud agent are separate components with different environments and feature sets.
Qoder CLI is a proprietary terminal coding agent oriented toward automation, code review, testing, and issue remediation. TypeScript and Python SDKs expose its agent through scripts and services, but the linked public repository is only a release and changelog surface.
An Oh My Pi extension that coordinates scope, research, planning, implementation and review while preserving versioned workflow state and explicit approvals.
Junie is JetBrains' task-oriented coding agent for terminal, JetBrains IDE, and CI/CD workflows. It can implement changes, analyze repositories, review code, and fix tests while allowing users to guide and approve its work or select an external model provider.
A Python CLI and MCP server for discovering, spawning and controlling coding agents in local or SSH-hosted tmux sessions.
Scala 3 toolkit combining typed LLM tools and provider clients with state-machine workflows built on Cats Effect, fs2 and http4s.
AgentOps records agent execution, model usage, costs, and failures so developers can inspect sessions and monitor applications across supported agent frameworks.
AgentScope builds tool-using agents and coordinated teams with reusable model abstractions, pipelines, runtime controls, and service deployment support.
A Python framework and runtime for building agents, coordinating teams and serving agent applications with persistent state.
AIBrix supplies cloud-native infrastructure components for deploying, routing, managing, and scaling language-model inference services, particularly around vLLM deployments.
aimock provides deterministic mock services for testing AI applications against model APIs, MCP, agent protocols, vector databases, and search.
AWS’s managed agent orchestration service for model-driven actions and knowledge retrieval; the classic service is closed to new customers.
A managed AWS service for screening AI inputs and responses using configurable content, topic, word, sensitive-information and grounding policies, including applications using models outside Bedrock.
A search and inference database combining full-text, vector and graph retrieval with document enrichment and RAG. Its Zig server exposes APIs and SDKs, while the core uses Elastic License 2.0 and selected clients and supporting components use Apache-2.0.
A Python-based platform for authoring, scheduling and monitoring batch workflows and data pipelines.
Managed execution and marketplace for web-data extraction and automation tools used by AI agents and data pipelines.
A self-hostable platform for tracing and evaluating AI applications using OpenTelemetry and OpenInference. Current Phoenix server code uses Elastic License 2.0, which restricts offering it as a competing hosted service.
assistant-ui provides composable TypeScript and React components for AI chat, streaming messages, tool interfaces, and human approvals.
A managed Azure service for detecting harmful text and images, prompt attacks and selected grounding or task-adherence problems through APIs and a web studio.
BeeAI Framework provides Python and TypeScript building blocks for tool-using agents and multi-agent workflows with configurable execution behavior.
BentoML turns Python inference code into deployable APIs and task queues, with model composition, dependency packaging, batching, and container-based deployment.
Bifrost is a Go AI gateway that standardizes provider access and supplies routing, failover, load balancing, caching, and a configuration interface.
BigCode Evaluation Harness generates and evaluates code-model outputs across programming benchmarks, with multi-GPU generation and container-based execution workflows.
Browser-based AI application builder for prompting, editing and publishing websites and applications.
Hosted platform for building AI agents with visual workflows, knowledge bases and integrations.
Braintrust JavaScript SDK adds tracing, logging, evaluation runs, and integration helpers to TypeScript and JavaScript AI applications.
AI-assisted web extraction and monitoring platform for creating structured datasets and API-accessible website workflows.
Browser Use lets language-model agents inspect and operate websites through a Python library and command-line workflows using local or remote browsers.
Managed browser infrastructure for AI agents and automated web workflows, with session inspection and developer APIs.
Burn is a Rust tensor and deep-learning framework for training and inference, with automatic differentiation and interchangeable CPU, GPU, and WebAssembly backends.
A Python framework for multi-agent systems, role-based collaboration and research into agent behavior at scale.
CLI that gives coding agents JVM library signatures, symbols, source and dependency trees as Markdown, using Maven coordinates or project classpaths.
Chroma stores embeddings, documents and metadata for AI retrieval, with a compact collection API and local or client-server usage.
Cloudflare Agents is a TypeScript SDK for persistent agents with state, scheduling, realtime communication, tools, and workflows on Durable Objects.
CodeBuddy Code is Tencent's proprietary terminal coding agent with repository context, interactive and headless workflows, native ACP and MCP integration, layered project memory, and configurable subagents. It is distributed through Node.js tooling, but its implementation source and language are not publicly established.
AI coding assistant for VS Code, JetBrains and Visual Studio, with bring-your-own-key and managed-plan options.
AI-assisted code review for pull requests and developer workflows, with contextual feedback and interactive review discussions.
Cognee turns documents, code, and conversations into persistent graph-based memory that AI agents can search and reuse across sessions.
Snowflake Cortex Code, now documented as Snowflake CoCo, is a data-focused agent for Snowflake engineering and analytics. It combines local code and shell work with schema-aware SQL execution, RBAC, and Snowflake-native workflows.
Cosh is a terminal coding agent with reusable Rust libraries, local context recall, and computer tools that operate graphical interfaces through accessibility information.
DeepEval brings repeatable tests and configurable evaluation metrics to LLM applications, retrieval pipelines, and agent trajectories through a Python testing interface.
DeepTeam generates adversarial tests for AI applications and agents, with attack strategies, vulnerability checks, and runtime guardrails in a Python framework.
E2B provides SDKs for starting isolated execution environments where AI agents can run commands, execute generated code, and interact with tools.
Eino is a Go framework for language-model applications, with reusable model and retrieval components, graph composition, and an agent development kit.
Elasticsearch combines full-text search, structured filtering and vector retrieval in a distributed search engine used for AI knowledge retrieval.
Developer platform for speech generation, transcription and conversational voice agents, with APIs and hosted tooling.
Embedchain is an archived Python RAG framework that packages data ingestion, chunking, embedding and retrieval behind a compact application API.
EvalPlus evaluates generated code with expanded test suites, model-generation support, and tooling for studying correctness and execution efficiency.
Evidently evaluates and monitors ML and LLM systems with configurable reports, test suites, data-quality checks, and a self-hostable monitoring interface.
Charm’s Go library for building tool-using AI agents through a common model-provider API. It powers Crush and supports dedicated provider adapters alongside a generic OpenAI-compatible layer for integrating additional model services.
FastMCP is a Python framework for building Model Context Protocol servers, clients, and applications. It turns Python functions into tools and handles schemas, validation, transports, and protocol lifecycle around application code.
FastMCP is a TypeScript framework for defining MCP tools, resources, prompts, and server transports with less setup code.
The official filesystem reference server exposes configurable file operations through Model Context Protocol. It provides a concrete Node.js implementation with allowed-directory controls for clients that need access to local files.
A Python retrieval toolkit for embedding and reranking inference, model fine-tuning, training-data preparation and retrieval evaluation.
garak probes language models for failure modes such as prompt injection, data leakage, hallucination, and unsafe generation using extensible test plugins.
Google’s open-source framework for building AI features with model adapters, tool calling, structured outputs and retrieval. It combines language SDKs with a local developer UI, tracing and evaluation tools; Go and JavaScript/TypeScript are established SDKs, with other languages at differing stages.
Giskard supplies modular evaluation, scenario testing, vulnerability scanning, and test generation for AI agents and retrieval applications.
Graphiti builds temporal knowledge graphs for AI agents, retaining relationships, source provenance, and changes to facts as new information arrives.
Ground Station captures coding-agent telemetry locally and presents model calls, tool calls, token usage, and execution trajectories through a CLI and browser interface.
A Python framework for validating LLM inputs and outputs, composing validators from Guardrails Hub, and producing structured responses with configured handling for validation failures.
Guidance combines Python control flow with constrained language-model generation, allowing reusable functions to shape outputs using grammars and regular expressions.
An LLM observability platform and AI gateway for logging requests, tracing agents and inspecting usage, cost and latency, with routing and fallback capabilities across model providers.
Hugging Face Hub is a platform for discovering, publishing, and collaborating on models, datasets, and machine-learning applications. It provides model cards, repository versioning, and hosted demos rather than being a single language model.
Inspect AI provides composable tasks, model integrations, tools, scorers, and result inspection for evaluating language models and agent behavior.
Instructor extracts validated structured data from language models using Pydantic schemas, provider adapters, automatic retries, and support for streamed objects.
A Go library for extracting structured data from language models into typed Go values. It derives schemas from structures and combines provider adapters, validation and retries to reduce hand-written response parsing.
JVM LLM inference engine with quantized models, embeddings, tool calling, an OpenAI-compatible server and distributed inference.
Java agent harness using ordinary classes and annotated methods, with structured output, tool loops, memory, MCP, tracing and evaluation.
json-render turns model-generated JSON specifications into interfaces assembled from a developer-defined component catalogue.
An event-driven orchestration platform using declarative workflows to coordinate data, AI and infrastructure tasks.
KServe provides Kubernetes resources and controllers for deploying and operating generative and predictive model inference across supported serving frameworks.
A commercial runtime screening service for prompt attacks, sensitive-data leakage, harmful content and unsafe agent interactions. Formerly branded Lakera Guard, its current documentation calls the runtime offering Check Point AI Guardrails.
LanceDB stores and searches vectors, metadata, and multimodal data using the Lance columnar format, with embedded SDKs and managed deployment options.
LangChain4j brings model-provider abstractions, tool use, retrieval pipelines, and agent building blocks to Java applications and enterprise JVM frameworks.
A Go implementation of LangChain concepts for building LLM applications from model adapters, chains, agents, tools and memory. Its integrations include hosted and local models, embedding providers and vector stores for retrieval-augmented generation.
A self-hostable LLM engineering platform connecting traces, prompt management, datasets, experiments and evaluation. Its core is MIT-licensed, while designated enterprise directories use separate commercial terms.
A commercial platform for tracing, debugging, evaluating and monitoring LLM applications and agents. Its MIT-licensed Python and TypeScript client SDKs work with LangChain and other application stacks.
LangWatch combines LLM tracing, agent simulation tests, evaluations, prompt management, and operational controls for applications and coding-agent usage.
Letta builds stateful agents with persistent memory and identity. Its current implementation is the Letta Code harness, with CLI, application and SDK interfaces.
Hugging Face Lighteval evaluates language models across multiple backends and retains sample-level results for debugging and comparing benchmark performance.
LightRAG builds retrieval-augmented applications by combining graph-based knowledge extraction with vector retrieval and document-management workflows.
LiteLLM provides a Python SDK and self-hosted AI gateway that normalize model-provider APIs and centralize routing, credentials, usage tracking, and spending controls.
LitServe builds custom Python inference services with application-defined loading and prediction logic, plus batching, streaming, concurrency, and deployment controls.
llama-swap starts, stops, and swaps local inference servers on demand behind a common API, helping several models share limited machine resources.
An event-driven Python workflow library, maintained within LlamaAgents, for orchestrating asynchronous AI application steps.
Java bindings to llama.cpp using the Foreign Function and Memory API, supporting local model loading, generation, embeddings and native acceleration.
EleutherAI LM Evaluation Harness runs reproducible benchmark tasks against language models through local inference backends or compatible hosted APIs.
LM Format Enforcer constrains token generation to match JSON schemas or regular expressions while integrating into supported model inference pipelines.
LocalAI exposes self-hosted model inference through compatible APIs, using separately loaded backends for language, vision, voice, image, and other supported workloads.
loqui provides local text-to-speech and speech recognition in Rust, usable in process, from a CLI, or through OpenAI-compatible audio endpoints.
AI application builder for creating and iterating on web applications through a conversational development interface.
A visual automation platform for connecting applications, routing data and incorporating AI agents into business workflows.
A Go implementation of the Model Context Protocol for connecting AI applications to tools, prompts and resources. It supplies client and server building blocks, transport support, session handling and hooks for application-specific behavior.
A Go command-line tool for deterministic testing of MCP servers over their actual transports. YAML assertions, schema linting, fuzzing and CI reports help check protocol behavior and tool responses without requiring an LLM judge.
Mem0 extracts and retrieves persistent user, session and agent memories for AI applications, with open-source libraries and a separate managed platform.
MemMachine supplies a persistent memory service for AI agents with episodic context, user profiles, and working memory exposed through client APIs.
Memobase derives persistent user profiles and event histories from conversations so applications can retrieve personalized context across sessions.
MemOS manages persistent memories for language-model applications and agents, with retrieval, editable knowledge, isolated memory collections, and integration plugins.
Microsandbox runs untrusted agent workloads in local microVMs, with a CLI and embeddable SDKs for sandbox lifecycle, images, snapshots, and execution.
A managed Azure platform for configuring, deploying and operating AI agents with integrated models, tools and enterprise controls.
Microsoft GraphRAG extracts knowledge graphs and summaries from text for graph-based retrieval and question answering. The research project is now in maintenance mode.
Microsoft’s automation platform for cloud workflows, desktop robotic process automation and AI-assisted business processes.
Visual AI-agent builder with model routing, data sources, custom functions and managed workflow execution.
mistral.rs runs supported language and multimodal models through a Rust inference engine, with command-line serving, Python and Rust SDKs, and compatible model APIs.
MLC LLM compiles language models for deployment across desktop, mobile, and browser hardware through a shared inference engine and application interfaces.
MongoDB Atlas Vector Search adds semantic retrieval to Atlas application data, combining embeddings with document metadata and aggregation workflows.
A visual workflow platform combining application integrations, custom code and AI agents, available for self-hosting or as a managed service.
Neo4j is a graph database for connected data, graph traversal and vector retrieval that can support knowledge-graph and RAG applications.
Nexus Agents coordinates coding agents with task routing, adversarial review, quality gates, and a tamper-evident audit trail.
A Python library for applying programmable input, output, retrieval, tool-execution and dialogue guardrails around LLM applications. Colang flows and configurable checks coordinate policy enforcement across model providers.
The official MCP Registry publishes structured metadata for discovering Model Context Protocol servers. It provides a shared registry API and publishing workflow that downstream directories and marketplaces can build upon.
The official MCP SDK catalogue links language-specific libraries for building Model Context Protocol clients and servers. It documents implementation tiers, protocol support, and entry points without treating all SDKs as one package.
ONNX Runtime executes machine-learning models across supported hardware and operating systems, with graph optimizations and hardware-specific execution providers.
OpenAI Evals is a repository-based framework and benchmark registry for testing language-model systems with reusable or private evaluation cases.
Asynchronous Scala clients for OpenAI and other LLM providers, including streaming, tool calls, embeddings, multimodal input and batch APIs.
OpenCompass provides a configurable model-evaluation platform covering many datasets, capabilities, and local or hosted language-model integrations.
Java agent SDK with ReAct agents, graph workflows, streaming execution, state persistence and resumable workflow runs.
OpenLIT uses OpenTelemetry to trace, evaluate, guard, and monitor AI applications, including coding agents, model costs, and supporting infrastructure.
OpenLLMetry adds OpenTelemetry instrumentation for model providers, vector databases, and AI frameworks, sending traces to compatible observability backends.
OpenSandbox offers unified APIs for running AI agents in isolated environments on local Docker or Kubernetes infrastructure, including code, browser, and desktop workloads.
OpenVINO optimizes and deploys models from supported training frameworks for inference on CPUs, Intel GPUs, and Intel NPUs across edge and server environments.
Opik combines application tracing, evaluation datasets, experiment comparison, prompt management, and production monitoring for LLM applications and agents.
ort exposes ONNX Runtime to Rust applications for hardware-accelerated model inference and training, with execution-provider controls and support for alternative runtime backends.
Outlines turns Python types and schemas into structured model outputs, offering a common interface across supported local and hosted model integrations.
PageIndex creates hierarchical document indexes and uses model reasoning to navigate them, offering a retrieval approach that does not require a vector database.
Pathway provides Python data pipelines backed by an incremental Rust engine, keeping transformations and retrieval-oriented datasets synchronized with changing source data.
pgvector adds vector similarity search to PostgreSQL, including exact search and approximate indexes alongside relational data and SQL.
Phren stores agent findings, tasks, and session context in Git-managed Markdown, with CLI, MCP, web, and editor interfaces.
Pinecone is a managed vector database for semantic retrieval, with metadata filtering and APIs for embedding-based search applications.
A hosted integration platform combining event triggers, reusable application actions and custom code for developer workflows.
Portkey AI Gateway routes requests across model providers with a common API, configurable fallbacks, load balancing, retries, caching, and guardrail integrations.
PostgreSQL provides transactional relational storage for AI application data and state; vector similarity search is available through extensions such as pgvector.
A Python workflow orchestration framework with scheduling, retries, caching and monitoring for data and application pipelines.
A toolkit for identifying, redacting and anonymizing sensitive data in text, images and structured records. Originally Microsoft Presidio, it is now maintained under the independent Data Privacy Stack organization.
Palo Alto Networks' security platform for AI models, applications and agents, combining discovery, posture assessment, model scanning, red teaming and runtime protection. It incorporates the acquired Protect AI capabilities.
A SentinelOne security offering for organizational AI use, focusing on visibility, sensitive-data leakage and runtime protection for AI applications and agents.
PromptBench is an archived Microsoft evaluation library for comparing language models, prompting methods, and robustness under prompt perturbations.
A CLI and library for evaluating prompts, models and AI applications and running automated red-team tests. Declarative test suites support local workflows, comparisons and continuous integration.
Pydantic Logfire instruments Python and AI applications with OpenTelemetry and provides a platform for exploring traces, logs, metrics, and evaluation evidence.
Microsoft PyRIT is a Python framework for identifying risks in generative AI systems and supporting repeatable security red-teaming work.
Qdrant is a vector database with metadata filtering, dense and sparse retrieval, and APIs for building semantic-search and RAG systems.
A Python evaluation toolkit for RAG and other LLM applications, with configurable metrics, test-data generation and experiment workflows. The current repository is maintained under Vibrant Labs.
RAGFlow combines document parsing, retrieval and agent workflows to turn heterogeneous files into searchable context with traceable citations.
Ray provides distributed tasks and stateful actors for AI workloads, with libraries for model serving, data processing, training, and tuning on clusters.
Redis is a data-structure server and query engine used for caching, session state and vector retrieval in AI applications.
Managed cloud workspaces for running coding agents against repositories with prepared dependencies, browser access and human review.
Managed AI inference platform for running model predictions through APIs without operating the underlying GPU servers.
Platform for building, deploying and monitoring conversational AI phone agents with call flows and external integrations.
Rig is a Rust framework for model-backed applications and agent workflows, with shared provider interfaces, streaming, tools, memory contracts, and vector-store integrations.
RMCP is the official Rust SDK for building Model Context Protocol clients and servers, with asynchronous transports, typed protocol structures, and tool-definition macros.
GPU cloud platform for deploying AI workloads on managed Pods and serverless inference endpoints.
Safe Docx lets coding agents read, search, edit, and compare Word and OpenDocument files through a local MCP server and CLI.
Experimental Scala tooling that gives coding agents compiler diagnostics, SemanticDB evidence, symbol queries and functional-programming analysis through a CLI and MCP.
Sentence Transformers is a Python library for running and training embedding and reranker models. It supports semantic search, similarity, retrieval, and model fine-tuning with a broad ecosystem of compatible checkpoints.
SGLang serves language and multimodal models with an inference engine designed for repeated prompts, agent workloads, and deployments from individual accelerators to distributed clusters.
Skyvern automates browser workflows using language models and computer vision, with a Playwright-compatible SDK and a visual workflow builder.
smolagents is a Python agent library whose agents can express tool actions as executable code, with model-provider adapters and integrations for external tools.
Spring AI adds model, tool, embedding, and vector-store abstractions to Spring applications using typed APIs and familiar Spring configuration patterns.
Stagehand provides browser-agent SDKs that combine familiar page operations with model-assisted actions, observation, and structured extraction.
Steel Browser provides a self-hostable browser API that manages sessions, Chrome processes, page extraction, and debugging for AI agents and automation tools.
An open-source agent SDK and harness supporting Python and TypeScript, model providers, tools and multi-agent coordination.
Scala toolkit for LLM clients and agents, with typed requests, tool calling, structured output, MCP and streaming integrations for several effect systems.
A durable execution platform for workflows that must survive failures, retries and long waits, including AI agent applications.
TensorRT-LLM provides NVIDIA GPU inference kernels, scheduling, and Python and native runtime components for serving large language and supported multimodal models.
Tessl is a proprietary context-and-skills platform for coding agents whose current catalog includes an explicit Spec-Driven Development workflow: agents gather requirements, write reviewable specs, wait for approval, implement against them and verify the resulting requirements.
Text Embeddings Inference serves supported embedding, reranking, and sequence-classification models with dynamic batching and HTTP or gRPC access.
TOON is a compact serialization format and TypeScript toolkit for passing JSON-shaped data to language models.
Triton Inference Server hosts models from multiple machine-learning frameworks behind inference APIs, with concurrent execution, batching, ensembles, and extensible backends.
An OpenTelemetry-native Python toolkit for instrumenting AI applications, scoring traces and comparing experiments, including feedback functions for retrieval quality, groundedness and answer relevance.
txtai combines an embeddings database, semantic search, model pipelines, and workflows in a Python framework with HTTP and MCP interfaces.
Unstructured partitions complex documents into structured elements that applications can clean, chunk and prepare for retrieval pipelines.
vectrize indexes Markdown folders for local hybrid search, combining embeddings and keyword retrieval with a background watcher and machine-readable results for agents.
Vercel AI SDK is a TypeScript toolkit for model calls, tool-using agents, structured responses, streaming, and AI application interfaces.
Vespa combines text, vector, tensor, and structured-data retrieval with configurable ranking and inference for search, recommendation, and retrieval-augmented applications.
vLLM is an open-source inference and serving engine for language and multimodal models. It combines efficient attention-memory management, continuous batching, and distributed execution with an OpenAI-compatible serving interface.
Weaviate combines vector retrieval, structured filtering and hybrid search in a database, with community, enterprise and hosted deployment options.
A toolkit and hosted workflow for tracing LLM applications, building repeatable evaluations and comparing application versions. The public toolkit is Apache-2.0; the Weights & Biases service has commercial terms.
XGrammar supplies optimized grammar-constrained decoding for structured generation, with APIs and integrations for model inference engines.
Xinference deploys and serves supported language, embedding, speech, and multimodal models through shared management tools and inference APIs.
A hosted automation platform connecting business applications, data and AI actions through configurable workflows.
Zep is a managed context and agent-memory service built around temporal knowledge graphs, with SDKs and integrations for AI applications.