Loading Tavory AI
Loading Tavory AI
Provider directory
Tavory currently lists 15 Fireworks models in its live catalog. Explore model capabilities, catalog signals and eligible workspace access. Tavory does not represent a manufacturer relationship unless explicitly stated.
text
DeepSeek V4 Flash 0731 is the official release of DeepSeek V4 Flash, superseding the preview version, with substantially enhanced agentic capabilities. A 304B parameter open-source Mixture-of-Experts model with speculative decoding, 1M token context, optimized for fast, cost-efficient inference with strong reasoning, coding, and function calling performance.
View modeltext
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
View modeltext
Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.
View modeltext
GPT-OSS-20B is OpenAI's lightweight open-weight model with 21.5 billion total parameters (3.6 billion active, Mixture-of-Experts). Optimized for lower latency and local or specialized use cases including agentic tasks, chain-of-thought reasoning, and web browsing. Released under Apache 2.0 license.
View modeltext
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. It features a 1028B Mixture-of-Experts architecture and supports vision, function calling, and agentic paradigms.
View modeltext
GPT-OSS-120B is OpenAI's open-weight large language model with 116 billion total parameters (5.1 billion active, Mixture-of-Experts). Designed for high-performance reasoning, agentic tasks, function calling, and general-purpose applications. Supports configurable reasoning depth and chain-of-thought. Released under Apache 2.0 license.
View modelvideo
MiniMax-M3 is a native multimodal model with 512K context running ~428B parameters and ~23B activated parameters. It brings native multimodality, enabling deeper semantic fusion across text, image, and video. M3 also introduces MiniMax Sparse Attention (MSA) to improve long context efficiency, achieving frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.
View modeltext
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.
View modeltext
GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths. A 743B-parameter model in Zhipu AI's GLM series, designed to plan, execute, and iterate autonomously on extended, engineering-grade tasks.
View modeltext
DeepSeek V4 Pro is a flagship open source Mixture of Experts model designed for frontier reasoning, advanced coding, and long context intelligence at scale (up to 1M tokens). It introduces a hybrid attention architecture that dramatically improves long context efficiency while reducing KV and compute overhead, along with stability and training enhancements for deep multi step reasoning. It represents a top tier open source system for complex agentic workflows, high precision reasoning, and demanding production workloads.
View modeltext
GLM 5.1 is Z.ai's next generation flagship model built for agentic engineering, with stronger coding capabilities and sustained performance over long horizon tasks with hundreds of iteration rounds. It's a 754B parameter MoE model.
View modeltext
DeepSeek V4 Flash is a streamlined open-source Mixture-of-Experts model (284B parameters) optimized for fast, cost-efficient inference while preserving strong reasoning and coding performance at 1M token context scale. It leverages hybrid attention innovations for lower latency and higher throughput, delivering near-Pro reasoning quality ideal for interactive agents and high-volume production workloads.
View modeltext
Mixture of Experts language model. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search.
View modeltext
GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths. A 743B-parameter model in Zhipu AI's GLM series, designed to plan, execute, and iterate autonomously on extended, engineering-grade tasks.
View modeltext
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its vision-language abilities while retaining full-stack, agent-level intelligence for coding, tool use, and productivity workflows. Its distinguishing trait is multi-modal interactive hybrid agent capability: it can perceive real-world scenes, read screens and interact with GUIs, generate code from visual references, and perform end-to-end navigation within mobile apps.
View modelChoose an eligible model from the catalog or let Smart Mode select a suitable available route. Tavory checks capabilities, plan access and current availability server-side when a request is sent.