Tavory wird geöffnet
Loading Tavory
Tavory wird geöffnet
Loading Tavory
Provider directory
Compare 5 distinct Fireworks models represented in Tavory across text and chat. Regional duplicates are consolidated into canonical model pages.
Tavory is an independent AI workspace and does not represent a manufacturer relationship unless explicitly stated. Model ownership and trademarks remain with their respective providers.
Distinct models
5
Categories
text
Catalog policy
Regional duplicates consolidated
Fireworks
DeepSeek V4 Flash Vision Exp is the first experimental multimodal model in the DeepSeek V4 family. It builds on DeepSeek V4 Flash with added visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text only agent tasks. 305B MoE, 1M token context. Served via Fireworks.
Fireworks
GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths. A 743B-parameter model in Zhipu AI's GLM series, designed to plan, execute, and iterate autonomously on extended, engineering-grade tasks.
Fireworks
GPT-OSS-20B is OpenAI's lightweight open-weight model with 21.5 billion total parameters (3.6 billion active, Mixture-of-Experts). Optimized for lower latency and local or specialized use cases including agentic tasks, chain-of-thought reasoning, and web browsing. Released under Apache 2.0 license.
Fireworks
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.
Fireworks
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
Choose an eligible model from the catalog or let Smart Mode select a suitable available route. Tavory checks capabilities, plan access and current availability server-side when a request is sent.