The W&B Weave homepage
Score7.2
Rank#3 of 32
From$60/mo
Free planYes
Free trialYes
Runs onAPI, Linux, macOS, Self-hosted, Web, Windows

Summary

W&B Weave provides observability and continuous improvement tools for production AI agents. It groups traces into sessions and turns, treating steps, tools, and sub-agents as distinct concepts. Built-in and custom signals capture and classify interactions, with alerts available through Slack notifications and webhooks. Its evaluation framework compares results and visualizes potential regressions before deployment, while the Playground lets teams compare models and prompting approaches against production traces. Pre-built scorers cover safety checks such as toxicity, bias, PII, and hallucinations, as well as quality measures including coherence and fluency. Weave works with frameworks such as LangChain and LlamaIndex and models from OpenAI, Meta, Amazon Bedrock, and Anthropic. Python and TypeScript libraries are available, along with OpenTelemetry trace submission to its OTLP endpoint. The Free plan includes 1 GB per month of ingestion and 5 GB of storage. Pro costs 60.00 USD per month, with 1.5 GB of monthly ingestion and additional ingestion at $0.10 per MB. A 30-day trial is listed.

Who it is for

Weave suits teams developing and monitoring production AI agents that need tracing, evaluations, and model or prompt comparisons. It also fits teams that use Python or TypeScript libraries or send OpenTelemetry traces.

What is good

  • Organizes traces into sessions, turns, steps, and tools.
  • Routes alerts through Slack and webhooks.
  • Includes safety and quality scorers.
  • Offers Python and TypeScript libraries.
  • Free plan includes 1 GB monthly ingestion and 5 GB storage.

What to know first

  • Free plan limits ingestion to 1 GB per month.
  • Pro costs 60.00 USD per month.
  • Additional Pro ingestion costs $0.10 per MB.

Everything Xiaomi review

W&B Weave: the full review

W&B Weave brings agent tracing, evaluations, and interaction monitoring together, with options for custom signals and alerts. Check the monthly ingestion allowances against expected usage, especially if choosing Pro.

Overview

W&B Weave is a production observability and improvement platform for AI agents, with tracing, monitoring and evaluation in one toolkit. It is best suited to teams building language-model applications that need to inspect agent behavior and assess changes before deployment. Its breadth is useful, but the monthly ingestion caps make usage costs worth checking before choosing a paid plan.

Founded in 2017 in San Francisco, Weights & Biases was acquired by CoreWeave in May 2025. Weave tracks agent activity alongside model costs, prompts and retrieval, making it relevant to AI Agent Observability Tools, AI LLM Observability Tools, LLM Observability Tools, Model Monitoring Software and AI Agent Evaluation Tools.

Key features

Agent tracing and monitoring

Weave organizes traces into sessions and turns, treating steps, tools and sub-agents as distinct concepts. That structure is useful when a team needs to follow a multi-step interaction rather than view model calls in isolation. Built-in and custom signals capture and classify interactions, while Slack notifications and webhook automations can route alerts to existing workflows.

Cost and token tracking, prompt management, LLM tracing and retrieval tracing broaden the view beyond agent steps. Python and TypeScript libraries, an OTLP endpoint for OpenTelemetry traces, and integrations with LangChain, LlamaIndex, OpenAI, Meta, Amazon Bedrock and Anthropic give teams several ways to connect their stack.

Evaluations, experimentation and guardrails

The evaluation framework offers comparisons and visualizations to help teams spot regressions before deployment. In the Playground, users can compare models and prompting approaches against production traces, tying experimentation to observed application behavior.

Pre-built safety scorers cover toxicity, bias, personally identifiable information and hallucinations; quality scorers assess coherence, fluency and context relevance. These provide useful starting points for assessment, while custom signals allow teams to reflect their own monitoring needs.

Security

W&B states that its platform is certified to ISO/IEC 27001:2022, ISO/IEC 27017:2015 and ISO/IEC 27018:2019, and compliant with SOC 2 Type 2 and HIPAA standards. Data is encrypted in transit with TLS 1.2+ and at rest with AES 256. Enterprise pricing includes options such as single sign-on, audit logs, secure private connectivity and customer-managed encryption keys, which matter to organizations with stricter access or data controls.

Pricing

Weave uses a freemium model, with a 30-day trial and paid plans from $60/mo.

PlanPrice and termsWhat it includes
Free0.00 USD per month1 GB/mo Weave data ingestion, 5 GB/mo storage, AI application evaluations, tracing and scorers.
Pro60.00 USD per month; billed monthly, with annual billing also availableStarts at $60/month, 1.5 GB/mo ingestion and additional ingestion at $0.10/MB; for fewer than 50 employees.
EnterpriseCustom pricing; invoiced annually upfrontCustomizable usage, enterprise support, and security and compliance features.

Free covers core evaluation and tracing, but its 1 GB monthly ingestion allowance and 5 GB storage cap may constrain applications with sustained trace volume. Pro adds just 0.5 GB/mo over Free, so the $0.10/MB charge for additional ingestion is a key cost consideration; its intended audience is organizations with fewer than 50 employees. Pro also includes priority email and chat support. Enterprise suits organizations needing configurable usage or stronger administrative and security controls, with an enterprise support package, but requires custom pricing and annual upfront invoicing.

Platforms

Weave supports API, Linux, macOS, self-hosted, web and Windows environments. Deployment options include both hosted and self-hosted approaches.

Who it's for

Weave is a strong fit for teams operating production AI agents that need to connect traces, evaluations, cost tracking and alerts in a shared workflow. The session-and-turn model and first-class treatment of tools and sub-agents are particularly relevant to applications with multi-step agent behavior. Teams with light usage can start on Free; teams expecting heavier ingestion should estimate volume against Pro’s modest allowance before committing. Organizations requiring enterprise access controls or security options can consider custom Enterprise pricing. It is less compelling for users who only need a standalone evaluator or who want predictable higher-volume ingestion at the entry paid tier.

Pros and cons

  • Pros: Agent traces distinguish sessions, turns, steps, tools and sub-agents, giving teams a structured way to inspect complex interactions.
  • Pros: Evaluations, production-trace comparisons, built-in scorers and alert routing connect assessment with ongoing monitoring.
  • Pros: Python and TypeScript SDKs, OpenTelemetry support and integrations across frameworks and model providers accommodate varied stacks.
  • Cons: Free is capped at 1 GB/mo ingestion and 5 GB/mo storage, which may be restrictive for busy applications.
  • Cons: Pro starts at $60/month for 1.5 GB/mo, and additional ingestion costs $0.10/MB, so usage can materially affect the bill.
  • Cons: Enterprise pricing is custom and invoiced annually upfront, a less flexible commitment for teams that need enterprise controls.

Alternatives

Evidently AI is a free, open-source Apache 2.0 framework for teams prioritizing an open-source option. OpenLIT is a free alternative whose OSS plan offers unlimited self-hosted usage, users, projects and environments, with community support through GitHub. Promptfoo may suit teams centered on evaluation and red-teaming: its free Community plan includes 10k red-team probes/month, all LLM evaluation features and local or self-hosted runs.

Opik is worth considering if running open-source observability and evaluation features locally is a priority. Langfuse offers a Hobby plan with 50k units/month, 30 days of data access and two users for teams comparing a defined free allowance. Arize AX provides a free SaaS tier with 25k trace spans/month, 1 GB ingestion/month, 15-day retention, unlimited users and evaluations, and a 10-issue Signal limit.

Confident AI is another freemium option. Galileo is another freemium option.

Verdict

Choose W&B Weave if your team needs production agent tracing, evaluation, monitoring and experimentation together, especially when tools and sub-agents make behavior difficult to follow. Its integrations and range of signals support a broad workflow, but the limited Pro ingestion allowance and per-MB overage mean it is not the safest choice for teams with high or unpredictable trace volume. Check expected usage first, and look elsewhere if open-source deployment or a generous fixed ingestion allowance is the priority.

W&B Weave plans and pricing

All plans
Free Free 1 GB/mo Weave data ingestion · 5 GB/mo storage · AI application evaluations, tracing, and scorers site.wandb.ai · 30 Sept 2026
Pro $60/mo billed monthly; annual billing also available Starts at $60/month · 1.5 GB/mo Weave data ingestion · additional ingestion $0.10/MB · fewer than 50 employees site.wandb.ai · 30 Sept 2026
Enterprise Not published Custom plans; invoiced annually upfront Customizable usage · enterprise support · security and compliance features site.wandb.ai · 30 Sept 2026

Compared on AI agent observability tools

Free plan
Yessite.wandb.ai
Paid from
$60/mosite.wandb.ai

Facts

Purpose
W&B Weave provides observability and continuous improvement tools for production AI agents.site.wandb.ai · 30 Sept 2026
Tracing
Weave organizes traces into sessions and turns and represents steps, tools, and sub-agents as first-class concepts.site.wandb.ai · 30 Sept 2026
Monitoring
Built-in and custom signals capture and classify agent interactions, and alerts can be routed through Slack notifications and webhook automations.site.wandb.ai · 30 Sept 2026
Evaluations
Weave provides an evaluation framework with comparisons and visualizations to help identify regressions before deployment.site.wandb.ai · 30 Sept 2026
Playground
Playground lets users compare models and prompting techniques against production traces.site.wandb.ai · 30 Sept 2026
Guardrails
Pre-built safety scorers include toxicity, bias, PII detection, and hallucinations, while quality scores include coherence, fluency, and context relevance.site.wandb.ai · 30 Sept 2026
Integrations
W&B says Weave works with popular LLM frameworks including LangChain and LlamaIndex and supports models from OpenAI, Meta, Amazon Bedrock, and Anthropic.site.wandb.ai · 30 Sept 2026
SDKs and API
Weave provides Python and TypeScript libraries and supports sending OpenTelemetry traces to its OTLP endpoint.docs.wandb.ai · 30 Sept 2026
Security
W&B states that its platform is certified to ISO/IEC 27001:2022, ISO/IEC 27017:2015, and ISO/IEC 27018:2019 and is compliant with SOC 2 Type 2 and HIPAA standards.site.wandb.ai · 30 Sept 2026
Encryption
W&B states that data is encrypted in transit using TLS 1.2+ and at rest using AES 256.site.wandb.ai · 30 Sept 2026
Enterprise controls
Enterprise pricing lists options including single sign-on, audit logs, secure private connectivity, and customer-managed encryption keys.site.wandb.ai · 30 Sept 2026
Support
The Pro plan includes priority email and chat support, while Enterprise includes an enterprise support package.site.wandb.ai · 30 Sept 2026
Usage limits
The pricing comparison lists 1 GB per month of Weave data ingestion on Free and 1.5 GB per month on Pro, with additional Pro ingestion priced at $0.10 per MB.site.wandb.ai · 30 Sept 2026
Company
Weights & Biases says it was founded in 2017 by Lukas Biewald, Chris Van Pelt, and Shawn Lewis and was acquired by CoreWeave in May 2025.site.wandb.ai · 30 Sept 2026

Company

Founded
2017site.wandb.ai · 28 Sept 2026
Headquarters
San Francisco, California, United Statessite.wandb.ai · 28 Sept 2026

Best W&B Weave alternatives

See all 12

Where it ranks on Everything Xiaomi

Is W&B Weave yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources