LangChain Reference home pageLangChain ReferenceLangChain Reference
  • GitHub
  • Main Docs
Deep Agents
LangChain
LangGraph
Integrations
LangSmith
  • Overview
  • Client
  • AsyncClient
  • Run Helpers
  • Run Trees
  • Evaluation
  • Schemas
  • Utilities
  • Wrappers
  • Anonymizer
  • Testing
  • Expect API
  • Middleware
  • Pytest Plugin
  • Deployment SDK
⌘I

LangChain Assistant

Ask a question to get started

Enter to send•Shift+Enter new line

Menu

OverviewClientAsyncClientRun HelpersRun TreesEvaluationSchemasUtilitiesWrappersAnonymizerTestingExpect APIMiddlewarePytest PluginDeployment SDK
Language
Theme
PythonlangsmithEvaluation

Evaluation

Tools for evaluating functions and models on datasets. Includes evaluators, scoring utilities, and dataset management.

Classes

Class

ExperimentResults

Results container for experiment data with stats and examples.

Breaking change in v0.4.32: The 'stats' field has been split into 'feedback_stats' and 'run_stats'.

Class

ExperimentResultRow

Class

ComparativeExperimentResults

Represents the results of an evaluate_comparative() call.

This class provides an iterator interface to iterate over the comparison results, indexed access by example ID, and properties to access the

Class

AsyncExperimentResults

Class

EvaluationResult

Evaluation result.

Class

EvaluationResults

Batch evaluation results.

This makes it easy for your evaluator to return multiple metrics at once.

Class

RunEvaluator

Evaluator interface class.

Class

DynamicRunEvaluator

A dynamic evaluator that wraps a function and transforms it into a RunEvaluator.

This class is designed to be used with the @run_evaluator decorator, allowing functions that take a Run and an o

Class

ComparisonEvaluationResult

Feedback scores for the results of comparative evaluations.

These are generated by functions that compare two or more runs, returning a ranking or other feedback.

Class

DynamicComparisonRunEvaluator

Compare predictions (as traces) from 2 or more runs.

Class

StringEvaluator

Grades the run's string input, output, and optional answer.

.. deprecated:: 0.5.0

StringEvaluator is deprecated. Use openevals instead: https://github.com/langchain-ai/openevals

Class

LLMEvaluator

A class for building LLM-as-a-judge evaluators.

.. deprecated:: 0.5.0

LLMEvaluator is deprecated. Use openevals instead: https://github.com/langchain-ai/openevals

Functions

Function

evaluate

Evaluate a target system on a given dataset.

Function

evaluate_existing

Evaluate existing experiment runs.

Function

evaluate_comparative

Evaluate existing experiment runs against each other.

This lets you use pairwise preference scoring to generate more reliable feedback in your experiments.

Function

aevaluate

Evaluate an async target system on a given dataset.

Function

aevaluate_existing

Evaluate existing experiment runs asynchronously.

Function

run_evaluator

Create a run evaluator from a function.

Decorator that transforms a function into a RunEvaluator.

Function

comparison_evaluator

Create a comaprison evaluator from a function.