Tools for evaluating functions and models on datasets. Includes evaluators, scoring utilities, and dataset management.
Results container for experiment data with stats and examples.
Breaking change in v0.4.32: The 'stats' field has been split into 'feedback_stats' and 'run_stats'.
Represents the results of an evaluate_comparative() call.
This class provides an iterator interface to iterate over the comparison results, indexed access by example ID, and properties to access the
Evaluation result.
Batch evaluation results.
This makes it easy for your evaluator to return multiple metrics at once.
Evaluator interface class.
A dynamic evaluator that wraps a function and transforms it into a RunEvaluator.
This class is designed to be used with the @run_evaluator decorator, allowing
functions that take a Run and an o
Feedback scores for the results of comparative evaluations.
These are generated by functions that compare two or more runs, returning a ranking or other feedback.
Compare predictions (as traces) from 2 or more runs.
Grades the run's string input, output, and optional answer.
.. deprecated:: 0.5.0
StringEvaluator is deprecated. Use openevals instead: https://github.com/langchain-ai/openevals
A class for building LLM-as-a-judge evaluators.
.. deprecated:: 0.5.0
LLMEvaluator is deprecated. Use openevals instead: https://github.com/langchain-ai/openevals
Evaluate a target system on a given dataset.
Evaluate existing experiment runs.
Evaluate existing experiment runs against each other.
This lets you use pairwise preference scoring to generate more reliable feedback in your experiments.
Evaluate an async target system on a given dataset.
Evaluate existing experiment runs asynchronously.
Create a run evaluator from a function.
Decorator that transforms a function into a RunEvaluator.
Create a comaprison evaluator from a function.