Langfuse v4: up to 165× faster · Read more
DocsData Model

Experiments Data Model

This page describes the data model for experiment-related objects in Langfuse. For an overview of how these objects work together, see the Concepts page. For score and score config objects, see the Scores data model.

For detailed reference please refer to

How experiments are created

You create experiment runs in Langfuse through one of these paths:

PathUse when
Experiments via SDKPython or JS/TS Experiment runner
Experiments via UIPrompt or model experiments from the dataset page
Experiments via OpenTelemetryDirect OTEL ingestion: other languages, custom OTLP pipelines, or re-ingesting experiment traces

To read experiment runs, items, and scores after they exist, use the Experiments API. There is no public REST endpoint for creating new experiment runs; the legacy POST /api/public/dataset-run-items path is deprecated.

Objects

Datasets

Datasets are a collection of inputs and, optionally, expected outputs that can be used during Dataset runs.

Datasets are a collection of DatasetItems.

Dataset object

Prop

Type

DatasetItem object

Prop

Type

DatasetItemMediaReference object

Dataset item media references point from a stored media token in input, expectedOutput, or metadata to a signed media download URL.

Prop

Type

The nested media object contains mediaId, contentType, contentLength, url, and urlExpiry. The url is a signed download URL and should be used before its expiration date. To refresh the signed URL, refetch the dataset.

DatasetRun (Experiment Run)

Dataset runs are used to run a dataset through your LLM application and optionally apply evaluation methods to the results. This is often referred to as Experiment run.


DatasetRun object

Prop

Type

DatasetRunItem object

Prop

Type

Langfuse currently assumes that experiments do not contain repetitions: each dataset item appears once per experiment. Accordingly, reads surface at most one experiment item per dataset item within an experiment. Repetition support is tracked in #5855.

Most of the time, we recommend that DatasetRunItems reference TraceIDs directly. The reference to ObservationID exists for backwards compatibility with older SDK versions.

End to End Data Relations

An experiment can combine a few Langfuse objects:

  • DatasetRuns (or Experiment runs) are created by looping through all or selected DatasetItems of a Dataset with your LLM application.
  • For each DatasetItem passed into the LLM application as an Input a DatasetRunItem & a Trace are created.
  • Optionally Scores can be added to the Traces to evaluate the output of the LLM application during the DatasetRun.

See the Concepts page for more information on how these objects work together conceptually. See the observability core concepts page for more details on traces and observations. See the Scores data model for more details on score and score config objects.

Function Definitions

When running experiments via the SDK, you define task and evaluator functions. These are user-defined functions that the experiment runner calls for each dataset item. For more information on how experiments work conceptually, see the Concepts page.

Task

A task is a function that takes a dataset item and returns an output during an experiment run.

See SDK references for function signatures and parameters:

Evaluator

An evaluator is a function that scores the output of a task for a single dataset item. Evaluators receive the input, output, expected output, and metadata, and return an Evaluation object that becomes a Score in Langfuse.

See SDK references for function signatures and parameters:

Run Evaluator

A run evaluator is a function that assesses the full experiment results and computes aggregate metrics. When run on Langfuse datasets, the resulting scores are attached to the dataset run.

See SDK references for function signatures and parameters:

For detailed usage examples of tasks and evaluators, see Experiments via SDK. For ingesting experiment traces without the SDK, see Experiments via OpenTelemetry.

Local Datasets

With Langfuse v4 and the current SDKs, experiments on local data appear under Experiments without a hosted dataset. Each task execution also creates a trace for debugging. See Compare experiments.


Was this page helpful?

Last updated on