diff --git a/.devcontainer/devcontainer.json b/.devcontainer/devcontainer.json
index f1baedb..5a81e51 100644
--- a/.devcontainer/devcontainer.json
+++ b/.devcontainer/devcontainer.json
@@ -5,7 +5,8 @@
// attach to container from docker compose setup
"dockerComposeFile": ["../docker-compose.dev.yml"],
- "service": "autonomy_benchmarks",
+ "service": "autonomy_evaluation",
+ "shutdownAction": "none",
// set up git pre-commit hooks
"onCreateCommand": "pip install pre-commit && apt update && apt install locales -y && locale-gen en_US.UTF-8",
diff --git a/.github/workflows/docker-ros.yml b/.github/workflows/docker-ros.yml
index 82f4d8d..bd99cc8 100644
--- a/.github/workflows/docker-ros.yml
+++ b/.github/workflows/docker-ros.yml
@@ -22,7 +22,7 @@ jobs:
target: dev,run
base-image: rwthika/ros2:jazzy
rmw-implementation: rmw_zenoh_cpp
- command: ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py
+ command: ros2 launch autonomy_evaluation autonomy_evaluation.launch.py
enable-slim: false
env:
TARGET_CMAKE_ARGS: -DCMAKE_EXPORT_COMPILE_COMMANDS=1
diff --git a/README.md b/README.md
index a008294..c528cf5 100644
--- a/README.md
+++ b/README.md
@@ -1,36 +1,37 @@
-# autonomy_benchmarks
+# autonomy_evaluation
-
-
+
+
-
-
-
-
-
+
+
+
+
+
-> This repository will be part of the **Autonomy.Hub Ecosystem**
+> This repository is part of the **Autonomy.Benchmarks** suite of the **Autonomy.Hub Ecosystem**
-As part of the Autonomy.Hub Ecosystem, **Autonomy.Benchmarks** enables the Automated Driving community to easily benchmark their automated driving building blocks across different tasks and datasets:
+Within the Autonomy.Benchmarks suite, **Autonomy.Evaluation** generates the metrics-based evidence for benchmarking automated driving deployments. It evaluates arbitrary ROS systems under test, either from their own topics alone, e.g. a closed-loop planner by its time to collision, or against ground truth, e.g. a perception algorithm against the labels of a dataset replayed by [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets), and reports the resulting metrics per scene and over all evaluated samples:
-- 🔄 **Unified ROS 2 Interface**: Work with multiple datasets using the benefits of the ROS 2 ecosystem
-- 📊 **Comprehensive Benchmarks**: Use the provided benchmarks with [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) to benchmark building blocks across different automated driving tasks
+- 🔄 **Unified ROS 2 Interface**: Evaluate any ROS system under test, on datasets, in simulation or live, using the benefits of the ROS 2 ecosystem
+- 🧩 **Pluggable Evaluations**: Select an evaluation by name, or bring your own from another package, reading any topics as inputs and, where needed, ground truth
+- 📊 **Established Metrics**: Use the provided evaluations, which follow the protocols of established challenges, with [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) across different automated driving tasks
- ⚡ **Efficient Data Pipeline**: Works seamlessly with preprocessed Rosbag files from [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) for fast execution during development
- 🐳 **Dockerized Environment**: Reproducible setup with all dependencies included
- 🔌 **Modular Architecture**: Easy integration with other ROS 2 packages
-## Supported Benchmarks
+## Supported Evaluations
-This repository supports various automated driving evaluation benchmarks.
+This repository supports evaluations of various automated driving tasks.
Detailed metric definitions and computation notes are documented in [docs/IMPLEMENTATION.md](docs/IMPLEMENTATION.md).
-> [**Contributions**](docs/IMPLEMENTATION.md#adding-more-benchmarks) adding more benchmarks are welcome
+> [**Contributions**](docs/IMPLEMENTATION.md#adding-more-evaluations) adding more evaluations are welcome
-| Benchmark | Challenge | Dataset | Task |
+| Evaluation | Challenge | Dataset | Task |
| --------- | --------- | ------- | ---- |
| [**nuScenes 3D Lidar Object Detection**](docs/IMPLEMENTATION.md#3d-lidar-object-detection) | [](https://www.nuscenes.org/object-detection) | [nuScenes](https://github.com/thinking-cars/autonomy_datasets) | 3D bounding box detection from lidar |
@@ -43,7 +44,7 @@ Detailed metric definitions and computation notes are documented in [docs/IMPLEM
Clone [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) and follow its setup instructions to prepare your dataset.
-Use the provided [docker-compose.yml](docker-compose.yml) to start the full pipeline — dataset publisher, system-under-test, and benchmark node:
+Use the provided [docker-compose.yml](docker-compose.yml) to start the full pipeline — dataset publisher, system under test, and evaluation node:
```bash
# enable GUI output from Docker container
@@ -57,13 +58,13 @@ docker compose up -d
docker compose down
```
-Configure the benchmark task and dataset via ROS launch arguments in [docker-compose.yml](docker-compose.yml):
+Configure the evaluation and dataset via ROS launch arguments in [docker-compose.yml](docker-compose.yml):
```yaml
-command: ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py benchmark:=nuscenes_lidar_object_detection prediction:=$your_prediction_topic label:=$your_label_topic request_samples:=/datasets/request_samples visualize:=true
+command: ros2 launch autonomy_evaluation autonomy_evaluation.launch.py evaluation:=nuscenes_lidar_object_detection prediction:=$your_prediction_topic label:=$your_label_topic request_samples:=/datasets/request_samples visualize:=true
```
-The benchmark node requests the samples it evaluates from the dataset node via its `request_samples` service, which publishes them and responds once they have been published. The dataset therefore publishes the next sample only once the system under test has processed the current one. As soon as all samples have been published, the benchmark aggregates its metrics per sample, per scene of the dataset and over the whole benchmark. See the [node documentation](autonomy_benchmarks/README.md#autonomy_benchmarks) for the sample request settings and the results.
+The evaluation node requests the samples it evaluates from the dataset node via its `request_samples` service, which publishes them and responds once they have been published. The dataset therefore publishes the next sample only once the system under test has processed the current one. As soon as all samples have been published, the node aggregates its metrics per scene of the dataset and over all evaluated samples. To evaluate samples published by others instead, e.g. by a closed-loop simulation, set `sample_source:=external`. See the [node documentation](autonomy_evaluation/README.md#autonomy_evaluation) for the topics of the evaluations, the sample settings and the results.
## 💻 Development
@@ -71,11 +72,11 @@ The benchmark node requests the samples it evaluates from the dataset node via i
1. Clone the repository.
```bash
- git clone https://github.com/thinking-cars/autonomy_benchmarks.git
+ git clone https://github.com/thinking-cars/autonomy_evaluation.git
```
1. Initialize the [`.openads-dev-environment`](https://github.com/openads-project/openads-dev-environment) submodule containing development environment configuration.
```bash
- cd autonomy_benchmarks
+ cd autonomy_evaluation
git submodule update --init --recursive
```
1. Open the repository in [Visual Studio Code](https://code.visualstudio.com).
@@ -108,11 +109,11 @@ colcon test-result --verbose
## 📝 Documentation
-Package and node interfaces are documented in the respective package READMEs listed below. Implementation details are found in the [Source Code Documentation](https://thinking-cars.github.io/autonomy_benchmarks).
+Package and node interfaces are documented in the respective package READMEs listed below. Implementation details are found in the [Source Code Documentation](https://thinking-cars.github.io/autonomy_evaluation).
| Package | Description |
| --- | --- |
-| [autonomy_benchmarks](autonomy_benchmarks/README.md) | Benchmarking suite for automated driving tasks |
+| [autonomy_evaluation](autonomy_evaluation/README.md) | Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments |
## ⚖️ Licensing
diff --git a/autonomy_benchmarks/README.md b/autonomy_benchmarks/README.md
deleted file mode 100644
index 33ac320..0000000
--- a/autonomy_benchmarks/README.md
+++ /dev/null
@@ -1,86 +0,0 @@
-# `autonomy_benchmarks`
-
-Benchmarking suite for automated driving tasks
-
-## Nodes
-
-### `autonomy_benchmarks`
-
-The node requests the samples it evaluates from the dataset, using the `request_samples` service
-of [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets), which publishes them
-and responds once they have been published. By default one sample is requested at a time, so the
-dataset only publishes the next sample once the system under test has delivered its output for
-the current one and the benchmark has evaluated it. Increase `samples_per_request` to publish
-samples in batches, set it to `0` to publish the whole dataset with a single request, or list the
-IDs of individual samples in `sample_ids` to evaluate only those.
-
-Once the dataset reports that all requested samples have been published, the per-sample metrics
-are aggregated and reported on three levels: for the whole benchmark (`aggregated_metrics`), for
-the samples of each scene the dataset published them from (`scene_results`, matched with the
-samples via the `published_scene_ids` of the responses), and for every single sample
-(`sample_results`). The dataset metrics are logged, and the results of all three levels are
-written to a JSON file if `results_path` is set. Samples that are not evaluated within
-`evaluation_timeout` seconds of being published, e.g. because the system under test skipped them,
-are left out.
-
-```bash
-ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \
- prediction:=/object_list/prediction \
- label:=/object_list/lidar_01 \
- request_samples:=/datasets/request_samples \
- results_path:=/results/nuscenes_lidar_object_detection.json
-```
-
-To look at the samples one by one, set `manual_playback` together with `visualize`. The node then requests no samples itself, and the samples are published with the [playback panel](https://github.com/thinking-cars/autonomy_datasets/blob/main/autonomy_datasets_rviz_plugins/README.md) in RViz instead, whose _Service_ field has to name the `request_samples` service of the dataset node (`/datasets/request_samples` by default). Every sample that arrives is evaluated and shown in RViz, and `samples_per_request`, `sample_ids` and `evaluation_timeout` have no effect. As the responses of the dataset only reach the panel, the node neither learns the scenes of the samples, so `scene_results` stays empty, nor when the dataset has ended: the results of the evaluated samples are reported once the node is stopped, e.g. with Ctrl-C, and are marked as incomplete.
-
-```bash
-ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \
- prediction:=/object_list/prediction \
- label:=/object_list/lidar_01 \
- visualize:=true \
- manual_playback:=true
-```
-
-```mermaid
-flowchart LR
- NODE("autonomy_benchmarks")
- NODE o--o|~/request_samples| SC0:::hidden
- classDef hidden display: none;
-```
-
-#### Service Clients
-
-| Service | Type | Description |
-| --- | --- | --- |
-| `~/request_samples` | `autonomy_datasets_msgs/srv/RequestSamples` | request samples from dataset |
-
-#### Parameters
-
-| Parameter | Type | Default | Description |
-| --- | --- | --- | --- |
-| `benchmark` | `string` | `nuscenes_lidar_object_detection` | benchmark name |
-| `visualize` | `bool` | `false` | publish the per-sample true positives, false positives and false negatives for RViz |
-| `manual_playback` | `bool` | `false` | leave requesting the samples of the dataset to the user |
-| `samples_per_request` | `int` | `1` | number of samples to request from the dataset at a time; 0 requests all remaining samples at once, 1 evaluates every sample before the next one is published |
-| `sample_ids` | `string` | - | comma-separated IDs of the dataset samples to evaluate (e.g. '0,10,20'); if empty, all samples of the dataset are evaluated |
-| `evaluation_timeout` | `float` | `60.0` | seconds to wait for a published sample to be evaluated before continuing without it |
-| `results_path` | `string` | - | path of the JSON file the benchmark results are written to; results are only logged if empty |
-
-## Launch Files
-
-### [`autonomy_benchmarks.launch.py`](launch/autonomy_benchmarks.launch.py)
-
-| Argument | Default | Description |
-| --- | --- | --- |
-| `request_samples` | `"~/request_samples"` | service of the dataset node used to request the samples to evaluate |
-| `benchmark` | `"nuscenes_lidar_object_detection"` | benchmark to run |
-| `name` | `"autonomy_benchmarks"` | node name |
-| `namespace` | `""` | node namespace |
-| `log_level` | `"info"` | ros logging level |
-| `use_sim_time` | `"true"` | use sim time |
-| `visualize` | `"false"` | publish the per-sample true positives, false positives and false negatives and open RViz |
-| `manual_playback` | `"false"` | request the samples to evaluate via the playback panel in RViz instead of one after another, and open RViz on it |
-| `samples_per_request` | `"1"` | number of samples to request from the dataset at a time (0 requests all remaining samples at once) |
-| `sample_ids` | `""` | comma-separated IDs of the dataset samples to evaluate (all samples if empty) |
-| `evaluation_timeout` | `"60.0"` | seconds to wait for a published sample to be evaluated before continuing without it |
-| `results_path` | `""` | path of the JSON file the benchmark results are written to (results are only logged if empty) |
diff --git a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py b/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py
deleted file mode 100644
index bdb62a9..0000000
--- a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py
+++ /dev/null
@@ -1,8 +0,0 @@
-# Copyright Thinking Cars GmbH
-# SPDX-License-Identifier: Apache-2.0
-
-"""Benchmark implementations for AutonomyHub."""
-
-from autonomy_benchmarks.benchmarks.AutonomyBenchmark import AutonomyBenchmark
-
-__all__ = ["AutonomyBenchmark"]
diff --git a/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py b/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py
deleted file mode 100644
index 9a567d6..0000000
--- a/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py
+++ /dev/null
@@ -1,144 +0,0 @@
-#!/usr/bin/env python3
-
-# Copyright Thinking Cars GmbH
-# SPDX-License-Identifier: Apache-2.0
-
-from launch import LaunchDescription
-from launch.actions import DeclareLaunchArgument
-from launch.conditions import IfCondition
-from launch.substitutions import LaunchConfiguration, PathJoinSubstitution
-from launch_ros.actions import Node, SetParameter
-from launch_ros.parameter_descriptions import ParameterValue
-from launch_ros.substitutions import FindPackageShare
-
-# Names of the benchmark's data inputs. Each name must match a key returned by
-# the benchmark's ``required_inputs()`` (and thus a ``compute_sample_metrics``
-# parameter). This single list is the source of truth: it drives both the
-# per-input launch arguments and the topic remappings below, so adding a new
-# input (e.g. "map", "radar") is a one-line change here.
-_INPUTS = ["prediction", "label"]
-
-# Inputs whose topic is derived from another input's topic instead of getting a
-# launch argument of its own, as ``input name -> (source input, topic suffix)``.
-# The dataset publishes the meta information of an object list next to it, on
-# "/meta_info".
-_DERIVED_INPUTS = {"label_meta_info": ("label", "/meta_info")}
-
-
-def generate_launch_description():
- """Create and return the launch description for the autonomy_benchmarks node."""
-
- # Service of the dataset node the benchmark requests the samples to evaluate from, remapped
- # onto the topic its argument resolves to just like the benchmark's data inputs.
- remappable_topics = [
- DeclareLaunchArgument(
- "request_samples",
- default_value="~/request_samples",
- description="service of the dataset node used to request the samples to evaluate",
- ),
- ]
-
- args = [
- *remappable_topics,
- DeclareLaunchArgument(
- "benchmark",
- default_value="nuscenes_lidar_object_detection",
- description="benchmark name",
- choices=["nuscenes_lidar_object_detection"],
- ),
- DeclareLaunchArgument("name", default_value="autonomy_benchmarks", description="node name"),
- DeclareLaunchArgument("namespace", default_value="", description="node namespace"),
- DeclareLaunchArgument(
- "log_level", default_value="info", description="ROS logging level (debug, info, warn, error, fatal)"
- ),
- DeclareLaunchArgument("use_sim_time", default_value="true", description="use simulation clock"),
- DeclareLaunchArgument(
- "visualize",
- default_value="false",
- choices=["true", "false"],
- description="publish the per-sample true positives, false positives and false negatives and open RViz on them",
- ),
- DeclareLaunchArgument(
- "manual_playback",
- default_value="false",
- choices=["true", "false"],
- description="request the samples to evaluate via the playback panel in RViz instead of one after another, "
- "and open RViz on it",
- ),
- DeclareLaunchArgument(
- "samples_per_request",
- default_value="1",
- description="number of samples to request from the dataset at a time (0 requests all remaining samples at once)",
- ),
- DeclareLaunchArgument(
- "sample_ids",
- default_value="",
- description="comma-separated IDs of the dataset samples to evaluate (all samples if empty)",
- ),
- DeclareLaunchArgument(
- "evaluation_timeout",
- default_value="60.0",
- description="seconds to wait for a published sample to be evaluated before continuing without it",
- ),
- DeclareLaunchArgument(
- "results_path",
- default_value="",
- description="path of the JSON file the benchmark results are written to (results are only logged if empty)",
- ),
- # One argument per benchmark input; defaults to the node-relative name so
- # an unset input is a no-op remap. Override with e.g. prediction:=/real/topic.
- *[
- DeclareLaunchArgument(
- name,
- default_value=f"~/{name}",
- description=f"real ROS topic feeding the benchmark's '{name}' input",
- )
- for name in _INPUTS
- ],
- ]
-
- # Remap each node-relative input name onto the topic its argument resolves to,
- # then onto the topic derived from another input for the derived inputs.
- remappings = [(name, LaunchConfiguration(name)) for name in _INPUTS]
- remappings += [(name, [LaunchConfiguration(source), suffix]) for name, (source, suffix) in _DERIVED_INPUTS.items()]
- remappings += [(la.default_value[0].text, LaunchConfiguration(la.name)) for la in remappable_topics]
-
- node = Node(
- package="autonomy_benchmarks",
- executable="autonomy_benchmarks",
- namespace=LaunchConfiguration("namespace"),
- name=LaunchConfiguration("name"),
- parameters=[
- {"benchmark": LaunchConfiguration("benchmark")},
- {"visualize": ParameterValue(LaunchConfiguration("visualize"), value_type=bool)},
- {"manual_playback": ParameterValue(LaunchConfiguration("manual_playback"), value_type=bool)},
- {"samples_per_request": ParameterValue(LaunchConfiguration("samples_per_request"), value_type=int)},
- {"sample_ids": ParameterValue(LaunchConfiguration("sample_ids"), value_type=str)},
- {"evaluation_timeout": ParameterValue(LaunchConfiguration("evaluation_timeout"), value_type=float)},
- {"results_path": ParameterValue(LaunchConfiguration("results_path"), value_type=str)},
- ],
- arguments=["--ros-args", "--log-level", LaunchConfiguration("log_level")],
- remappings=remappings,
- output="screen",
- emulate_tty=True,
- )
- rviz = Node(
- package="rviz2",
- executable="rviz2",
- name="rviz2",
- arguments=[
- "-d",
- PathJoinSubstitution([FindPackageShare("autonomy_benchmarks"), "config", "conf.rviz"]),
- ],
- condition=IfCondition(LaunchConfiguration("visualize")),
- output="screen",
- )
-
- return LaunchDescription(
- [
- *args,
- SetParameter("use_sim_time", LaunchConfiguration("use_sim_time")),
- node,
- rviz,
- ]
- )
diff --git a/autonomy_benchmarks/setup.cfg b/autonomy_benchmarks/setup.cfg
deleted file mode 100644
index 7e58c70..0000000
--- a/autonomy_benchmarks/setup.cfg
+++ /dev/null
@@ -1,4 +0,0 @@
-[develop]
-script_dir=$base/lib/autonomy_benchmarks
-[install]
-install_scripts=$base/lib/autonomy_benchmarks
diff --git a/autonomy_benchmarks/tests/benchmarks/__init__.py b/autonomy_benchmarks/tests/benchmarks/__init__.py
deleted file mode 100644
index 5b0f7f7..0000000
--- a/autonomy_benchmarks/tests/benchmarks/__init__.py
+++ /dev/null
@@ -1,4 +0,0 @@
-# Copyright Thinking Cars GmbH
-# SPDX-License-Identifier: Apache-2.0
-
-"""Benchmark-oriented test modules for autonomy_benchmarks."""
diff --git a/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py b/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py
deleted file mode 100644
index 4bc831e..0000000
--- a/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py
+++ /dev/null
@@ -1,194 +0,0 @@
-# Copyright Thinking Cars GmbH
-# SPDX-License-Identifier: Apache-2.0
-
-"""Tests for the result store of the AutonomyBenchmark base class.
-
-Metrics are reported on three levels: for every single sample, aggregated over the samples of each
-scene of the dataset, and aggregated over all evaluated samples. A minimal benchmark whose metrics
-are trivial to predict is used, so that the tests cover the grouping and not a metric definition.
-"""
-
-from __future__ import annotations
-
-import json
-from typing import Any, Dict, List
-
-from autonomy_benchmarks.benchmarks.AutonomyBenchmark import AutonomyBenchmark
-
-
-class CountingBenchmark(AutonomyBenchmark):
- """Minimal benchmark counting the objects of a sample and summing them up when aggregating."""
-
- def __init__(self) -> None:
- """Name the benchmark."""
- super().__init__(name="counting", description="counts objects")
-
- def required_inputs(self) -> Dict[str, Any]:
- """Declare the inputs, which this benchmark does not read from ROS messages."""
- return {"prediction": object, "label": object}
-
- def compute_sample_metrics(self, prediction: Any, label: Any, sample_id: str = None) -> Dict[str, Any]:
- """Report the given prediction and label counts of a single sample."""
- return {"num_predictions": prediction, "num_labels": label}
-
- def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Dict[str, Any]:
- """Sum the counts of the given samples."""
- return {
- "num_predictions": sum(entry["metrics"]["num_predictions"] for entry in sample_results),
- "num_labels": sum(entry["metrics"]["num_labels"] for entry in sample_results),
- }
-
-
-def _benchmark_of(samples) -> CountingBenchmark:
- """Record ``(sample_id, scene_id, num_predictions, num_labels)`` samples in a benchmark."""
- benchmark = CountingBenchmark()
- for sample_id, scene_id, num_predictions, num_labels in samples:
- benchmark.record_sample(prediction=num_predictions, label=num_labels, sample_id=sample_id, scene_id=scene_id)
- return benchmark
-
-
-class TestSampleResults:
- """Tests recording samples with the scene of the dataset they belong to."""
-
- # The recorded samples themselves are no longer reported alongside the aggregated
- # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize()
- # def test_records_sample_with_its_scene(self):
- # """A recorded sample keeps its ID, its scene and its metrics."""
- # benchmark = _benchmark_of([("0", "scene_a", 2, 3)])
- #
- # assert benchmark.finalize()["sample_results"] == [
- # {"sample_id": "0", "scene_id": "scene_a", "metrics": {"num_predictions": 2, "num_labels": 3}}
- # ]
-
- def test_scene_can_be_set_after_the_sample_was_recorded(self):
- """An evaluation loop that learns the scene late sets it on the returned entry."""
- benchmark = CountingBenchmark()
-
- entry = benchmark.record_sample(prediction=1, label=1, sample_id="0")
- assert entry["scene_id"] is None
- entry["scene_id"] = "scene_a"
-
- assert benchmark.sample_results_by_scene() == {"scene_a": [entry]}
-
-
-class TestFinalize:
- """Tests aggregating the recorded samples per scene and over the whole benchmark."""
-
- def test_aggregates_per_sample_scene_and_benchmark(self):
- """Metrics are reported for every sample, every scene and all samples together."""
- results = _benchmark_of(
- [
- ("0", "scene_a", 1, 1),
- ("1", "scene_a", 2, 3),
- ("2", "scene_b", 4, 5),
- ]
- ).finalize()
-
- assert results["num_samples"] == 3
- assert results["num_scenes"] == 2
- assert results["aggregated_metrics"] == {"num_predictions": 7, "num_labels": 9}
- assert results["scene_results"]["scene_a"] == {
- "num_samples": 2,
- "sample_ids": ["0", "1"],
- "aggregated_metrics": {"num_predictions": 3, "num_labels": 4},
- }
- assert results["scene_results"]["scene_b"]["aggregated_metrics"] == {"num_predictions": 4, "num_labels": 5}
- # The metrics of the single samples are no longer reported alongside the aggregated
- # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize()
- # assert [entry["metrics"] for entry in results["sample_results"]] == [
- # {"num_predictions": 1, "num_labels": 1},
- # {"num_predictions": 2, "num_labels": 3},
- # {"num_predictions": 4, "num_labels": 5},
- # ]
-
- def test_groups_samples_of_a_scene_that_are_not_recorded_consecutively(self):
- """Samples are grouped by their scene, not by the order they were recorded in."""
- results = _benchmark_of(
- [
- ("0", "scene_a", 1, 0),
- ("1", "scene_b", 2, 0),
- ("2", "scene_a", 4, 0),
- ]
- ).finalize()
-
- assert results["scene_results"]["scene_a"]["sample_ids"] == ["0", "2"]
- assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 5
- assert results["scene_results"]["scene_b"]["sample_ids"] == ["1"]
-
- def test_samples_without_a_scene_are_only_aggregated_over_the_benchmark(self):
- """A sample that cannot be attributed to a scene still counts for the whole benchmark."""
- results = _benchmark_of([("0", "scene_a", 1, 0), ("1", None, 2, 0)]).finalize()
-
- assert results["num_samples"] == 2
- assert results["num_scenes"] == 1
- assert results["aggregated_metrics"]["num_predictions"] == 3
- assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 1
-
- def test_reports_no_scene_without_recorded_scenes(self):
- """Samples recorded without a scene aggregate to no scene results at all."""
- results = _benchmark_of([("0", None, 1, 0)]).finalize()
-
- assert results["num_scenes"] == 0
- assert results["scene_results"] == {}
-
-
-class TestSaveResults:
- """Tests writing the results of all three levels to a JSON file."""
-
- def test_writes_sample_scene_and_benchmark_metrics(self, tmp_path):
- """The stored results hold the metrics of every sample, every scene and the benchmark."""
- benchmark = _benchmark_of([("0", "scene_a", 1, 1), ("1", "scene_b", 2, 2)])
-
- output_path = benchmark.save_results(str(tmp_path / "results" / "counting.json"))
-
- stored = json.loads(open(output_path).read())
- assert stored["aggregated_metrics"] == {"num_predictions": 3, "num_labels": 3}
- assert sorted(stored["scene_results"]) == ["scene_a", "scene_b"]
- # The single samples are no longer written alongside the aggregated
- # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize()
- # assert [entry["sample_id"] for entry in stored["sample_results"]] == ["0", "1"]
-
- def test_writes_previously_computed_results(self, tmp_path):
- """Results that have already been computed are written as they are."""
- benchmark = _benchmark_of([("0", "scene_a", 1, 1)])
- results = benchmark.finalize()
-
- output_path = benchmark.save_results(str(tmp_path / "counting.json"), results=results)
-
- assert json.loads(open(output_path).read())["aggregated_metrics"] == results["aggregated_metrics"]
-
-
-class TestIncompleteResults:
- """Tests marking the results of an evaluation that did not process all samples."""
-
- def test_results_of_all_samples_are_complete(self):
- """Results aggregated after the last sample of the benchmark are marked as complete."""
- assert _benchmark_of([("0", "scene_a", 1, 1)]).finalize()["complete"] is True
-
- def test_interrupted_results_are_marked_incomplete(self):
- """Results aggregated before the last sample, e.g. after Ctrl-C, are marked as incomplete."""
- results = _benchmark_of([("0", "scene_a", 1, 1)]).finalize(complete=False)
-
- assert results["complete"] is False
- # the samples that were evaluated are still reported
- assert results["num_samples"] == 1
- assert results["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1}
-
- def test_writes_incomplete_results_to_the_results_file(self, tmp_path):
- """The results file of an interrupted evaluation marks the results it holds as incomplete."""
- benchmark = _benchmark_of([("0", "scene_a", 1, 1)])
-
- output_path = benchmark.save_results(str(tmp_path / "counting.json"), complete=False)
-
- stored = json.loads(open(output_path).read())
- assert stored["complete"] is False
- assert stored["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1}
-
- def test_written_results_keep_the_flag_they_were_finalized_with(self, tmp_path):
- """A given payload is written as it is, marked the way it was finalized."""
- benchmark = _benchmark_of([("0", "scene_a", 1, 1)])
- results = benchmark.finalize(complete=False)
-
- output_path = benchmark.save_results(str(tmp_path / "counting.json"), results=results)
-
- assert json.loads(open(output_path).read())["complete"] is False
diff --git a/autonomy_benchmarks/tests/test_autonomy_benchmarks.py b/autonomy_benchmarks/tests/test_autonomy_benchmarks.py
deleted file mode 100644
index 651eed7..0000000
--- a/autonomy_benchmarks/tests/test_autonomy_benchmarks.py
+++ /dev/null
@@ -1,307 +0,0 @@
-# Copyright Thinking Cars GmbH
-# SPDX-License-Identifier: Apache-2.0
-
-"""Tests for the helpers of the autonomy_benchmarks node.
-
-The node itself drives the evaluation via the ``request_samples`` service of the dataset and is
-covered by running it against a dataset. Tested here are the parsing of the samples to evaluate,
-as an unparsable value stops the node, the matching of the received input messages into the
-samples to evaluate, which has to hold up when the dataset continues with a scene that was
-recorded before the scene played before it, the evaluation of a sample, which requests the next
-samples unless the user controls the playback manually, and the finalization of the results,
-which reports the samples of an interrupted evaluation as incomplete.
-"""
-
-from __future__ import annotations
-
-from collections import deque
-from types import SimpleNamespace
-
-import pytest
-from autonomy_benchmarks.autonomy_benchmarks import AutonomyBenchmarks, parse_sample_ids, SampleSynchronizer
-
-_TOPICS = ["prediction", "label", "label_meta_info"]
-
-# stamps of the last sample of a scene and of the first samples of the scene the dataset continues
-# with, which nuScenes recorded years earlier
-_PREVIOUS_SCENE = (1537853053, 397270000)
-_NEXT_SCENE = [(1531885320, 49418000), (1531885320, 548742000), (1531885321, 48634000)]
-
-
-def _message(stamp: tuple[int, int]) -> SimpleNamespace:
- """Fake an input message stamped with the recording time of its sample."""
- return SimpleNamespace(header=SimpleNamespace(stamp=SimpleNamespace(sec=stamp[0], nanosec=stamp[1])))
-
-
-def _synchronizer(queue_size: int = 10) -> tuple[SampleSynchronizer, list]:
- """Create a synchronizer of the benchmark inputs next to the list of the samples it matched."""
- matched_samples: list = []
- synchronizer = SampleSynchronizer(_TOPICS, callback=lambda *msgs: matched_samples.append(msgs), queue_size=queue_size)
- return synchronizer, matched_samples
-
-
-def _publish_sample(synchronizer: SampleSynchronizer, stamp: tuple[int, int], topics=None) -> dict:
- """Add one message per given input (all of them by default), all stamped with the same time."""
- messages = {topic: _message(stamp) for topic in topics or _TOPICS}
- for topic, message in messages.items():
- synchronizer.add(topic, message)
- return messages
-
-
-class TestParseSampleIds:
- """Tests parsing the 'sample_ids' parameter into the sample IDs to request."""
-
- def test_parses_comma_separated_ids(self):
- """Sample IDs are parsed in the given order."""
- assert parse_sample_ids("0,10,20") == [0, 10, 20]
-
- def test_parses_single_id(self):
- """A single ID is a valid request."""
- assert parse_sample_ids("7") == [7]
-
- def test_ignores_surrounding_whitespace(self):
- """Sample IDs separated by ', ' are parsed like IDs separated by ','."""
- assert parse_sample_ids(" 1, 2 ,3 ") == [1, 2, 3]
-
- @pytest.mark.parametrize("sample_ids", ["", " ", ","])
- def test_no_ids_evaluate_the_whole_dataset(self, sample_ids):
- """An empty value requests no specific samples, so the whole dataset is evaluated."""
- assert parse_sample_ids(sample_ids) == []
-
- @pytest.mark.parametrize("sample_ids", ["1;2", "first", "1.5", "1-2"])
- def test_rejects_values_that_are_no_ids(self, sample_ids):
- """A value that is no comma-separated list of IDs is rejected."""
- with pytest.raises(ValueError):
- parse_sample_ids(sample_ids)
-
-
-class TestSampleSynchronizer:
- """Tests matching the messages of the benchmark inputs into the samples to evaluate."""
-
- def test_matches_the_messages_of_a_sample_in_input_order(self):
- """A sample is reported once every input has been received, in the order of the inputs."""
- synchronizer, matched_samples = _synchronizer()
-
- messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=list(reversed(_TOPICS)))
-
- assert matched_samples == [tuple(messages[topic] for topic in _TOPICS)]
- assert not synchronizer.incomplete_samples
-
- def test_waits_for_the_missing_inputs_of_a_sample(self):
- """A sample of which an input is missing is not reported yet."""
- synchronizer, matched_samples = _synchronizer()
-
- _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"])
-
- assert matched_samples == []
-
- def test_matches_samples_of_a_scene_recorded_before_the_previous_scene(self):
- """The dataset continues with an older scene, whose samples are matched all the same."""
- synchronizer, matched_samples = _synchronizer()
-
- _publish_sample(synchronizer, _PREVIOUS_SCENE)
- messages = _publish_sample(synchronizer, _NEXT_SCENE[0])
-
- assert len(matched_samples) == 2
- assert matched_samples[-1] == tuple(messages[topic] for topic in _TOPICS)
-
- def test_keeps_a_waiting_sample_older_than_a_matched_one(self):
- """A sample of a new, older scene is not dropped by a late sample of the previous scene."""
- synchronizer, matched_samples = _synchronizer()
- # the last sample of the previous scene still waits for the system under test, while the
- # first sample of the next scene, recorded years earlier, is published already
- pending = _publish_sample(synchronizer, _PREVIOUS_SCENE, topics=["label", "label_meta_info"])
- messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"])
-
- synchronizer.add("prediction", _message(_PREVIOUS_SCENE))
- synchronizer.add("prediction", _message(_NEXT_SCENE[0]))
-
- assert len(matched_samples) == 2
- assert matched_samples[0][1:] == (pending["label"], pending["label_meta_info"])
- assert matched_samples[1][1:] == (messages["label"], messages["label_meta_info"])
-
- def test_gives_up_on_the_sample_waiting_the_longest(self):
- """Samples are dropped in the order they arrived, not by their stamp."""
- synchronizer, matched_samples = _synchronizer(queue_size=2)
-
- # a sample of the previous scene waits first, followed by two samples of the older scene
- for stamp in [_PREVIOUS_SCENE, *_NEXT_SCENE[:2]]:
- _publish_sample(synchronizer, stamp, topics=["label"])
-
- # the queue size is exceeded, so the sample that has been waiting the longest is given up
- # on, even though the samples kept for evaluation were recorded years before it
- assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[:2]
-
- _publish_sample(synchronizer, _NEXT_SCENE[2], topics=["label"])
-
- assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[1:]
- assert matched_samples == []
-
-
-class _FakeLogger:
- """Collect the messages the node logs instead of publishing them to ROS."""
-
- def __init__(self):
- """Start with an empty log."""
- self.messages: list = []
-
- def info(self, message: str, **kwargs) -> None:
- """Record a logged message, whatever its severity."""
- self.messages.append(message)
-
- debug = warn = error = info
-
-
-class _FakeBenchmarkHandler:
- """Stand in for the benchmark whose results the node aggregates and writes."""
-
- def __init__(self):
- """Start without finalized or written results."""
- self.finalized_complete = None
- self.written_results = None
-
- def finalize(self, complete: bool = True) -> dict:
- """Report results that are marked the way the node asked for."""
- self.finalized_complete = complete
- return {"num_samples": 2, "num_scenes": 1, "complete": complete, "aggregated_metrics": {}}
-
- def save_results(self, output_path: str, results: dict = None) -> str:
- """Keep the results instead of writing them to a file."""
- self.written_results = results
- return output_path
-
-
-class _RecordingBenchmarkHandler:
- """Stand in for the benchmark that records the samples the node evaluates."""
-
- def __init__(self):
- """Start without recorded samples."""
- self.recorded_samples: list = []
-
- def record_sample(self, sample_id: str, **messages) -> dict:
- """Record a sample the way the benchmark stores it, still without a scene."""
- entry = {"sample_id": sample_id, "scene_id": None, "metrics": {}}
- self.recorded_samples.append(entry)
- return entry
-
-
-def _evaluating_node(manual_playback: bool) -> SimpleNamespace:
- """Stub the node state that evaluating a sample reads, counting the attempts to continue the benchmark."""
- node = SimpleNamespace(
- manual_playback=manual_playback,
- input_topics=_TOPICS,
- benchmark_handler=_RecordingBenchmarkHandler(),
- num_evaluated_samples=0,
- scenes_awaiting_sample=deque(),
- samples_awaiting_scene=deque(),
- evaluation_timeout=60.0,
- evaluation_deadline=None,
- results_path="",
- visualization_publishers={},
- num_advances=0,
- get_logger=lambda logger=_FakeLogger(): logger,
- )
- node.advance_benchmark = lambda: setattr(node, "num_advances", node.num_advances + 1)
- return node
-
-
-def _sample_messages(stamp: tuple[int, int]) -> tuple:
- """Fake the synchronized input messages of one sample, in the order of the inputs."""
- return tuple(_message(stamp) for _ in _TOPICS)
-
-
-class TestEvaluateSample:
- """Tests evaluating a sample with the playback driven by the benchmark or by the user in RViz."""
-
- def test_benchmark_requests_the_next_samples_after_evaluating_one(self):
- """Without manual playback, the benchmark continues with the next samples on its own."""
- node = _evaluating_node(manual_playback=False)
-
- AutonomyBenchmarks.evaluate_sample(node, *_sample_messages(_NEXT_SCENE[0]))
-
- assert node.num_advances == 1
- # the dataset reports the scene of the sample with the response to the request
- assert list(node.samples_awaiting_scene) == node.benchmark_handler.recorded_samples
-
- def test_manual_playback_leaves_requesting_samples_to_the_user(self):
- """With manual playback, samples are evaluated as they arrive, without requesting further ones."""
- node = _evaluating_node(manual_playback=True)
-
- for stamp in _NEXT_SCENE:
- AutonomyBenchmarks.evaluate_sample(node, *_sample_messages(stamp))
-
- assert node.num_evaluated_samples == len(_NEXT_SCENE)
- assert node.num_advances == 0
- # the scenes are only reported to the playback panel, so no sample waits for one
- assert not node.samples_awaiting_scene
-
-
-def _node(num_evaluated_samples: int = 2, results_path: str = "/results/benchmark.json") -> SimpleNamespace:
- """Stub the node state that finalizing a benchmark reads, without initializing ROS."""
- return SimpleNamespace(
- benchmark="counting",
- benchmark_finished=False,
- benchmark_handler=_FakeBenchmarkHandler(),
- num_evaluated_samples=num_evaluated_samples,
- results_path=results_path,
- request_timer=SimpleNamespace(cancel=lambda: None),
- get_logger=lambda logger=_FakeLogger(): logger,
- )
-
-
-class TestFinalizeBenchmark:
- """Tests reporting the results of a benchmark that ran to its end or was interrupted."""
-
- def test_finished_benchmark_writes_complete_results(self):
- """A benchmark that evaluated all its samples reports complete results."""
- node = _node()
-
- AutonomyBenchmarks.finalize_benchmark(node)
-
- assert node.benchmark_handler.finalized_complete is True
- assert node.benchmark_handler.written_results["complete"] is True
-
- def test_interrupted_benchmark_writes_incomplete_results(self):
- """An evaluation stopped before its last sample, e.g. with Ctrl-C, still writes its results."""
- node = _node()
-
- AutonomyBenchmarks.finalize_benchmark(node, complete=False)
-
- assert node.benchmark_handler.finalized_complete is False
- assert node.benchmark_handler.written_results["complete"] is False
-
- def test_finished_benchmark_is_not_finalized_again_on_shutdown(self):
- """Shutting down after the last sample must not overwrite the results with incomplete ones."""
- node = _node()
- AutonomyBenchmarks.finalize_benchmark(node)
-
- AutonomyBenchmarks.finalize_benchmark(node, complete=False)
-
- assert node.benchmark_handler.finalized_complete is True
- assert node.benchmark_handler.written_results["complete"] is True
-
- def test_manual_playback_writes_its_results_when_stopped(self):
- """With manual playback, which runs no request timer, the results are written once the node is stopped."""
- node = _node()
- node.request_timer = None
-
- AutonomyBenchmarks.finalize_benchmark(node, complete=False)
-
- assert node.benchmark_handler.written_results["complete"] is False
-
- def test_interrupted_benchmark_without_samples_writes_nothing(self):
- """An evaluation interrupted before its first sample has no results to write."""
- node = _node(num_evaluated_samples=0)
-
- AutonomyBenchmarks.finalize_benchmark(node, complete=False)
-
- assert node.benchmark_handler.written_results is None
-
- def test_results_are_only_logged_without_a_results_path(self):
- """Without 'results_path' the interrupted results are logged instead of written."""
- node = _node(results_path="")
-
- AutonomyBenchmarks.finalize_benchmark(node, complete=False)
-
- assert node.benchmark_handler.finalized_complete is False
- assert node.benchmark_handler.written_results is None
diff --git a/autonomy_evaluation/README.md b/autonomy_evaluation/README.md
new file mode 100644
index 0000000..5ea37a4
--- /dev/null
+++ b/autonomy_evaluation/README.md
@@ -0,0 +1,98 @@
+# `autonomy_evaluation`
+
+Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments
+
+`autonomy_evaluation` is the part of the **Autonomy.Benchmarks** suite that turns the output of a system under test into
+metrics. It evaluates the samples that [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) replays
+against the labels of the dataset, and reports the metrics per scene and over all evaluated samples, as one part of the evidence
+an automated driving deployment is benchmarked on.
+
+## Nodes
+
+### `autonomy_evaluation`
+
+The node runs the evaluation selected by `evaluation`, either one of this package by its name or one implemented in another package as `:` (see [Adding more Evaluations](../docs/IMPLEMENTATION.md#adding-more-evaluations)). Each evaluation declares the topics it reads: its _inputs_ from the system under test, and, if it compares them with a reference, its _ground truth_. The node subscribes to all of them, each on its node-relative name, which the launch file remaps onto the topic given by the launch argument of the same name, e.g. `prediction:=/object_list/prediction`. A topic without such an argument is subscribed in the private namespace of the node. The node logs which topic it evaluates as input and which as ground truth.
+
+| Evaluation | Inputs | Ground truth |
+| --- | --- | --- |
+| `nuscenes_lidar_object_detection` | `prediction` (`perception_msgs/ObjectList`) | `label` (`perception_msgs/ObjectList`), `label_meta_info` (`autonomy_datasets_msgs/ObjectListMetaInfo`, subscribed next to `label` on `/meta_info`) |
+
+The messages of all topics that belong to the same sample are matched by their stamp and evaluated together. By default, only messages with exactly the same stamp are matched, as the dataset stamps all messages of a sample with its recording time and a system under test stamps its output with the stamp of the input it processed. Set `sync_tolerance` to match topics whose stamps differ, e.g. those a simulation publishes at slightly different times or at different rates. A message without a header, e.g. a `std_msgs/Bool`, is stamped with the time it is received at.
+
+With `sample_source` set to `dataset`, the default, the node requests the samples it evaluates from the dataset, using the `request_samples` service of [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets), which publishes them and responds once they have been published. By default one sample is requested at a time, so the dataset only publishes the next sample once the system under test has delivered its output for the current one and the node has evaluated it. Increase `samples_per_request` to publish samples in batches, set it to `0` to publish the whole dataset with a single request, or list the IDs of individual samples in `sample_ids` to evaluate only those.
+
+Once the dataset reports that all requested samples have been published, or its node has shut down after publishing its last sample, the per-sample metrics are aggregated over all evaluated samples (`aggregated_metrics`) and for the samples of each scene the dataset published them from (`scene_results`, matched with the samples via the `published_scene_ids` of the responses). The metrics are logged, and written to a JSON file if `results_path` is set. Samples that are not evaluated within `evaluation_timeout` seconds of being published, e.g. because the system under test skipped them, are left out.
+
+```bash
+ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
+ prediction:=/object_list/prediction \
+ label:=/object_list/lidar_01 \
+ request_samples:=/datasets/request_samples \
+ results_path:=/results/nuscenes_lidar_object_detection.json
+```
+
+With `sample_source` set to `external`, the node requests no samples itself and evaluates every sample that others publish, e.g. a closed-loop simulation, a live system, or the dataset stepped through with the [playback panel](https://github.com/thinking-cars/autonomy_datasets/blob/main/autonomy_datasets_rviz_plugins/README.md) in RViz, whose _Service_ field has to name the `request_samples` service of the dataset node (`/datasets/request_samples` by default). `samples_per_request`, `sample_ids` and `evaluation_timeout` then have no effect. As the node receives no responses of a dataset, it neither learns the scenes of the samples, so `scene_results` stays empty, nor when publishing has ended: the results of the evaluated samples are reported once the node is stopped, e.g. with Ctrl-C, and are marked as incomplete.
+
+```bash
+# look at the samples of a dataset one by one, stepping through them in RViz
+ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
+ prediction:=/object_list/prediction \
+ label:=/object_list/lidar_01 \
+ visualize:=true \
+ sample_source:=external
+
+# evaluate a closed-loop planner with an evaluation of another package, which reads inputs only
+ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
+ evaluation:=my_planner_evaluations.time_to_collision:TimeToCollision \
+ ego_data:=/ego_data \
+ objects:=/simulation/objects \
+ sample_source:=external \
+ sync_tolerance:=0.05 \
+ results_path:=/results/time_to_collision.json
+```
+
+```mermaid
+flowchart LR
+ NODE("autonomy_evaluation")
+ NODE o--o|~/request_samples| SC0:::hidden
+ classDef hidden display: none;
+```
+
+#### Service Clients
+
+| Service | Type | Description |
+| --- | --- | --- |
+| `~/request_samples` | `autonomy_datasets_msgs/srv/RequestSamples` | request samples from dataset |
+
+#### Parameters
+
+| Parameter | Type | Default | Description |
+| --- | --- | --- | --- |
+| `evaluation` | `string` | `nuscenes_lidar_object_detection` | name of an evaluation of this package, or ':' of an evaluation implemented in another package |
+| `visualize` | `bool` | `false` | publish the per-sample visualization of the evaluation for RViz, e.g. the true positives, false positives and false negatives of an object detection |
+| `sample_source` | `string` | `dataset` | 'dataset' requests the samples to evaluate from the dataset one after another; 'external' evaluates the samples published by others, e.g. by a simulation, a live system or the playback panel in RViz, and reports the results once the node is stopped |
+| `sync_tolerance` | `float` | `0.0` | seconds by which the stamps of the messages of a sample may differ; 0 only matches messages with exactly the same stamp, as the dataset and a system under test echoing its stamps publish them |
+| `samples_per_request` | `int` | `1` | number of samples to request from the dataset at a time; 0 requests all remaining samples at once, 1 evaluates every sample before the next one is published |
+| `sample_ids` | `string` | - | comma-separated IDs of the dataset samples to evaluate (e.g. '0,10,20'); if empty, all samples of the dataset are evaluated |
+| `evaluation_timeout` | `float` | `60.0` | seconds to wait for a published sample to be evaluated before continuing without it |
+| `results_path` | `string` | - | path of the JSON file the evaluation results are written to; results are only logged if empty |
+
+## Launch Files
+
+### [`autonomy_evaluation.launch.py`](launch/autonomy_evaluation.launch.py)
+
+| Argument | Default | Description |
+| --- | --- | --- |
+| `request_samples` | `"~/request_samples"` | service of the dataset node used to request the samples to evaluate |
+| `evaluation` | `"nuscenes_lidar_object_detection"` | evaluation to run |
+| `name` | `"autonomy_evaluation"` | node name |
+| `namespace` | `""` | node namespace |
+| `log_level` | `"info"` | ros logging level |
+| `use_sim_time` | `"true"` | use sim time |
+| `visualize` | `"false"` | publish the per-sample true positives, false positives and false negatives and open RViz |
+| `sample_source` | `"dataset"` | 'dataset' requests the samples to evaluate from the dataset one after another; 'external' evaluates the samples published by others, e.g. by a simulation or the playback panel in RViz |
+| `sync_tolerance` | `"0.0"` | seconds by which the stamps of the messages of a sample may differ (0 matches exact stamps only) |
+| `samples_per_request` | `"1"` | number of samples to request from the dataset at a time (0 requests all remaining samples at once) |
+| `sample_ids` | `""` | comma-separated IDs of the dataset samples to evaluate (all samples if empty) |
+| `evaluation_timeout` | `"60.0"` | seconds to wait for a published sample to be evaluated before continuing without it |
+| `results_path` | `""` | path of the JSON file the evaluation results are written to (results are only logged if empty) |
diff --git a/autonomy_evaluation/autonomy_evaluation/__init__.py b/autonomy_evaluation/autonomy_evaluation/__init__.py
new file mode 100644
index 0000000..de1b036
--- /dev/null
+++ b/autonomy_evaluation/autonomy_evaluation/__init__.py
@@ -0,0 +1,4 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Metrics-based evaluation of automated driving tasks."""
diff --git a/autonomy_benchmarks/autonomy_benchmarks/autonomy_benchmarks.py b/autonomy_evaluation/autonomy_evaluation/autonomy_evaluation.py
similarity index 52%
rename from autonomy_benchmarks/autonomy_benchmarks/autonomy_benchmarks.py
rename to autonomy_evaluation/autonomy_evaluation/autonomy_evaluation.py
index ff0b450..fe8a584 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/autonomy_benchmarks.py
+++ b/autonomy_evaluation/autonomy_evaluation/autonomy_evaluation.py
@@ -2,6 +2,7 @@
# SPDX-License-Identifier: Apache-2.0
import json
+import signal
import time
from collections import deque, OrderedDict
from functools import partial
@@ -10,6 +11,7 @@
import rclpy
import rclpy.exceptions
from autonomy_datasets_msgs.srv import RequestSamples
+from autonomy_evaluation.evaluations import load_evaluation
from rcl_interfaces.msg import FloatingPointRange, IntegerRange, ParameterDescriptor, SetParametersResult
from rclpy.clock import Clock, ClockType
from rclpy.executors import ExternalShutdownException
@@ -20,17 +22,30 @@
from rclpy.task import Future
from rclpy.timer import Timer
-# Interval in seconds at which the benchmark checks whether it can request further samples; the
-# benchmark also advances whenever a request is answered or a sample has been evaluated
+# Interval in seconds at which the evaluation checks whether it can request further samples; the
+# evaluation also advances whenever a request is answered or a sample has been evaluated
_REQUEST_TIMER_PERIOD_S = 0.5
# Interval in seconds at which waiting for the sample request service of the dataset is logged
_SERVICE_WAIT_LOG_INTERVAL_S = 10.0
+# Seconds for which the sample request service of a dataset that has been available must be gone
+# before the dataset node is considered to have shut down, which lets a response that the dataset
+# sent just before shutting down still arrive
+_DATASET_SHUTDOWN_GRACE_PERIOD_S = 2.0
+
# Number of samples whose messages are kept while they wait for the messages of their remaining
-# benchmark inputs
+# evaluation inputs
_SYNCHRONIZER_QUEUE_SIZE = 10
+# Values of the 'sample_source' parameter: the node requests the samples to evaluate from the
+# dataset, or evaluates whatever samples reach it, e.g. from a simulation or the RViz playback panel
+SAMPLE_SOURCE_DATASET = "dataset"
+SAMPLE_SOURCE_EXTERNAL = "external"
+
+# Stamp of a message as (seconds, nanoseconds)
+Stamp = tuple[int, int]
+
def parse_sample_ids(sample_ids: str) -> list[int]:
"""Parses the IDs of the dataset samples to evaluate
@@ -48,87 +63,168 @@ def parse_sample_ids(sample_ids: str) -> list[int]:
class SampleSynchronizer:
- """Matches the messages of the benchmark inputs that belong to the same dataset sample.
-
- The dataset stamps all messages of a sample with the recording time of that sample, so the
- messages of a sample are matched by their exact header stamp. Messages of a sample that never
- completes are dropped in the order they arrived, never by comparing their stamps: the dataset
- replays one scene after the other, and a scene can have been recorded days before the scene
- played before it, so the stamp of a message says nothing about how recently it was received.
- (``message_filters.TimeSynchronizer`` drops by stamp instead, and therefore discards the
- messages of a scene that starts before the end of the preceding one.)
+ """Matches the messages of the evaluation inputs that belong to the same sample.
+
+ Messages are matched by their stamp. By default, only messages with exactly the same stamp are
+ matched: the dataset stamps all messages of a sample with the recording time of that sample,
+ and a system under test stamps its output with the stamp of the input it processed. With a
+ tolerance, a message joins the waiting sample whose stamp is closest to its own, as long as
+ both stamps differ by no more than the tolerance and the sample still misses a message of its
+ input. This matches topics that a simulation or a live system publishes at slightly different
+ times or at different rates, of which the first message within the tolerance is matched.
+
+ Messages of a sample that never completes are dropped in the order they arrived, never by
+ comparing their stamps: the dataset replays one scene after the other, and a scene can have
+ been recorded days before the scene played before it, so the stamp of a message says nothing
+ about how recently it was received. (``message_filters.TimeSynchronizer`` drops by stamp
+ instead, and therefore discards the messages of a scene that starts before the end of the
+ preceding one.)
Messages are added from subscription callbacks, which the node executor runs one after
another, so no locking is needed.
"""
- def __init__(self, topics: Sequence[str], callback: Callable[..., None], queue_size: int = _SYNCHRONIZER_QUEUE_SIZE):
+ def __init__(
+ self,
+ topics: Sequence[str],
+ callback: Callable[[Stamp, dict[str, Any]], None],
+ queue_size: int = _SYNCHRONIZER_QUEUE_SIZE,
+ tolerance: float = 0.0,
+ ):
"""Constructor
Args:
topics (Sequence[str]): input names to match, in the order their messages are passed
to the callback
- callback (Callable[..., None]): called with the messages of every completed sample
+ callback (Callable[[Stamp, dict[str, Any]], None]): called for every completed sample
+ with the stamp of its message of the first input and its messages by input name
queue_size (int, optional): number of samples to keep while they wait for the messages
of their remaining inputs
+ tolerance (float, optional): seconds by which the stamps of the messages of a sample
+ may differ; 0 only matches messages with exactly the same stamp
"""
self.topics = list(topics)
self.callback = callback
self.queue_size = queue_size
- # messages of the samples that are still missing inputs, by header stamp and in the order
- # the samples were first received on any input
- self.incomplete_samples: OrderedDict[tuple[int, int], dict[str, Any]] = OrderedDict()
+ self.tolerance_ns = round(tolerance * 1e9)
+ # stamped messages of the samples that are still missing inputs, by the stamp of the
+ # message that opened the sample and in the order the samples were opened
+ self.incomplete_samples: OrderedDict[Stamp, dict[str, tuple[Stamp, Any]]] = OrderedDict()
- def add(self, topic: str, message: Any):
+ def add(self, topic: str, message: Any, stamp: Stamp):
"""Adds a received message and reports the sample it completes to the callback
Args:
topic (str): input name the message was received on
- message (Any): received message, stamped with the time of its sample
+ message (Any): received message
+ stamp (Stamp): stamp of the message, i.e. of the sample it belongs to
"""
- stamp = (message.header.stamp.sec, message.header.stamp.nanosec)
- messages = self.incomplete_samples.setdefault(stamp, {})
- messages[topic] = message
+ sample_stamp = self.find_sample(topic, stamp)
+ messages = self.incomplete_samples.setdefault(sample_stamp, {})
+ messages[topic] = (stamp, message)
if len(messages) == len(self.topics):
- del self.incomplete_samples[stamp]
- self.callback(*(messages[topic] for topic in self.topics))
+ del self.incomplete_samples[sample_stamp]
+ self.callback(messages[self.topics[0]][0], {topic: messages[topic][1] for topic in self.topics})
return
# give up on the sample that has been waiting for its remaining inputs the longest
while len(self.incomplete_samples) > self.queue_size:
self.incomplete_samples.popitem(last=False)
+ def find_sample(self, topic: str, stamp: Stamp) -> Stamp:
+ """Finds the waiting sample a message belongs to
+
+ A message joins the sample of exactly its stamp, replacing a message the sample already
+ holds for its input, or else the waiting sample within the tolerance whose stamp is
+ closest to its own and that still misses a message of its input.
+
+ Args:
+ topic (str): input name the message was received on
+ stamp (Stamp): stamp of the message
+
+ Returns:
+ Stamp: stamp of the sample the message belongs to, which is its own stamp if it opens
+ a new sample
+ """
+ if stamp in self.incomplete_samples or not self.tolerance_ns:
+ return stamp
+ nanoseconds = _to_nanoseconds(stamp)
+ candidates = [
+ (abs(_to_nanoseconds(sample_stamp) - nanoseconds), sample_stamp)
+ for sample_stamp, messages in self.incomplete_samples.items()
+ if topic not in messages
+ ]
+ candidates = [candidate for candidate in candidates if candidate[0] <= self.tolerance_ns]
+ return min(candidates)[1] if candidates else stamp
+
+
+def _to_nanoseconds(stamp: Stamp) -> int:
+ """Converts a stamp into nanoseconds
-class AutonomyBenchmarks(Node):
- """ROS 2 node for benchmarking automated driving tasks."""
+ Args:
+ stamp (Stamp): stamp as (seconds, nanoseconds)
+
+ Returns:
+ int: nanoseconds
+ """
+ return stamp[0] * 1_000_000_000 + stamp[1]
+
+
+class AutonomyEvaluation(Node):
+ """ROS 2 node evaluating automated driving tasks, generating metrics-based evidence for benchmarking them."""
def __init__(self):
"""Constructor"""
- super().__init__("autonomy_benchmarks")
+ super().__init__("autonomy_evaluation")
self.auto_reconfigurable_params: list[str] = []
- self.benchmark = self.declare_and_load_parameter(
- name="benchmark",
+ self.evaluation = self.declare_and_load_parameter(
+ name="evaluation",
param_type=rclpy.Parameter.Type.STRING,
- description="benchmark name",
+ description="name of an evaluation of this package, or ':' of an evaluation implemented in "
+ "another package",
default="nuscenes_lidar_object_detection",
+ add_to_auto_reconfigurable_params=False,
+ read_only=True,
)
self.visualize = self.declare_and_load_parameter(
name="visualize",
param_type=rclpy.Parameter.Type.BOOL,
- description="publish the per-sample true positives, false positives and false negatives for RViz",
+ description="publish the per-sample visualization of the evaluation for RViz, e.g. the true positives, false "
+ "positives and false negatives of an object detection",
default=False,
)
- self.manual_playback = self.declare_and_load_parameter(
- name="manual_playback",
- param_type=rclpy.Parameter.Type.BOOL,
- description="leave requesting the samples of the dataset to the user",
- default=False,
+ self.sample_source = self.declare_and_load_parameter(
+ name="sample_source",
+ param_type=rclpy.Parameter.Type.STRING,
+ description="'dataset' requests the samples to evaluate from the dataset one after another; 'external' "
+ "evaluates the samples published by others, e.g. by a simulation, a live system or the playback panel in RViz, "
+ "and reports the results once the node is stopped",
+ default="dataset",
+ add_to_auto_reconfigurable_params=False,
+ read_only=True,
+ )
+ if self.sample_source not in (SAMPLE_SOURCE_DATASET, SAMPLE_SOURCE_EXTERNAL):
+ self.get_logger().fatal(
+ f"Parameter 'sample_source' is neither '{SAMPLE_SOURCE_DATASET}' nor '{SAMPLE_SOURCE_EXTERNAL}': "
+ f"'{self.sample_source}'"
+ )
+ raise SystemExit(1)
+ self.requests_samples = self.sample_source == SAMPLE_SOURCE_DATASET
+
+ self.sync_tolerance = self.declare_and_load_parameter(
+ name="sync_tolerance",
+ param_type=rclpy.Parameter.Type.DOUBLE,
+ description="seconds by which the stamps of the messages of a sample may differ; 0 only matches messages "
+ "with exactly the same stamp, as the dataset and a system under test echoing its stamps publish them",
+ default=0.0,
add_to_auto_reconfigurable_params=False,
read_only=True,
+ from_value=0.0,
+ to_value=60.0,
)
self.samples_per_request = self.declare_and_load_parameter(
@@ -168,7 +264,7 @@ def __init__(self):
self.results_path = self.declare_and_load_parameter(
name="results_path",
param_type=rclpy.Parameter.Type.STRING,
- description="path of the JSON file the benchmark results are written to; results are only logged if empty",
+ description="path of the JSON file the evaluation results are written to; results are only logged if empty",
default="",
)
@@ -274,30 +370,28 @@ def setup(self):
self.data_subscriptions: dict[str, Subscription] = {}
- # get handler for specified benchmark
- benchmark_handler = None
- if self.benchmark == "nuscenes_lidar_object_detection":
- from autonomy_benchmarks.benchmarks.lidar_object_detection.NuscenesLidarObjectDetection import (
- NuscenesLidarObjectDetection,
- )
-
- benchmark_handler = NuscenesLidarObjectDetection()
- else:
- self.get_logger().fatal(f"Benchmark '{self.benchmark}' not recognized, exiting")
+ # load the selected evaluation and the topics it reads
+ try:
+ evaluation_handler = load_evaluation(self.evaluation)
+ inputs = evaluation_handler.all_inputs()
+ except ValueError as exception:
+ self.get_logger().fatal(f"{exception}, exiting")
raise SystemExit(1)
- # create subscriptions for benchmark data inputs, whose messages are matched into the
- # samples to evaluate by their header stamp
+ # create subscriptions for the topics of the evaluation, whose messages are matched into
+ # the samples to evaluate by their stamp
self.message_synchronizer = SampleSynchronizer(
- topics=list(benchmark_handler.required_inputs()),
+ topics=list(inputs),
callback=self.evaluate_sample,
queue_size=_SYNCHRONIZER_QUEUE_SIZE,
+ tolerance=self.sync_tolerance,
)
- for msg_topic, msg_type in benchmark_handler.required_inputs().items():
- self.data_subscriptions[msg_topic] = self.create_subscription(
+ derived_topics = evaluation_handler.derived_topics()
+ for name, msg_type in inputs.items():
+ self.data_subscriptions[name] = self.create_subscription(
msg_type,
- msg_topic,
- partial(self.message_synchronizer.add, msg_topic),
+ self.input_topic(name, derived_topics),
+ partial(self.receive_message, name),
qos_profile=QoSProfile(
reliability=ReliabilityPolicy.RELIABLE,
durability=DurabilityPolicy.VOLATILE,
@@ -305,11 +399,27 @@ def setup(self):
depth=10,
),
)
+ ground_truth = evaluation_handler.required_ground_truth()
+ for role, names in (("input", [name for name in inputs if name not in ground_truth]), ("ground truth", ground_truth)):
+ if names:
+ topics = ", ".join(f"'{name}' from '{self.data_subscriptions[name].topic_name}'" for name in names)
+ self.get_logger().info(f"Evaluating {role} {topics}")
+
+ # Messages without a header are stamped when they are received, which the messages of
+ # other inputs can only match within a tolerance
+ unstamped_inputs = [name for name, msg_type in inputs.items() if "header" not in msg_type.get_fields_and_field_types()]
+ if unstamped_inputs:
+ self.get_logger().info(f"Stamping the messages of {unstamped_inputs} on reception, as they carry no header")
+ if len(inputs) > 1 and not self.sync_tolerance:
+ self.get_logger().warn(
+ f"Messages of {unstamped_inputs} can only be matched with those of other inputs if they are "
+ "received at exactly their stamp, set 'sync_tolerance' to match them within a tolerance"
+ )
- # create publishers visualizing the benchmark's per-sample matching outcome
+ # create publishers visualizing the evaluation's per-sample matching outcome
self.visualization_publishers: dict[str, Publisher] = {}
if self.visualize:
- for msg_topic, msg_type in benchmark_handler.visualization_outputs().items():
+ for msg_topic, msg_type in evaluation_handler.visualization_outputs().items():
self.visualization_publishers[msg_topic] = self.create_publisher(
msg_type,
f"~/{msg_topic}",
@@ -320,11 +430,9 @@ def setup(self):
depth=10,
),
)
- self.get_logger().info(f"Visualizing benchmark results on: {sorted(self.visualization_publishers)}")
+ self.get_logger().info(f"Visualizing evaluation results on: {sorted(self.visualization_publishers)}")
- # store handler and ordered topic list for use in evaluate_sample
- self.benchmark_handler = benchmark_handler
- self.input_topics: list = list(benchmark_handler.required_inputs().keys())
+ self.evaluation_handler = evaluation_handler
self.published_sample_ids: list[int] = []
# A sample may already be evaluated before the dataset node answers the request that
@@ -336,58 +444,126 @@ def setup(self):
self.pending_request: Optional[Future] = None
self.evaluation_deadline: Optional[float] = None
self.publishing_finished = False
- self.benchmark_finished = False
+ # whether the sample request service of the dataset has been available, and since when it
+ # is gone, to tell a dataset node that has not started yet from one that has shut down
+ self.dataset_available = False
+ self.dataset_unavailable_since: Optional[float] = None
+ self.evaluation_finished = False
self.num_evaluated_samples = 0
- if self.manual_playback:
+ # With an external sample source, the samples are evaluated as they arrive. Whoever
+ # publishes them receives the responses of the dataset, if any, so the evaluation neither
+ # learns the scenes of the samples nor when publishing has ended, and reports its results
+ # once the node is stopped.
+ if not self.requests_samples:
self.request_timer: Optional[Timer] = None
if self.requested_sample_ids:
- self.get_logger().warn("Parameter 'sample_ids' is ignored, as the samples are requested manually")
- self.get_logger().info("Evaluating the samples requested manually, stop the node to report the results")
+ self.get_logger().warn(
+ f"Parameter 'sample_ids' is ignored, as only sample source '{SAMPLE_SOURCE_DATASET}' requests samples"
+ )
+ self.get_logger().info("Evaluating the samples published by others, stop the node to report the results")
return
self.sample_request_client = self.create_client(RequestSamples, "~/request_samples")
- # driven by a steady clock, so that the benchmark also advances while the simulation clock
+ # name of the service with remappings applied, for the log messages
+ self.sample_request_service = self.resolve_service_name(self.sample_request_client.srv_name)
+ # driven by a steady clock, so that the evaluation also advances while the simulation clock
# of the dataset stands still, i.e. while no sample is being published
self.request_timer = self.create_timer(
_REQUEST_TIMER_PERIOD_S,
- self.advance_benchmark,
+ self.advance_evaluation,
clock=Clock(clock_type=ClockType.STEADY_TIME),
)
- self.get_logger().info(f"Requesting samples to evaluate from '{self.sample_request_client.srv_name}'")
+ self.get_logger().info(f"Requesting samples to evaluate from '{self.sample_request_service}'")
- def advance_benchmark(self):
- """Requests the next samples to evaluate, or finalizes the benchmark once all were published
+ def input_topic(self, name: str, derived_topics: dict[str, tuple[str, str]]) -> str:
+ """Determines the topic an input of the evaluation is subscribed on
+
+ An input is subscribed on its node-relative name, which is remapped onto the topic of the
+ system under test or of the ground truth. An input published next to another one follows
+ the topic of that input, unless it is remapped itself.
+
+ Args:
+ name (str): input name
+ derived_topics (dict[str, tuple[str, str]]): inputs published next to another input,
+ as input name -> (name of the other input, suffix of its topic)
+
+ Returns:
+ str: topic to subscribe on
+ """
+ if name not in derived_topics or self.resolve_topic_name(name) != self.resolve_topic_name(name, only_expand=True):
+ return name
+ source, suffix = derived_topics[name]
+ return self.resolve_topic_name(source) + suffix
+
+ def advance_evaluation(self):
+ """Requests the next samples to evaluate, or finalizes the evaluation once all were published
Called periodically as well as whenever a request has been answered or a sample has been
- evaluated, and does nothing while the benchmark is waiting for one of those.
+ evaluated, and does nothing while the evaluation is waiting for one of those. The service
+ of the dataset is only needed to request further samples, so the evaluation finishes even
+ after the dataset node has shut down.
"""
- if self.benchmark_finished or self.pending_request is not None:
+ if self.evaluation_finished:
+ return
+ self.track_dataset_availability()
+ if self.pending_request is not None or self.awaiting_evaluations():
return
- if not self.sample_request_client.service_is_ready():
+ if self.publishing_finished:
+ self.finalize_evaluation()
+ self.shutdown()
+ return
+ if not self.dataset_available:
self.get_logger().warn(
- f"Waiting for service '{self.sample_request_client.srv_name}' to request samples of the dataset...",
+ f"Waiting for service '{self.sample_request_service}' to request samples of the dataset...",
throttle_duration_sec=_SERVICE_WAIT_LOG_INTERVAL_S,
)
return
- if self.awaiting_evaluations():
+ if self.dataset_unavailable_since is None:
+ self.request_samples()
+
+ def track_dataset_availability(self):
+ """Ends publishing once the dataset node has shut down
+
+ The dataset node shuts down after it has published its last sample, without necessarily
+ reporting the end of the dataset: a request that publishes the last sample is answered
+ before the dataset notices that no sample follows. Once the service of a dataset that has
+ been available is gone for longer than a grace period, no further samples can be
+ published, so the samples published so far are the ones to evaluate. A request that the
+ dataset did not answer before it went away is given up on.
+ """
+ if self.sample_request_client.service_is_ready():
+ self.dataset_available = True
+ self.dataset_unavailable_since = None
+ return
+ if not self.dataset_available or self.publishing_finished:
return
- if self.publishing_finished:
- self.finalize_benchmark()
- self.shutdown()
+ now = time.monotonic()
+ if self.dataset_unavailable_since is None:
+ self.dataset_unavailable_since = now
+ if now - self.dataset_unavailable_since < _DATASET_SHUTDOWN_GRACE_PERIOD_S:
return
- self.request_samples()
+
+ if self.pending_request is not None:
+ self.sample_request_client.remove_pending_request(self.pending_request)
+ self.pending_request = None
+ self.get_logger().warn("The dataset node shut down before answering the last request for samples")
+ self.get_logger().info(
+ f"The dataset node providing '{self.sample_request_service}' has shut down, finishing with the "
+ "samples it published"
+ )
+ self.publishing_finished = True
def awaiting_evaluations(self) -> bool:
"""Reports whether published samples are still waiting to be evaluated
- A sample is evaluated once all benchmark inputs have been received for it, which happens
+ A sample is evaluated once all evaluation inputs have been received for it, which happens
once the system under test has processed the sample the dataset published. Samples that
are not evaluated within 'evaluation_timeout' seconds are given up on, so that a system
- under test which skips samples does not stall the benchmark.
+ under test which skips samples does not stall the evaluation.
Returns:
- bool: whether the benchmark waits for published samples to be evaluated
+ bool: whether the evaluation waits for published samples to be evaluated
"""
outstanding_evaluations = len(self.scenes_awaiting_sample) - len(self.samples_awaiting_scene)
if outstanding_evaluations <= 0:
@@ -425,7 +601,7 @@ def request_samples(self):
self.pending_request.add_done_callback(self.samples_published_callback)
def samples_published_callback(self, future: Future):
- """Records the samples the dataset node has published and continues the benchmark
+ """Records the samples the dataset node has published and continues the evaluation
Args:
future (Future): future of the request, holding the response of the dataset node
@@ -436,7 +612,7 @@ def samples_published_callback(self, future: Future):
except Exception as exception:
self.get_logger().error(f"Requesting samples of the dataset failed: {exception}")
self.publishing_finished = True
- self.advance_benchmark()
+ self.advance_evaluation()
return
published_sample_ids = [int(sample_id) for sample_id in response.published_sample_ids]
@@ -455,12 +631,12 @@ def samples_published_callback(self, future: Future):
self.publishing_finished = (
response.end_of_dataset or not response.success or bool(self.requested_sample_ids) or self.samples_per_request <= 0
)
- self.advance_benchmark()
+ self.advance_evaluation()
def assign_scenes(self, published_scene_ids: Sequence[str]):
"""Attributes the scenes of published samples to the samples that are evaluated for them
- The benchmark aggregates the metrics of the samples of a scene, so every evaluated sample
+ The evaluation aggregates the metrics of the samples of a scene, so every evaluated sample
needs the scene the dataset published it from. Samples are evaluated in the order the
dataset published them, so both are matched in that order; a scene whose sample has not
been evaluated yet waits for it, and vice versa.
@@ -475,12 +651,12 @@ def assign_scenes(self, published_scene_ids: Sequence[str]):
else:
self.scenes_awaiting_sample.append(str(scene_id))
- def finalize_benchmark(self, complete: bool = True):
- """Aggregates the metrics of the evaluated samples per scene and over the whole benchmark
+ def finalize_evaluation(self, complete: bool = True):
+ """Aggregates the metrics of the evaluated samples per scene and over the whole evaluation
The results hold the metrics of every single sample, of the samples of each scene, and of
- all evaluated samples, of which the metrics over all samples are logged. Calling this on a
- benchmark that has already been finalized does nothing, so that a benchmark which ran to
+ all evaluated samples, of which the metrics over all samples are logged. Calling this on an
+ evaluation that has already been finalized does nothing, so that an evaluation which ran to
its end is not finalized a second time when the node shuts down.
Args:
@@ -488,31 +664,31 @@ def finalize_benchmark(self, complete: bool = True):
results of an evaluation that was interrupted before its last sample, e.g. with
Ctrl-C, are marked as incomplete via '"complete": false'
"""
- if self.benchmark_finished:
+ if self.evaluation_finished:
return
- self.benchmark_finished = True
+ self.evaluation_finished = True
if self.request_timer is not None:
self.request_timer.cancel()
if not self.num_evaluated_samples:
- self.get_logger().warn(f"Benchmark '{self.benchmark}' evaluated no sample, no metrics are aggregated")
+ self.get_logger().warn(f"Evaluation '{self.evaluation}' evaluated no sample, no metrics are aggregated")
return
- results = self.benchmark_handler.finalize(complete=complete)
+ results = self.evaluation_handler.finalize(complete=complete)
aggregated_metrics = json.dumps(results["aggregated_metrics"], indent=2, default=str)
evaluated_samples = f"{results['num_samples']} evaluated sample(s) of {results['num_scenes']} scene(s)"
if complete:
- self.get_logger().info(f"Benchmark '{self.benchmark}' finished after {evaluated_samples}.")
+ self.get_logger().info(f"Evaluation '{self.evaluation}' finished after {evaluated_samples}.")
else:
self.get_logger().warn(
- f"Benchmark '{self.benchmark}' was interrupted after {evaluated_samples}, "
+ f"Evaluation '{self.evaluation}' was interrupted after {evaluated_samples}, "
"its results are marked as incomplete."
)
if self.results_path:
- reported_results = "benchmark results" if complete else "incomplete benchmark results"
+ reported_results = "evaluation results" if complete else "incomplete evaluation results"
try:
- results_path = self.benchmark_handler.save_results(self.results_path, results=results)
+ results_path = self.evaluation_handler.save_results(self.results_path, results=results)
self.get_logger().info(f"Wrote {reported_results} to '{results_path}'")
except OSError as exception:
self.get_logger().error(f"Failed to write {reported_results} to '{self.results_path}': {exception}")
@@ -520,50 +696,55 @@ def finalize_benchmark(self, complete: bool = True):
self.get_logger().info(f"Aggregated dataset metrics:\n{aggregated_metrics}")
def shutdown(self):
- """Stops the node, as there is nothing left to evaluate once the benchmark has finished
+ """Stops the node, as there is nothing left to evaluate once the evaluation has finished
Shutting down the ROS context ends the spinning of the node in 'main', which lets the
- process exit with the benchmark results reported.
+ process exit with the evaluation results reported.
"""
self.get_logger().info("Nothing left to evaluate, shutting down")
rclpy.try_shutdown()
- def evaluate_sample(self, *args):
- """Callback to evaluate a single sample when all required input messages have been received.
+ def receive_message(self, name: str, message: Any):
+ """Stamps a received message of an input and hands it to the matching of the samples
- The positional *args* are the synchronized ROS messages in the same order
- as the input names (dict keys) returned by
- ``benchmark_handler.required_inputs()``. Those keys match the
- ``compute_sample_metrics`` parameter names, so the raw ``ObjectList``
- messages are forwarded to ``benchmark_handler.record_sample`` by keyword;
- the benchmark extracts the fields it needs inside
- ``compute_sample_metrics``.
+ A message is stamped with the stamp of its header, which the dataset sets to the
+ recording time of its sample. A message without a header is stamped with the time it is
+ received at.
- Samples are identified by the ROS header stamp of their messages, which is the stamp the
- dataset recorded them with, and are attributed to the scene the dataset reported for them,
- which can arrive after the sample has been evaluated.
+ Args:
+ name (str): input name the message was received on
+ message (Any): received message
"""
- self.get_logger().debug("Received synchronized input messages, evaluating sample...")
+ header = getattr(message, "header", None)
+ if header is not None:
+ stamp = (header.stamp.sec, header.stamp.nanosec)
+ else:
+ stamp = self.get_clock().now().seconds_nanoseconds()
+ self.message_synchronizer.add(name, message, stamp)
- if len(args) != len(self.input_topics):
- self.get_logger().error(f"Expected {len(self.input_topics)} messages but received {len(args)}; skipping sample.")
- return
+ def evaluate_sample(self, stamp: Stamp, messages: dict[str, Any]):
+ """Evaluates a single sample once the messages of all its inputs have been received
- # Map each synchronized message to its required-input name (which matches
- # the compute_sample_metrics parameter names).
- messages = dict(zip(self.input_topics, args))
+ The messages are passed to the evaluation by the names of their inputs, which match the
+ parameters of its ``compute_sample_metrics``.
- # Use the ROS header stamp of the first message as sample ID.
- stamp = args[0].header.stamp
- sample_id = f"{stamp.sec}.{stamp.nanosec:09d}"
- self.get_logger().debug(f"Sample ID: '{sample_id}'")
+ Samples are identified by the stamp of their message of the first input, which is the
+ stamp the dataset recorded them with, and are attributed to the scene the dataset
+ reported for them, which can arrive after the sample has been evaluated.
+
+ Args:
+ stamp (Stamp): stamp of the sample
+ messages (dict[str, Any]): messages of the sample by input name
+ """
+ sample_id = f"{stamp[0]}.{stamp[1]:09d}"
+ self.get_logger().debug(f"Evaluating sample '{sample_id}'")
- result = self.benchmark_handler.record_sample(sample_id=sample_id, **messages)
+ result = self.evaluation_handler.record_sample(sample_id=sample_id, **messages)
self.num_evaluated_samples += 1
# attribute the sample to the scene the dataset published it from, which the dataset may
- # only report after the sample has been evaluated; with the playback controlled manually,
- # the scene is only reported to the playback panel
- if not self.manual_playback:
+ # only report after the sample has been evaluated; with an external sample source, the
+ # scene is only reported to whoever requested the sample, if at all
+ if self.requests_samples:
if self.scenes_awaiting_sample:
result["scene_id"] = self.scenes_awaiting_sample.popleft()
else:
@@ -574,33 +755,36 @@ def evaluate_sample(self, *args):
if not self.results_path:
self.get_logger().debug(f"Sample '{sample_id}' result: {result}")
- # publish the sample's matching outcome for inspection in RViz
+ # publish the sample's visualization for inspection in RViz
if self.visualization_publishers:
- visualization = self.benchmark_handler.visualize_sample(sample_id=sample_id, **messages)
+ visualization = self.evaluation_handler.visualize_sample(sample_id=sample_id, **messages)
for msg_topic, publisher in self.visualization_publishers.items():
publisher.publish(visualization[msg_topic])
# request the next samples, or aggregate the dataset metrics if this was the last one; with
- # the playback controlled manually, the user requests the next samples instead
- if not self.manual_playback:
- self.advance_benchmark()
+ # an external sample source, others publish the next samples instead
+ if self.requests_samples:
+ self.advance_evaluation()
def main():
"""Initializes ROS, runs the node event loop, and performs shutdown cleanup."""
rclpy.init()
- node = AutonomyBenchmarks()
+ node = AutonomyEvaluation()
try:
rclpy.spin(node)
except (KeyboardInterrupt, ExternalShutdownException):
pass
finally:
# An evaluation that is stopped before its last sample, e.g. with Ctrl-C, still reports the
- # samples it did evaluate, marked as incomplete results. A benchmark that ran to its end
- # has been finalized already and is left untouched.
+ # samples it did evaluate, marked as incomplete results. An evaluation that ran to its end
+ # has been finalized already and is left untouched. Ctrl-C reaches the node from the
+ # terminal and once more from the launch system, so further interrupts are ignored while
+ # the results are written, which they would otherwise abort.
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
try:
- node.finalize_benchmark(complete=False)
+ node.finalize_evaluation(complete=False)
finally:
node.destroy_node()
rclpy.try_shutdown()
diff --git a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/AutonomyBenchmark.py b/autonomy_evaluation/autonomy_evaluation/evaluations/Evaluation.py
similarity index 52%
rename from autonomy_benchmarks/autonomy_benchmarks/benchmarks/AutonomyBenchmark.py
rename to autonomy_evaluation/autonomy_evaluation/evaluations/Evaluation.py
index 2f405a7..8e54991 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/AutonomyBenchmark.py
+++ b/autonomy_evaluation/autonomy_evaluation/evaluations/Evaluation.py
@@ -1,11 +1,14 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Abstract base class for all AutonomyHub benchmarks.
-
-Each benchmark defines how to compute per-sample and aggregated metrics for a
-specific perception task (e.g. 2-D / 3-D object detection). Concrete
-subclasses must override the abstract methods.
+"""Abstract base class for all evaluations of automated driving tasks.
+
+Each evaluation defines which topics it reads and how to compute per-sample and
+aggregated metrics from their messages. An evaluation may read the topics of a
+system under test only, e.g. the ego state and the surrounding objects of a
+closed-loop planner to compute its time to collision, or compare them with
+ground-truth topics, e.g. the predictions of a perception algorithm with the
+labels of a dataset. Concrete subclasses must override the abstract methods.
"""
from __future__ import annotations
@@ -13,50 +16,92 @@
import json
import os
from abc import ABC, abstractmethod
-from typing import Any, Dict, List, Optional
+from typing import Any, Dict, List, Optional, Tuple
-class AutonomyBenchmark(ABC):
- """Meta-class (abstract base class) for AutonomyHub benchmarks.
+class Evaluation(ABC):
+ """Meta-class (abstract base class) for the evaluations of automated driving tasks.
- A benchmark is responsible for:
- * extracting the model input from a dataset sample,
- * computing per-sample metrics from a prediction and the ground-truth label,
+ An evaluation is responsible for:
+ * declaring the topics it reads, split into the inputs from the system
+ under test and, if it compares them with a reference, the ground truth,
+ * computing per-sample metrics from the messages of these topics,
* aggregating per-sample metrics into scene-level and dataset-level metrics,
* persisting results to JSON.
+
+ The messages of all topics that belong to the same sample are passed to
+ :meth:`compute_sample_metrics` as keyword arguments named after the topics
+ (see :meth:`all_inputs`), so an implementation names its parameters like
+ its topics.
"""
def __init__(self, name: str, description: str = "") -> None:
- """Initialize a benchmark definition and empty result store."""
+ """Initialize an evaluation definition and empty result store."""
self.name: str = name
self.description: str = description
self._sample_results: List[Dict[str, Any]] = []
# ------------------------------------------------------------------
- # Abstract interface – must be implemented by every concrete benchmark
+ # Abstract interface – must be implemented by every concrete evaluation
# ------------------------------------------------------------------
@abstractmethod
def required_inputs(self) -> Dict[str, Any]:
- """Define expected input ROS message types.
+ """Define the topics of the system under test that are evaluated.
+
+ These are the outputs of the system under test, e.g. the predictions
+ of a perception algorithm, and, for an evaluation that needs no
+ reference, the topics of the scenario it runs in, e.g. the ego state
+ and the surrounding objects a closed-loop planner is evaluated on.
+
+ Returns
+ -------
+ A dictionary mapping input names to their ROS message types. The
+ names are the keyword arguments of :meth:`compute_sample_metrics`
+ and the node-relative topics the node subscribes to.
+ """
+
+ def required_ground_truth(self) -> Dict[str, Any]:
+ """Define the ground-truth topics the inputs are compared with.
+
+ Ground truth is a reference the output of the system under test is
+ measured against, e.g. the labels of a dataset. An evaluation that
+ computes its metrics from the inputs alone needs none.
Returns
-------
- A dictionary mapping topic names to their expected message types.
+ A dictionary mapping ground-truth names to their ROS message types,
+ empty by default. The names are used like those of
+ :meth:`required_inputs`.
"""
+ return {}
+
+ def derived_topics(self) -> Dict[str, Tuple[str, str]]:
+ """Define the inputs that are published next to the topic of another input.
+
+ A topic that is published next to another one, e.g. the meta
+ information of an object list on ``/meta_info``,
+ follows that topic instead of having to be configured on its own.
+
+ Returns
+ -------
+ A dictionary mapping the name of such an input to the name of the
+ input it is published next to and the suffix appended to that
+ input's topic, empty by default.
+ """
+ return {}
@abstractmethod
- def compute_sample_metrics(self, prediction: Any, label: Any, sample_id: Optional[str] = None) -> Dict[str, Any]:
- """Compute metrics for a single (prediction, label) pair.
+ def compute_sample_metrics(self, sample_id: Optional[str] = None, **messages: Any) -> Dict[str, Any]:
+ """Compute metrics for a single sample.
Parameters
----------
- prediction:
- The model's output for one sample.
- label:
- The ground-truth for the same sample.
sample_id:
An optional identifier for the sample.
+ messages:
+ One message per topic of :meth:`all_inputs`, passed by the name
+ of its topic.
Returns
-------
@@ -79,31 +124,62 @@ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Di
mean-AP, precision, recall).
"""
+ # ------------------------------------------------------------------
+ # Topics
+ # ------------------------------------------------------------------
+
+ def all_inputs(self) -> Dict[str, Any]:
+ """Combine the inputs and the ground truth into the topics the evaluation reads.
+
+ Returns
+ -------
+ A dictionary mapping the names of the inputs, followed by those of the
+ ground truth, to their ROS message types.
+
+ Raises
+ ------
+ ValueError
+ If the evaluation reads no topic, an input and a ground-truth topic
+ share a name, or a derived topic refers to a topic it does not read.
+ """
+ inputs = dict(self.required_inputs())
+ ground_truth = self.required_ground_truth()
+ shared_names = sorted(set(inputs) & set(ground_truth))
+ if shared_names:
+ raise ValueError(f"Evaluation '{self.name}' declares {shared_names} as input and as ground truth")
+ inputs.update(ground_truth)
+ if not inputs:
+ raise ValueError(f"Evaluation '{self.name}' declares no topic to evaluate")
+ for name, (source, _) in self.derived_topics().items():
+ if name not in inputs or source not in inputs or source == name:
+ raise ValueError(f"Evaluation '{self.name}' derives the topic of '{name}' from that of '{source}'")
+ return inputs
+
# ------------------------------------------------------------------
# Optional visualization interface
# ------------------------------------------------------------------
def visualization_outputs(self) -> Dict[str, Any]:
- """Define the ROS message types this benchmark publishes for visualization.
+ """Define the ROS message types this evaluation publishes for visualization.
Returns
-------
A dictionary mapping output names to their ROS message types, empty for
- a benchmark that offers no visualization. The names are node-relative
+ an evaluation that offers no visualization. The names are node-relative
topics and match the keys of :meth:`visualize_sample`.
"""
return {}
- def visualize_sample(self, prediction: Any, label: Any, sample_id: Optional[str] = None, **auxiliary: Any) -> Dict[str, Any]:
- """Build the visualization messages for a single (prediction, label) pair.
+ def visualize_sample(self, sample_id: Optional[str] = None, **messages: Any) -> Dict[str, Any]:
+ """Build the visualization messages for a single sample.
Called once per sample while visualization is enabled, with the same
- arguments as :meth:`record_sample`.
+ arguments as :meth:`compute_sample_metrics`.
Returns
-------
A dictionary mapping the names of :meth:`visualization_outputs` to ready
- ROS messages, empty for a benchmark that offers no visualization.
+ ROS messages, empty for an evaluation that offers no visualization.
"""
return {}
@@ -111,24 +187,17 @@ def visualize_sample(self, prediction: Any, label: Any, sample_id: Optional[str]
# Concrete helpers
# ------------------------------------------------------------------
- def record_sample(
- self,
- prediction: Any,
- label: Any,
- sample_id: Optional[str] = None,
- scene_id: Optional[str] = None,
- **auxiliary: Any,
- ) -> Dict[str, Any]:
+ def record_sample(self, sample_id: Optional[str] = None, scene_id: Optional[str] = None, **messages: Any) -> Dict[str, Any]:
"""Compute and store per-sample metrics.
- This is the main entry point used by the evaluation loop. Any
- keyword arguments in *auxiliary* are forwarded verbatim to
- :meth:`compute_sample_metrics`, allowing benchmarks to receive
- auxiliary per-sample data (e.g. static-map annotations) without
- changing the abstract interface.
+ This is the main entry point used by the evaluation loop. The
+ *messages* of the sample are forwarded verbatim to
+ :meth:`compute_sample_metrics`.
Parameters
----------
+ sample_id:
+ An optional identifier for the sample.
scene_id:
The scene of the dataset the sample belongs to, which
:meth:`finalize` aggregates the samples by. An evaluation loop
@@ -140,7 +209,7 @@ def record_sample(
The stored entry of the sample, as ``{"sample_id", "scene_id",
"metrics"}``.
"""
- metrics = self.compute_sample_metrics(prediction, label, sample_id, **auxiliary)
+ metrics = self.compute_sample_metrics(sample_id=sample_id, **messages)
entry: Dict[str, Any] = {"sample_id": sample_id, "scene_id": scene_id, "metrics": metrics}
self._sample_results.append(entry)
return entry
@@ -171,7 +240,7 @@ def finalize(self, complete: bool = True) -> Dict[str, Any]:
Parameters
----------
complete:
- Whether all samples of the benchmark have been evaluated. An
+ Whether all samples of the evaluation have been evaluated. An
evaluation that was interrupted, e.g. with Ctrl-C, still reports the
samples it did evaluate, marked as ``"complete": false`` so that
they are not mistaken for the results over the whole dataset.
@@ -179,7 +248,7 @@ def finalize(self, complete: bool = True) -> Dict[str, Any]:
aggregated = self.compute_aggregated_metrics(self._sample_results)
scenes = self.sample_results_by_scene()
return {
- "benchmark": self.name,
+ "evaluation": self.name,
"description": self.description,
"complete": complete,
"num_samples": len(self._sample_results),
@@ -207,7 +276,7 @@ def save_results(self, output_path: str, results: Optional[Dict[str, Any]] = Non
A results payload previously obtained from :meth:`finalize`, which
is computed here when omitted.
complete:
- Whether all samples of the benchmark have been evaluated, see
+ Whether all samples of the evaluation have been evaluated, see
:meth:`finalize`. Only used while the results are computed here; a
given payload is written with the flag it was finalized with.
diff --git a/autonomy_evaluation/autonomy_evaluation/evaluations/__init__.py b/autonomy_evaluation/autonomy_evaluation/evaluations/__init__.py
new file mode 100644
index 0000000..972640b
--- /dev/null
+++ b/autonomy_evaluation/autonomy_evaluation/evaluations/__init__.py
@@ -0,0 +1,9 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Evaluations of automated driving tasks, each computing the metrics of one task."""
+
+from autonomy_evaluation.evaluations.Evaluation import Evaluation
+from autonomy_evaluation.evaluations.registry import EVALUATIONS, load_evaluation
+
+__all__ = ["Evaluation", "EVALUATIONS", "load_evaluation"]
diff --git a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/NuscenesLidarObjectDetection.py b/autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/NuscenesLidarObjectDetection.py
similarity index 96%
rename from autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/NuscenesLidarObjectDetection.py
rename to autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/NuscenesLidarObjectDetection.py
index 519c975..0b7fb0e 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/NuscenesLidarObjectDetection.py
+++ b/autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/NuscenesLidarObjectDetection.py
@@ -1,7 +1,7 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""nuScenes - Lidar object detection benchmark.
+"""nuScenes - Lidar object detection evaluation.
Evaluates 3D bounding-box predictions from lidar point cloud input against
nuScenes ground-truth labels using the official nuScenes evaluation protocol.
@@ -50,9 +50,9 @@
from typing import Any, Dict, List, Optional, Tuple
import numpy as np
-from autonomy_benchmarks.benchmarks.AutonomyBenchmark import AutonomyBenchmark
-from autonomy_benchmarks.utils.ObjectDetectionUtils import ObjectDetectionUtils
from autonomy_datasets_msgs.msg import ObjectListMetaInfo
+from autonomy_evaluation.evaluations.Evaluation import Evaluation
+from autonomy_evaluation.utils.ObjectDetectionUtils import ObjectDetectionUtils
from perception_msgs.msg import ObjectClassification, ObjectList
from perception_msgs_utils.state_getters import (
get_height,
@@ -186,8 +186,8 @@ class DetectionClass:
_SUPPORTED_CLASSES: frozenset = frozenset(set(_PERCEPTION_TYPE_TO_CLASS.values()) - {_CLASS_IGNORE}) & _DETECTION_CLASSES
-class NuscenesLidarObjectDetection(AutonomyBenchmark):
- """Benchmark for lidar 3D object detection on the nuScenes dataset.
+class NuscenesLidarObjectDetection(Evaluation):
+ """Evaluation of lidar 3D object detection on the nuScenes dataset.
See the module docstring for full evaluation settings. Objects are keyed by
their canonical nuScenes detection class name. The classes evaluated with a
@@ -202,7 +202,7 @@ def __init__(self) -> None:
"""Configure nuScenes thresholds, class ranges, and TP metric rules."""
super().__init__(
name="nuscenes_lidar_object_detection",
- description=("3D bounding-box object detection benchmark from lidar point clouds using the nuScenes dataset."),
+ description=("3D bounding-box object detection evaluation from lidar point clouds using the nuScenes dataset."),
)
# BEV distance thresholds used for AP and TP-metric computation.
@@ -250,26 +250,44 @@ def __init__(self) -> None:
# ------------------------------------------------------------------
def required_inputs(self) -> Dict[str, Any]:
- """Define expected input ROS message types.
+ """Define the evaluated output of the system under test.
+
+ Returns:
+ Input name to ROS message type: the ``prediction`` object list of
+ the detector, matching the ``compute_sample_metrics`` parameter.
+ """
+
+ return {"prediction": ObjectList}
+
+ def required_ground_truth(self) -> Dict[str, Any]:
+ """Define the dataset labels the predictions are compared with.
``label_meta_info`` carries the dataset annotations of the ``label``
- object list (``original_class``, ``attribute``, point counts). The
- dataset publishes it next to the object list, on
- ``/meta_info``, so the launch file derives its topic from
- the ``label`` argument instead of exposing one of its own.
+ object list (``original_class``, ``attribute``, point counts).
Returns:
- Input name to ROS message type. The keys match the
- ``compute_sample_metrics`` parameters and are remapped to real ROS
- topics in the launch file.
+ Ground-truth name to ROS message type, matching the
+ ``compute_sample_metrics`` parameters.
"""
return {
- "prediction": ObjectList,
"label": ObjectList,
"label_meta_info": ObjectListMetaInfo,
}
+ def derived_topics(self) -> Dict[str, Tuple[str, str]]:
+ """Follow the topic of the labels with the topic of their meta information.
+
+ The dataset publishes the meta information of an object list next to
+ it, on ``/meta_info``, so only the ``label`` topic needs
+ to be configured.
+
+ Returns:
+ ``label_meta_info`` derived from ``label`` with suffix ``/meta_info``.
+ """
+
+ return {"label_meta_info": ("label", "/meta_info")}
+
@staticmethod
def _index_meta_info(meta_info: Any) -> Dict[int, Dict[str, List[str]]]:
"""Index an ``ObjectListMetaInfo`` message by object ID.
@@ -514,7 +532,7 @@ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Di
Returns:
A dict containing ``threshold_metrics`` (per-threshold AP / mAP with
the P-R filter), ``tp_metrics`` (per-class ATE, ASE, AOE, AVE, AAE
- at the 2 m threshold), and ``benchmark_score`` (flat mAP and NDS
+ at the 2 m threshold), and ``score`` (flat mAP and NDS
summary).
"""
all_match_records = [entry["metrics"]["match_records"] for entry in sample_results]
@@ -526,13 +544,13 @@ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Di
threshold_metrics = self._compute_threshold_metrics(cumulative_results)
# Compute TP error metrics at the 2 m threshold.
tp_metrics = self._compute_tp_metrics(merged_records)
- # Compute benchmark score (mAP and NDS).
- benchmark_score = self._compute_benchmark_score(threshold_metrics, tp_metrics)
+ # Compute the score (mAP and NDS).
+ score = self._compute_score(threshold_metrics, tp_metrics)
return {
"threshold_metrics": threshold_metrics,
"tp_metrics": tp_metrics,
- "benchmark_score": benchmark_score,
+ "score": score,
}
# ------------------------------------------------------------------
@@ -1055,7 +1073,7 @@ class with no GT/predictions falls back to worst-case ``1.0`` errors.
}
return tp_metrics
- def _compute_benchmark_score(
+ def _compute_score(
self,
threshold_metrics: Dict[str, Any],
tp_metrics: Dict[str, Dict[str, Any]],
diff --git a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/__init__.py b/autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/__init__.py
similarity index 51%
rename from autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/__init__.py
rename to autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/__init__.py
index 3da1540..de18a9e 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/lidar_object_detection/__init__.py
+++ b/autonomy_evaluation/autonomy_evaluation/evaluations/lidar_object_detection/__init__.py
@@ -1,9 +1,9 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""LiDAR object detection benchmarks."""
+"""LiDAR object detection evaluations."""
-from autonomy_benchmarks.benchmarks.lidar_object_detection.NuscenesLidarObjectDetection import (
+from autonomy_evaluation.evaluations.lidar_object_detection.NuscenesLidarObjectDetection import (
NuscenesLidarObjectDetection,
)
diff --git a/autonomy_evaluation/autonomy_evaluation/evaluations/registry.py b/autonomy_evaluation/autonomy_evaluation/evaluations/registry.py
new file mode 100644
index 0000000..379ac56
--- /dev/null
+++ b/autonomy_evaluation/autonomy_evaluation/evaluations/registry.py
@@ -0,0 +1,62 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Registry of the evaluations the node runs, selected by their name.
+
+Besides the evaluations of this package, an evaluation implemented in another
+package is selected as ``:``, e.g.
+``my_package.evaluations:TimeToCollision``, so that a system under test can be
+evaluated without adding its evaluation to this package.
+"""
+
+from __future__ import annotations
+
+import importlib
+from typing import Dict, Tuple
+
+from autonomy_evaluation.evaluations.Evaluation import Evaluation
+
+# Evaluations of this package by the name they are selected with, as (module, class). A module is
+# only imported once its evaluation is selected, so that running one evaluation does not require
+# the dependencies of all others.
+EVALUATIONS: Dict[str, Tuple[str, str]] = {
+ "nuscenes_lidar_object_detection": (
+ "autonomy_evaluation.evaluations.lidar_object_detection.NuscenesLidarObjectDetection",
+ "NuscenesLidarObjectDetection",
+ ),
+}
+
+
+def load_evaluation(name: str) -> Evaluation:
+ """Instantiate the evaluation selected by its name.
+
+ Parameters
+ ----------
+ name:
+ Name of an evaluation of this package (see :data:`EVALUATIONS`), or
+ ``:`` of an :class:`Evaluation` subclass implemented in
+ another package, which is instantiated without arguments.
+
+ Returns
+ -------
+ The selected evaluation.
+
+ Raises
+ ------
+ ValueError
+ If no evaluation of this name exists, or its class cannot be loaded.
+ """
+ if name in EVALUATIONS:
+ module_name, class_name = EVALUATIONS[name]
+ else:
+ module_name, _, class_name = name.partition(":")
+ if not module_name or not class_name:
+ raise ValueError(f"Unknown evaluation '{name}', expected one of {sorted(EVALUATIONS)} or ':'")
+
+ try:
+ evaluation_class = getattr(importlib.import_module(module_name), class_name)
+ except (ImportError, AttributeError) as exception:
+ raise ValueError(f"Cannot load evaluation '{name}' from '{module_name}:{class_name}': {exception}") from exception
+ if not isinstance(evaluation_class, type) or not issubclass(evaluation_class, Evaluation):
+ raise ValueError(f"Evaluation '{name}' is no subclass of {Evaluation.__module__}.{Evaluation.__name__}")
+ return evaluation_class()
diff --git a/autonomy_benchmarks/autonomy_benchmarks/utils/ObjectDetectionUtils.py b/autonomy_evaluation/autonomy_evaluation/utils/ObjectDetectionUtils.py
similarity index 97%
rename from autonomy_benchmarks/autonomy_benchmarks/utils/ObjectDetectionUtils.py
rename to autonomy_evaluation/autonomy_evaluation/utils/ObjectDetectionUtils.py
index e610bbf..a381835 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/utils/ObjectDetectionUtils.py
+++ b/autonomy_evaluation/autonomy_evaluation/utils/ObjectDetectionUtils.py
@@ -1,9 +1,9 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Utility helpers for object-detection benchmarks.
+"""Utility helpers for object-detection evaluations.
-Contains reusable functions shared across all benchmark implementations:
+Contains reusable functions shared across all evaluation implementations:
- **AP integration** (``compute_ap_101_point``): 101-point recall interpolation.
- **TP metric integration** (``compute_tp_101_point``): average TP metric over a
@@ -30,9 +30,9 @@
class ObjectDetectionUtils:
- """Static helpers shared across the object-detection benchmarks.
+ """Static helpers shared across the object-detection evaluations.
- Groups two kinds of stateless utilities used by every benchmark:
+ Groups two kinds of stateless utilities used by every evaluation:
* Geometry on plain ``dict`` object records: 2D and volumetric 3D IoU,
the rotated BEV footprint polygon, BEV center distance, and BEV
diff --git a/autonomy_benchmarks/autonomy_benchmarks/utils/__init__.py b/autonomy_evaluation/autonomy_evaluation/utils/__init__.py
similarity index 51%
rename from autonomy_benchmarks/autonomy_benchmarks/utils/__init__.py
rename to autonomy_evaluation/autonomy_evaluation/utils/__init__.py
index 2089998..469f404 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/utils/__init__.py
+++ b/autonomy_evaluation/autonomy_evaluation/utils/__init__.py
@@ -1,8 +1,8 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Utility functions for AutonomyHub benchmarks."""
+"""Utility functions shared by the evaluations."""
-from autonomy_benchmarks.utils.ObjectDetectionUtils import ObjectDetectionUtils
+from autonomy_evaluation.utils.ObjectDetectionUtils import ObjectDetectionUtils
__all__ = ["ObjectDetectionUtils"]
diff --git a/autonomy_benchmarks/config/conf.rviz b/autonomy_evaluation/config/conf.rviz
similarity index 99%
rename from autonomy_benchmarks/config/conf.rviz
rename to autonomy_evaluation/config/conf.rviz
index af56693..571471e 100644
--- a/autonomy_benchmarks/config/conf.rviz
+++ b/autonomy_evaluation/config/conf.rviz
@@ -372,7 +372,7 @@ Visualization Manager:
Filter size: 10
History Policy: Keep Last
Reliability Policy: Reliable
- Value: /autonomy_benchmarks/true_positives
+ Value: /autonomy_evaluation/true_positives
Value: true
Velocity arrow:
Use velocity color: false
@@ -447,7 +447,7 @@ Visualization Manager:
Filter size: 10
History Policy: Keep Last
Reliability Policy: Reliable
- Value: /autonomy_benchmarks/false_positives
+ Value: /autonomy_evaluation/false_positives
Value: true
Velocity arrow:
Use velocity color: false
@@ -522,7 +522,7 @@ Visualization Manager:
Filter size: 10
History Policy: Keep Last
Reliability Policy: Reliable
- Value: /autonomy_benchmarks/false_negatives
+ Value: /autonomy_evaluation/false_negatives
Value: true
Velocity arrow:
Use velocity color: false
diff --git a/autonomy_evaluation/launch/autonomy_evaluation.launch.py b/autonomy_evaluation/launch/autonomy_evaluation.launch.py
new file mode 100644
index 0000000..27edab0
--- /dev/null
+++ b/autonomy_evaluation/launch/autonomy_evaluation.launch.py
@@ -0,0 +1,167 @@
+#!/usr/bin/env python3
+
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+from typing import Mapping
+
+from autonomy_evaluation.evaluations import Evaluation, load_evaluation
+from launch import LaunchContext, LaunchDescription
+from launch.actions import DeclareLaunchArgument, OpaqueFunction
+from launch.conditions import IfCondition
+from launch.substitutions import LaunchConfiguration, PathJoinSubstitution
+from launch_ros.actions import Node, SetParameter
+from launch_ros.parameter_descriptions import ParameterValue
+from launch_ros.substitutions import FindPackageShare
+
+
+def input_remappings(evaluation: Evaluation, launch_configurations: Mapping[str, str]) -> list[tuple[str, str]]:
+ """Remap the topics the evaluation reads onto the topics given by the launch arguments of their names.
+
+ Every input and ground-truth topic of the evaluation is configured by a launch argument of its name,
+ e.g. ``prediction:=/object_list/prediction``. A topic without such an argument is subscribed in the
+ private namespace of the node, except for a topic published next to another one, which the node derives
+ from the topic of that one.
+
+ Args:
+ evaluation (Evaluation): evaluation to run
+ launch_configurations (Mapping[str, str]): values of the launch arguments by name
+
+ Returns:
+ list[tuple[str, str]]: remappings of the node-relative input names
+ """
+ derived_topics = evaluation.derived_topics()
+ remappings = []
+ for name in evaluation.all_inputs():
+ if name in launch_configurations:
+ remappings.append((name, launch_configurations[name]))
+ elif name not in derived_topics:
+ remappings.append((name, f"~/{name}"))
+ return remappings
+
+
+def launch_node(context: LaunchContext, remappings: list) -> list[Node]:
+ """Launch the node with the topics of the selected evaluation remapped.
+
+ Args:
+ context (LaunchContext): launch context holding the launch arguments
+ remappings (list): remappings of the names other than the topics of the evaluation
+
+ Returns:
+ list[Node]: node to launch
+ """
+ evaluation = load_evaluation(LaunchConfiguration("evaluation").perform(context))
+ node = Node(
+ package="autonomy_evaluation",
+ executable="autonomy_evaluation",
+ namespace=LaunchConfiguration("namespace"),
+ name=LaunchConfiguration("name"),
+ parameters=[
+ {"evaluation": LaunchConfiguration("evaluation")},
+ {"visualize": ParameterValue(LaunchConfiguration("visualize"), value_type=bool)},
+ {"sample_source": LaunchConfiguration("sample_source")},
+ {"sync_tolerance": ParameterValue(LaunchConfiguration("sync_tolerance"), value_type=float)},
+ {"samples_per_request": ParameterValue(LaunchConfiguration("samples_per_request"), value_type=int)},
+ {"sample_ids": ParameterValue(LaunchConfiguration("sample_ids"), value_type=str)},
+ {"evaluation_timeout": ParameterValue(LaunchConfiguration("evaluation_timeout"), value_type=float)},
+ {"results_path": ParameterValue(LaunchConfiguration("results_path"), value_type=str)},
+ ],
+ arguments=["--ros-args", "--log-level", LaunchConfiguration("log_level")],
+ remappings=input_remappings(evaluation, context.launch_configurations) + remappings,
+ output="screen",
+ emulate_tty=True,
+ )
+ return [node]
+
+
+def generate_launch_description():
+ """Create and return the launch description for the autonomy_evaluation node."""
+
+ # Service of the dataset node the evaluation requests the samples to evaluate from, remapped
+ # onto the topic its argument resolves to just like the evaluation's data inputs.
+ remappable_topics = [
+ DeclareLaunchArgument(
+ "request_samples",
+ default_value="~/request_samples",
+ description="service of the dataset node used to request the samples to evaluate",
+ ),
+ ]
+
+ args = [
+ *remappable_topics,
+ DeclareLaunchArgument(
+ "evaluation",
+ default_value="nuscenes_lidar_object_detection",
+ description="name of an evaluation of this package, or ':' of an evaluation implemented in "
+ "another package; each topic the evaluation reads is set by an argument of its name, e.g. prediction:=/topic",
+ ),
+ DeclareLaunchArgument("name", default_value="autonomy_evaluation", description="node name"),
+ DeclareLaunchArgument("namespace", default_value="", description="node namespace"),
+ DeclareLaunchArgument(
+ "log_level", default_value="info", description="ROS logging level (debug, info, warn, error, fatal)"
+ ),
+ DeclareLaunchArgument("use_sim_time", default_value="true", description="use simulation clock"),
+ DeclareLaunchArgument(
+ "visualize",
+ default_value="false",
+ choices=["true", "false"],
+ description="publish the per-sample visualization of the evaluation and open RViz on it",
+ ),
+ DeclareLaunchArgument(
+ "sample_source",
+ default_value="dataset",
+ choices=["dataset", "external"],
+ description="'dataset' requests the samples to evaluate from the dataset one after another; 'external' "
+ "evaluates the samples published by others, e.g. by a simulation or the playback panel in RViz",
+ ),
+ DeclareLaunchArgument(
+ "sync_tolerance",
+ default_value="0.0",
+ description="seconds by which the stamps of the messages of a sample may differ (0 matches exact stamps only)",
+ ),
+ DeclareLaunchArgument(
+ "samples_per_request",
+ default_value="1",
+ description="number of samples to request from the dataset at a time (0 requests all remaining samples at once)",
+ ),
+ DeclareLaunchArgument(
+ "sample_ids",
+ default_value="",
+ description="comma-separated IDs of the dataset samples to evaluate (all samples if empty)",
+ ),
+ DeclareLaunchArgument(
+ "evaluation_timeout",
+ default_value="60.0",
+ description="seconds to wait for a published sample to be evaluated before continuing without it",
+ ),
+ DeclareLaunchArgument(
+ "results_path",
+ default_value="",
+ description="path of the JSON file the evaluation results are written to (results are only logged if empty)",
+ ),
+ ]
+
+ # The topics the evaluation reads depend on the selected evaluation, so the node is only
+ # created once the launch arguments are known
+ remappings = [(la.default_value[0].text, LaunchConfiguration(la.name)) for la in remappable_topics]
+ node = OpaqueFunction(function=launch_node, kwargs={"remappings": remappings})
+ rviz = Node(
+ package="rviz2",
+ executable="rviz2",
+ name="rviz2",
+ arguments=[
+ "-d",
+ PathJoinSubstitution([FindPackageShare("autonomy_evaluation"), "config", "conf.rviz"]),
+ ],
+ condition=IfCondition(LaunchConfiguration("visualize")),
+ output="screen",
+ )
+
+ return LaunchDescription(
+ [
+ *args,
+ SetParameter("use_sim_time", LaunchConfiguration("use_sim_time")),
+ node,
+ rviz,
+ ]
+ )
diff --git a/autonomy_benchmarks/package.xml b/autonomy_evaluation/package.xml
similarity index 86%
rename from autonomy_benchmarks/package.xml
rename to autonomy_evaluation/package.xml
index 2121ce3..614af09 100644
--- a/autonomy_benchmarks/package.xml
+++ b/autonomy_evaluation/package.xml
@@ -2,9 +2,9 @@
- autonomy_benchmarks
+ autonomy_evaluation
1.0.0
- Benchmarking suite for automated driving tasks
+ Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments
Haohao Hu
Raphael van Kempen
diff --git a/autonomy_benchmarks/resource/autonomy_benchmarks b/autonomy_evaluation/resource/autonomy_evaluation
similarity index 100%
rename from autonomy_benchmarks/resource/autonomy_benchmarks
rename to autonomy_evaluation/resource/autonomy_evaluation
diff --git a/autonomy_evaluation/setup.cfg b/autonomy_evaluation/setup.cfg
new file mode 100644
index 0000000..59b63fe
--- /dev/null
+++ b/autonomy_evaluation/setup.cfg
@@ -0,0 +1,4 @@
+[develop]
+script_dir=$base/lib/autonomy_evaluation
+[install]
+install_scripts=$base/lib/autonomy_evaluation
diff --git a/autonomy_benchmarks/setup.py b/autonomy_evaluation/setup.py
similarity index 86%
rename from autonomy_benchmarks/setup.py
rename to autonomy_evaluation/setup.py
index c481265..4043461 100644
--- a/autonomy_benchmarks/setup.py
+++ b/autonomy_evaluation/setup.py
@@ -6,7 +6,7 @@
from setuptools import find_packages, setup
-package_name = "autonomy_benchmarks"
+package_name = "autonomy_evaluation"
setup(
name=package_name,
@@ -26,6 +26,6 @@
license="TODO: License declaration",
tests_require=["pytest"],
entry_points={
- "console_scripts": ["autonomy_benchmarks = autonomy_benchmarks.autonomy_benchmarks:main"],
+ "console_scripts": ["autonomy_evaluation = autonomy_evaluation.autonomy_evaluation:main"],
},
)
diff --git a/autonomy_benchmarks/tests/__init__.py b/autonomy_evaluation/tests/__init__.py
similarity index 56%
rename from autonomy_benchmarks/tests/__init__.py
rename to autonomy_evaluation/tests/__init__.py
index 23e5cd9..f3e5696 100644
--- a/autonomy_benchmarks/tests/__init__.py
+++ b/autonomy_evaluation/tests/__init__.py
@@ -1,4 +1,4 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Test suite for the autonomy_benchmarks package."""
+"""Test suite for the autonomy_evaluation package."""
diff --git a/autonomy_benchmarks/tests/utils/__init__.py b/autonomy_evaluation/tests/evaluations/__init__.py
similarity index 56%
rename from autonomy_benchmarks/tests/utils/__init__.py
rename to autonomy_evaluation/tests/evaluations/__init__.py
index 75af748..ae03f57 100644
--- a/autonomy_benchmarks/tests/utils/__init__.py
+++ b/autonomy_evaluation/tests/evaluations/__init__.py
@@ -1,4 +1,4 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Utility test modules for autonomy_benchmarks."""
+"""Tests of the evaluations of autonomy_evaluation."""
diff --git a/autonomy_benchmarks/tests/benchmarks/lidar_object_detection/__init__.py b/autonomy_evaluation/tests/evaluations/lidar_object_detection/__init__.py
similarity index 59%
rename from autonomy_benchmarks/tests/benchmarks/lidar_object_detection/__init__.py
rename to autonomy_evaluation/tests/evaluations/lidar_object_detection/__init__.py
index 65f82a8..cfb62fd 100644
--- a/autonomy_benchmarks/tests/benchmarks/lidar_object_detection/__init__.py
+++ b/autonomy_evaluation/tests/evaluations/lidar_object_detection/__init__.py
@@ -1,4 +1,4 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Tests for the nuScenes lidar benchmark."""
+"""Tests for the nuScenes lidar evaluation."""
diff --git a/autonomy_benchmarks/tests/benchmarks/lidar_object_detection/test_NuscenesLidarObjectDetection.py b/autonomy_evaluation/tests/evaluations/lidar_object_detection/test_NuscenesLidarObjectDetection.py
similarity index 91%
rename from autonomy_benchmarks/tests/benchmarks/lidar_object_detection/test_NuscenesLidarObjectDetection.py
rename to autonomy_evaluation/tests/evaluations/lidar_object_detection/test_NuscenesLidarObjectDetection.py
index a073600..6ce04ee 100644
--- a/autonomy_benchmarks/tests/benchmarks/lidar_object_detection/test_NuscenesLidarObjectDetection.py
+++ b/autonomy_evaluation/tests/evaluations/lidar_object_detection/test_NuscenesLidarObjectDetection.py
@@ -1,10 +1,10 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Tests for NuscenesLidarObjectDetection benchmark.
+"""Tests for NuscenesLidarObjectDetection evaluation.
Tests focus on public API methods: compute_sample_metrics(),
-compute_aggregated_metrics(), and benchmark configuration.
+compute_aggregated_metrics(), and evaluation configuration.
Inputs are real ROS messages (built via the helpers below), matching how
``_extract_objects`` reads them: geometry through the ``perception_msgs_utils``
@@ -21,10 +21,10 @@
from typing import List, Optional, Tuple
import pytest
-from autonomy_benchmarks.benchmarks.lidar_object_detection.NuscenesLidarObjectDetection import (
+from autonomy_datasets_msgs.msg import ObjectListMetaInfo, ObjectMetaInfo
+from autonomy_evaluation.evaluations.lidar_object_detection.NuscenesLidarObjectDetection import (
NuscenesLidarObjectDetection,
)
-from autonomy_datasets_msgs.msg import ObjectListMetaInfo, ObjectMetaInfo
from diagnostic_msgs.msg import KeyValue
from perception_msgs.msg import HEXAMOTION, Object, ObjectClassification, ObjectList, ObjectState
@@ -164,20 +164,20 @@ def _label(
class TestNuscenesLidarObjectDetection:
- """Tests for NuscenesLidarObjectDetection benchmark.
+ """Tests for NuscenesLidarObjectDetection evaluation.
Covers sample metrics computation (Pass 1), aggregated metrics
(Pass 2), and NDS score calculation.
"""
def setup_method(self):
- """Create a fresh benchmark instance for each test."""
+ """Create a fresh evaluation instance for each test."""
self.bm = NuscenesLidarObjectDetection()
def _metrics(self, pred_objs, gts, reverse_meta: bool = False) -> dict:
"""Compute sample metrics from prediction objects and ``_gt`` pairs.
- Mirrors the node's synchronized callback, which hands the benchmark the
+ Mirrors the node's synchronized callback, which hands the evaluation the
label object list and its meta info message together.
"""
label, label_meta_info = _label(gts, reverse_meta=reverse_meta)
@@ -367,19 +367,19 @@ def test_perfect_matching_yields_high_map(self):
preds = [_pred(x=float(i), confidence=1.0 - i * 0.01) for i in range(n)]
gts = [_gt(x=float(i), num_lidar_pts=5) for i in range(n)]
result = self.bm.compute_aggregated_metrics([self._make_sample_result(preds, gts)])
- assert result["benchmark_score"]["map"] >= 0.8
+ assert result["score"]["map"] >= 0.8
def test_no_pred_yields_zero_map(self):
"""Verify missing predictions yield zero mAP."""
gts = [_gt(x=0.0, num_lidar_pts=5)]
result = self.bm.compute_aggregated_metrics([self._make_sample_result([], gts)])
- assert result["benchmark_score"]["map"] == 0.0
+ assert result["score"]["map"] == 0.0
def test_all_fp_yields_zero_map(self):
"""Verify all false positives yield zero mAP."""
preds = [_pred(x=float(i), confidence=0.9) for i in range(5)]
result = self.bm.compute_aggregated_metrics([self._make_sample_result(preds, [])])
- assert result["benchmark_score"]["map"] == 0.0
+ assert result["score"]["map"] == 0.0
def test_unsupported_class_gt_excluded_from_map(self):
"""A GT class the interface cannot express (truck) must not dilute mAP.
@@ -393,7 +393,7 @@ def test_unsupported_class_gt_excluded_from_map(self):
gts.append(_gt(x=5.0, y=20.0, original_class="vehicle.truck", num_lidar_pts=5))
result = self.bm.compute_aggregated_metrics([self._make_sample_result(preds, gts)])
# truck is unsupported -> excluded; mAP stays high (car only).
- assert result["benchmark_score"]["map"] >= 0.8
+ assert result["score"]["map"] >= 0.8
assert "truck" not in self.bm.supported_classes
def test_supported_but_undetected_class_still_counts_as_zero(self):
@@ -408,13 +408,13 @@ def test_supported_but_undetected_class_still_counts_as_zero(self):
gts.append(_gt(x=5.0, y=20.0, original_class="human.pedestrian.adult", num_lidar_pts=5))
result = self.bm.compute_aggregated_metrics([self._make_sample_result(preds, gts)])
# pedestrian is supported -> AP 0 pulls the mean well below the car-only case.
- assert result["benchmark_score"]["map"] < 0.6
+ assert result["score"]["map"] < 0.6
assert "pedestrian" in self.bm.supported_classes
def test_capability_lists_reported(self):
"""The score declares which classes the interface can / cannot express."""
gts = [_gt(x=0.0, num_lidar_pts=5)]
- score = self.bm.compute_aggregated_metrics([self._make_sample_result([], gts)])["benchmark_score"]
+ score = self.bm.compute_aggregated_metrics([self._make_sample_result([], gts)])["score"]
assert score["supported_classes"] == sorted(self.bm.supported_classes)
assert "car" in score["supported_classes"]
assert "truck" in score["unsupported_classes"]
@@ -428,10 +428,31 @@ def test_multi_frame_accumulation(self):
merged = self.bm._merge_match_records([r1["metrics"]["match_records"], r2["metrics"]["match_records"]])
assert merged[0.5]["car"]["gt_count"] == 2
+ # --- topics ---
+
+ def test_compares_the_predictions_with_the_labels_as_ground_truth(self):
+ """The prediction of the detector is evaluated against the labels and their meta information."""
+ assert self.bm.required_inputs() == {"prediction": ObjectList}
+ assert self.bm.required_ground_truth() == {"label": ObjectList, "label_meta_info": ObjectListMetaInfo}
+
+ def test_label_meta_info_follows_the_label_topic(self):
+ """The dataset publishes the meta information next to the labels, on '/meta_info'."""
+ assert self.bm.derived_topics() == {"label_meta_info": ("label", "/meta_info")}
+
+ def test_sample_metrics_are_computed_from_the_messages_by_topic_name(self):
+ """The node passes the messages of a sample by the names of their topics."""
+ label, label_meta_info = _label([_gt(x=0.0, num_lidar_pts=5)])
+
+ prediction = _msg([_pred(x=0.0)])
+
+ result = self.bm.record_sample(sample_id="0", prediction=prediction, label=label, label_meta_info=label_meta_info)
+
+ assert result["metrics"]["sample_ground_truth_num"] == 1
+
# --- visualization of the per-sample matching outcome ---
def test_visualization_outputs_are_object_lists(self):
- """The benchmark declares one ObjectList output per matching outcome."""
+ """The evaluation declares one ObjectList output per matching outcome."""
assert self.bm.visualization_outputs() == {
"true_positives": ObjectList,
"false_positives": ObjectList,
@@ -488,7 +509,7 @@ def _tm(overall_map: float) -> dict:
def test_nds_with_valid_aae(self):
"""Verify NDS uses full 6-term formula when AAE data is available."""
tp_metrics = {"car": {"ate": 0.2, "ase": 0.1, "aoe": 0.3, "ave": 0.4, "aae": 0.05}}
- score = self.bm._compute_benchmark_score(self._tm(0.5), tp_metrics)
+ score = self.bm._compute_score(self._tm(0.5), tp_metrics)
expected = (
5.0 * 0.5
+ max(1.0 - 0.2, 0.0)
@@ -503,13 +524,13 @@ def test_nds_without_aae_is_lower(self):
"""Verify missing AAE data (falls back to 1.0) reduces NDS score."""
tp_with = {"car": {"ate": 0.0, "ase": 0.0, "aoe": 0.0, "ave": 0.0, "aae": 0.0}}
tp_without = {"car": {"ate": 0.0, "ase": 0.0, "aoe": 0.0, "ave": 0.0, "aae": None}}
- score_with = self.bm._compute_benchmark_score(self._tm(1.0), tp_with)
- score_without = self.bm._compute_benchmark_score(self._tm(1.0), tp_without)
+ score_with = self.bm._compute_score(self._tm(1.0), tp_with)
+ score_without = self.bm._compute_score(self._tm(1.0), tp_without)
assert score_with["nds"] == pytest.approx(1.0, abs=1e-4)
assert score_without["nds"] == pytest.approx(0.9, abs=1e-4)
def test_nds_clamps_tp_errors(self):
"""Verify TP errors above 1.0 are clamped in NDS computation."""
tp_metrics = {"car": {"ate": 2.0, "ase": 3.0, "aoe": 5.0, "ave": 1.5, "aae": 2.0}}
- score = self.bm._compute_benchmark_score(self._tm(0.0), tp_metrics)
+ score = self.bm._compute_score(self._tm(0.0), tp_metrics)
assert score["nds"] >= 0.0
diff --git a/autonomy_evaluation/tests/evaluations/test_Evaluation.py b/autonomy_evaluation/tests/evaluations/test_Evaluation.py
new file mode 100644
index 0000000..9fcea9b
--- /dev/null
+++ b/autonomy_evaluation/tests/evaluations/test_Evaluation.py
@@ -0,0 +1,288 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Tests for the topics and the result store of the Evaluation base class.
+
+An evaluation reads the topics of a system under test and, if it compares them with a reference,
+ground-truth topics. Metrics are reported on three levels: for every single sample, aggregated over
+the samples of each scene of the dataset, and aggregated over all evaluated samples. Minimal
+evaluations whose metrics are trivial to predict are used, so that the tests cover the topics and
+the grouping and not a metric definition.
+"""
+
+from __future__ import annotations
+
+import json
+from typing import Any, Dict, List, Optional
+
+import pytest
+from autonomy_evaluation.evaluations.Evaluation import Evaluation
+
+
+class CountingEvaluation(Evaluation):
+ """Minimal evaluation counting the objects of a sample and summing them up when aggregating."""
+
+ def __init__(self) -> None:
+ """Name the evaluation."""
+ super().__init__(name="counting", description="counts objects")
+
+ def required_inputs(self) -> Dict[str, Any]:
+ """Declare the evaluated input, which this evaluation does not read from ROS messages."""
+ return {"prediction": object}
+
+ def required_ground_truth(self) -> Dict[str, Any]:
+ """Declare the ground truth the input is compared with."""
+ return {"label": object}
+
+ def compute_sample_metrics(self, prediction: Any, label: Any, sample_id: Optional[str] = None) -> Dict[str, Any]:
+ """Report the given prediction and label counts of a single sample."""
+ return {"num_predictions": prediction, "num_labels": label}
+
+ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Dict[str, Any]:
+ """Sum the counts of the given samples."""
+ return {
+ "num_predictions": sum(entry["metrics"]["num_predictions"] for entry in sample_results),
+ "num_labels": sum(entry["metrics"]["num_labels"] for entry in sample_results),
+ }
+
+
+class ClosestObjectEvaluation(Evaluation):
+ """Minimal evaluation of inputs only, reporting the closest object without any ground truth."""
+
+ def __init__(self) -> None:
+ """Name the evaluation."""
+ super().__init__(name="closest_object")
+
+ def required_inputs(self) -> Dict[str, Any]:
+ """Declare the ego position and the object distances, as a closed-loop planner is evaluated on."""
+ return {"ego_position": object, "object_positions": object}
+
+ def compute_sample_metrics(self, ego_position: float, object_positions: List[float], sample_id: Optional[str] = None):
+ """Report the distance of the closest object of a single sample."""
+ return {"min_distance": min(abs(position - ego_position) for position in object_positions)}
+
+ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Dict[str, Any]:
+ """Report the closest distance of all given samples."""
+ return {"min_distance": min(entry["metrics"]["min_distance"] for entry in sample_results)}
+
+
+class _TopicsEvaluation(CountingEvaluation):
+ """Counting evaluation with freely declared topics, to test their validation."""
+
+ def __init__(self, inputs: Dict[str, Any], ground_truth: Dict[str, Any], derived_topics=None) -> None:
+ """Declare the given topics."""
+ super().__init__()
+ self._inputs, self._ground_truth, self._derived_topics = inputs, ground_truth, derived_topics or {}
+
+ def required_inputs(self) -> Dict[str, Any]:
+ """Declare the given inputs."""
+ return self._inputs
+
+ def required_ground_truth(self) -> Dict[str, Any]:
+ """Declare the given ground truth."""
+ return self._ground_truth
+
+ def derived_topics(self):
+ """Declare the given derived topics."""
+ return self._derived_topics
+
+
+def _evaluation_of(samples) -> CountingEvaluation:
+ """Record ``(sample_id, scene_id, num_predictions, num_labels)`` samples in an evaluation."""
+ evaluation = CountingEvaluation()
+ for sample_id, scene_id, num_predictions, num_labels in samples:
+ evaluation.record_sample(prediction=num_predictions, label=num_labels, sample_id=sample_id, scene_id=scene_id)
+ return evaluation
+
+
+class TestTopics:
+ """Tests declaring the topics an evaluation reads."""
+
+ def test_inputs_are_followed_by_the_ground_truth(self):
+ """The topics are passed to the evaluation in the order inputs first, ground truth second."""
+ assert list(CountingEvaluation().all_inputs()) == ["prediction", "label"]
+
+ def test_evaluation_of_inputs_only_needs_no_ground_truth(self):
+ """An evaluation computing its metrics from the system under test alone declares no ground truth."""
+ evaluation = ClosestObjectEvaluation()
+
+ assert evaluation.required_ground_truth() == {}
+ assert evaluation.derived_topics() == {}
+ assert list(evaluation.all_inputs()) == ["ego_position", "object_positions"]
+
+ def test_rejects_a_topic_declared_as_input_and_as_ground_truth(self):
+ """A topic name identifies a single message of a sample, so it cannot play both roles."""
+ with pytest.raises(ValueError, match="as input and as ground truth"):
+ _TopicsEvaluation({"objects": object}, {"objects": object}).all_inputs()
+
+ def test_rejects_an_evaluation_without_topics(self):
+ """An evaluation that reads no topic would never evaluate a sample."""
+ with pytest.raises(ValueError, match="no topic"):
+ _TopicsEvaluation({}, {}).all_inputs()
+
+ @pytest.mark.parametrize(
+ "derived_topics",
+ [{"meta_info": ("label", "/meta_info")}, {"label": ("objects", "/meta_info")}, {"label": ("label", "/meta_info")}],
+ )
+ def test_rejects_a_derived_topic_of_unknown_inputs(self, derived_topics):
+ """A derived topic and the topic it is derived from must both be read by the evaluation."""
+ with pytest.raises(ValueError, match="derives the topic"):
+ _TopicsEvaluation({"prediction": object}, {"label": object}, derived_topics).all_inputs()
+
+
+class TestSampleResults:
+ """Tests recording samples with the scene of the dataset they belong to."""
+
+ # The recorded samples themselves are no longer reported alongside the aggregated
+ # results, while 'sample_results' is commented out in Evaluation.finalize()
+ # def test_records_sample_with_its_scene(self):
+ # """A recorded sample keeps its ID, its scene and its metrics."""
+ # evaluation = _evaluation_of([("0", "scene_a", 2, 3)])
+ #
+ # assert evaluation.finalize()["sample_results"] == [
+ # {"sample_id": "0", "scene_id": "scene_a", "metrics": {"num_predictions": 2, "num_labels": 3}}
+ # ]
+
+ def test_scene_can_be_set_after_the_sample_was_recorded(self):
+ """An evaluation loop that learns the scene late sets it on the returned entry."""
+ evaluation = CountingEvaluation()
+
+ entry = evaluation.record_sample(prediction=1, label=1, sample_id="0")
+ assert entry["scene_id"] is None
+ entry["scene_id"] = "scene_a"
+
+ assert evaluation.sample_results_by_scene() == {"scene_a": [entry]}
+
+
+class TestFinalize:
+ """Tests aggregating the recorded samples per scene and over the whole evaluation."""
+
+ def test_aggregates_per_sample_scene_and_evaluation(self):
+ """Metrics are reported for every sample, every scene and all samples together."""
+ results = _evaluation_of(
+ [
+ ("0", "scene_a", 1, 1),
+ ("1", "scene_a", 2, 3),
+ ("2", "scene_b", 4, 5),
+ ]
+ ).finalize()
+
+ assert results["num_samples"] == 3
+ assert results["num_scenes"] == 2
+ assert results["aggregated_metrics"] == {"num_predictions": 7, "num_labels": 9}
+ assert results["scene_results"]["scene_a"] == {
+ "num_samples": 2,
+ "sample_ids": ["0", "1"],
+ "aggregated_metrics": {"num_predictions": 3, "num_labels": 4},
+ }
+ assert results["scene_results"]["scene_b"]["aggregated_metrics"] == {"num_predictions": 4, "num_labels": 5}
+ # The metrics of the single samples are no longer reported alongside the aggregated
+ # results, while 'sample_results' is commented out in Evaluation.finalize()
+ # assert [entry["metrics"] for entry in results["sample_results"]] == [
+ # {"num_predictions": 1, "num_labels": 1},
+ # {"num_predictions": 2, "num_labels": 3},
+ # {"num_predictions": 4, "num_labels": 5},
+ # ]
+
+ def test_groups_samples_of_a_scene_that_are_not_recorded_consecutively(self):
+ """Samples are grouped by their scene, not by the order they were recorded in."""
+ results = _evaluation_of(
+ [
+ ("0", "scene_a", 1, 0),
+ ("1", "scene_b", 2, 0),
+ ("2", "scene_a", 4, 0),
+ ]
+ ).finalize()
+
+ assert results["scene_results"]["scene_a"]["sample_ids"] == ["0", "2"]
+ assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 5
+ assert results["scene_results"]["scene_b"]["sample_ids"] == ["1"]
+
+ def test_samples_without_a_scene_are_only_aggregated_over_the_evaluation(self):
+ """A sample that cannot be attributed to a scene still counts for the whole evaluation."""
+ results = _evaluation_of([("0", "scene_a", 1, 0), ("1", None, 2, 0)]).finalize()
+
+ assert results["num_samples"] == 2
+ assert results["num_scenes"] == 1
+ assert results["aggregated_metrics"]["num_predictions"] == 3
+ assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 1
+
+ def test_aggregates_an_evaluation_of_inputs_only(self):
+ """Samples of an evaluation without ground truth are recorded from their inputs alone."""
+ evaluation = ClosestObjectEvaluation()
+ evaluation.record_sample(sample_id="0", scene_id="scene_a", ego_position=0.0, object_positions=[4.0, -2.5])
+ evaluation.record_sample(sample_id="1", scene_id="scene_a", ego_position=1.0, object_positions=[4.0])
+
+ results = evaluation.finalize()
+
+ assert results["aggregated_metrics"] == {"min_distance": 2.5}
+ assert results["scene_results"]["scene_a"]["num_samples"] == 2
+
+ def test_reports_no_scene_without_recorded_scenes(self):
+ """Samples recorded without a scene aggregate to no scene results at all."""
+ results = _evaluation_of([("0", None, 1, 0)]).finalize()
+
+ assert results["num_scenes"] == 0
+ assert results["scene_results"] == {}
+
+
+class TestSaveResults:
+ """Tests writing the results of all three levels to a JSON file."""
+
+ def test_writes_sample_scene_and_evaluation_metrics(self, tmp_path):
+ """The stored results hold the metrics of every sample, every scene and the evaluation."""
+ evaluation = _evaluation_of([("0", "scene_a", 1, 1), ("1", "scene_b", 2, 2)])
+
+ output_path = evaluation.save_results(str(tmp_path / "results" / "counting.json"))
+
+ stored = json.loads(open(output_path).read())
+ assert stored["aggregated_metrics"] == {"num_predictions": 3, "num_labels": 3}
+ assert sorted(stored["scene_results"]) == ["scene_a", "scene_b"]
+ # The single samples are no longer written alongside the aggregated
+ # results, while 'sample_results' is commented out in Evaluation.finalize()
+ # assert [entry["sample_id"] for entry in stored["sample_results"]] == ["0", "1"]
+
+ def test_writes_previously_computed_results(self, tmp_path):
+ """Results that have already been computed are written as they are."""
+ evaluation = _evaluation_of([("0", "scene_a", 1, 1)])
+ results = evaluation.finalize()
+
+ output_path = evaluation.save_results(str(tmp_path / "counting.json"), results=results)
+
+ assert json.loads(open(output_path).read())["aggregated_metrics"] == results["aggregated_metrics"]
+
+
+class TestIncompleteResults:
+ """Tests marking the results of an evaluation that did not process all samples."""
+
+ def test_results_of_all_samples_are_complete(self):
+ """Results aggregated after the last sample of the evaluation are marked as complete."""
+ assert _evaluation_of([("0", "scene_a", 1, 1)]).finalize()["complete"] is True
+
+ def test_interrupted_results_are_marked_incomplete(self):
+ """Results aggregated before the last sample, e.g. after Ctrl-C, are marked as incomplete."""
+ results = _evaluation_of([("0", "scene_a", 1, 1)]).finalize(complete=False)
+
+ assert results["complete"] is False
+ # the samples that were evaluated are still reported
+ assert results["num_samples"] == 1
+ assert results["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1}
+
+ def test_writes_incomplete_results_to_the_results_file(self, tmp_path):
+ """The results file of an interrupted evaluation marks the results it holds as incomplete."""
+ evaluation = _evaluation_of([("0", "scene_a", 1, 1)])
+
+ output_path = evaluation.save_results(str(tmp_path / "counting.json"), complete=False)
+
+ stored = json.loads(open(output_path).read())
+ assert stored["complete"] is False
+ assert stored["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1}
+
+ def test_written_results_keep_the_flag_they_were_finalized_with(self, tmp_path):
+ """A given payload is written as it is, marked the way it was finalized."""
+ evaluation = _evaluation_of([("0", "scene_a", 1, 1)])
+ results = evaluation.finalize(complete=False)
+
+ output_path = evaluation.save_results(str(tmp_path / "counting.json"), results=results)
+
+ assert json.loads(open(output_path).read())["complete"] is False
diff --git a/autonomy_evaluation/tests/evaluations/test_registry.py b/autonomy_evaluation/tests/evaluations/test_registry.py
new file mode 100644
index 0000000..50f2369
--- /dev/null
+++ b/autonomy_evaluation/tests/evaluations/test_registry.py
@@ -0,0 +1,51 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Tests for selecting the evaluation the node runs by its name.
+
+Evaluations of this package are selected by their registered name, evaluations implemented in
+another package by ``:``. A name that selects no evaluation stops the node, so it
+has to be rejected with a message naming the problem.
+"""
+
+from __future__ import annotations
+
+import pytest
+from autonomy_evaluation.evaluations import Evaluation, EVALUATIONS, load_evaluation
+from autonomy_evaluation.evaluations.lidar_object_detection.NuscenesLidarObjectDetection import (
+ NuscenesLidarObjectDetection,
+)
+
+
+class TestLoadEvaluation:
+ """Tests instantiating an evaluation by its name."""
+
+ @pytest.mark.parametrize("name", sorted(EVALUATIONS))
+ def test_loads_every_registered_evaluation(self, name):
+ """Every evaluation of this package can be selected by its registered name."""
+ assert isinstance(load_evaluation(name), Evaluation)
+
+ def test_loads_a_registered_evaluation_by_its_name(self):
+ """The registered name selects the evaluation it is registered for."""
+ assert isinstance(load_evaluation("nuscenes_lidar_object_detection"), NuscenesLidarObjectDetection)
+
+ def test_loads_an_evaluation_by_its_module_and_class(self):
+ """An evaluation of another package is selected as ':'."""
+ evaluation = load_evaluation(f"{NuscenesLidarObjectDetection.__module__}:{NuscenesLidarObjectDetection.__name__}")
+
+ assert isinstance(evaluation, NuscenesLidarObjectDetection)
+
+ @pytest.mark.parametrize(
+ "name, reason",
+ [
+ ("unknown_evaluation", "Unknown evaluation"),
+ ("autonomy_evaluation.evaluations:", "Unknown evaluation"),
+ ("no_such_package.evaluations:Evaluation", "Cannot load"),
+ ("autonomy_evaluation.evaluations:NoSuchEvaluation", "Cannot load"),
+ ("autonomy_evaluation.utils.ObjectDetectionUtils:ObjectDetectionUtils", "no subclass"),
+ ],
+ )
+ def test_rejects_a_name_that_selects_no_evaluation(self, name, reason):
+ """Names of no evaluation are rejected with the reason they cannot be loaded."""
+ with pytest.raises(ValueError, match=reason):
+ load_evaluation(name)
diff --git a/autonomy_evaluation/tests/test_autonomy_evaluation.py b/autonomy_evaluation/tests/test_autonomy_evaluation.py
new file mode 100644
index 0000000..914e7a3
--- /dev/null
+++ b/autonomy_evaluation/tests/test_autonomy_evaluation.py
@@ -0,0 +1,572 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Tests for the helpers of the autonomy_evaluation node.
+
+The node itself drives the evaluation via the ``request_samples`` service of the dataset and is
+covered by running it against a dataset. Tested here are the parsing of the samples to evaluate,
+as an unparsable value stops the node, the topics the inputs of an evaluation are subscribed on,
+the stamping of received messages, the matching of the received input messages into the samples
+to evaluate, which has to hold up when the dataset continues with a scene that was recorded before
+the scene played before it and has to match topics of a simulation within a tolerance, the
+evaluation of a sample, which requests the next samples unless others publish them, advancing the
+evaluation until it finishes, which it also has to once the dataset node has shut down after its
+last sample, and the finalization of the results, which reports the samples of an interrupted
+evaluation as incomplete.
+"""
+
+from __future__ import annotations
+
+from collections import deque
+from types import SimpleNamespace
+
+import pytest
+from autonomy_evaluation import autonomy_evaluation
+from autonomy_evaluation.autonomy_evaluation import AutonomyEvaluation, parse_sample_ids, SampleSynchronizer
+
+_TOPICS = ["prediction", "label", "label_meta_info"]
+
+# stamps of the last sample of a scene and of the first samples of the scene the dataset continues
+# with, which nuScenes recorded years earlier
+_PREVIOUS_SCENE = (1537853053, 397270000)
+_NEXT_SCENE = [(1531885320, 49418000), (1531885320, 548742000), (1531885321, 48634000)]
+
+
+def _message(stamp: tuple[int, int]) -> SimpleNamespace:
+ """Fake an input message stamped with the recording time of its sample."""
+ return SimpleNamespace(header=SimpleNamespace(stamp=SimpleNamespace(sec=stamp[0], nanosec=stamp[1])))
+
+
+def _synchronizer(queue_size: int = 10, topics=None, tolerance: float = 0.0) -> tuple[SampleSynchronizer, list]:
+ """Create a synchronizer of the evaluation inputs next to the list of the messages of the samples it matched."""
+ matched_samples: list = []
+ synchronizer = SampleSynchronizer(
+ topics or _TOPICS,
+ callback=lambda stamp, messages: matched_samples.append(tuple(messages.values())),
+ queue_size=queue_size,
+ tolerance=tolerance,
+ )
+ return synchronizer, matched_samples
+
+
+def _publish_sample(synchronizer: SampleSynchronizer, stamp: tuple[int, int], topics=None) -> dict:
+ """Add one message per given input (all of them by default), all stamped with the same time."""
+ messages = {topic: _message(stamp) for topic in topics or _TOPICS}
+ for topic, message in messages.items():
+ synchronizer.add(topic, message, stamp)
+ return messages
+
+
+def _later(stamp: tuple[int, int], milliseconds: int) -> tuple[int, int]:
+ """Shift a stamp by the given milliseconds."""
+ nanoseconds = stamp[0] * 1_000_000_000 + stamp[1] + milliseconds * 1_000_000
+ return divmod(nanoseconds, 1_000_000_000)
+
+
+class TestParseSampleIds:
+ """Tests parsing the 'sample_ids' parameter into the sample IDs to request."""
+
+ def test_parses_comma_separated_ids(self):
+ """Sample IDs are parsed in the given order."""
+ assert parse_sample_ids("0,10,20") == [0, 10, 20]
+
+ def test_parses_single_id(self):
+ """A single ID is a valid request."""
+ assert parse_sample_ids("7") == [7]
+
+ def test_ignores_surrounding_whitespace(self):
+ """Sample IDs separated by ', ' are parsed like IDs separated by ','."""
+ assert parse_sample_ids(" 1, 2 ,3 ") == [1, 2, 3]
+
+ @pytest.mark.parametrize("sample_ids", ["", " ", ","])
+ def test_no_ids_evaluate_the_whole_dataset(self, sample_ids):
+ """An empty value requests no specific samples, so the whole dataset is evaluated."""
+ assert parse_sample_ids(sample_ids) == []
+
+ @pytest.mark.parametrize("sample_ids", ["1;2", "first", "1.5", "1-2"])
+ def test_rejects_values_that_are_no_ids(self, sample_ids):
+ """A value that is no comma-separated list of IDs is rejected."""
+ with pytest.raises(ValueError):
+ parse_sample_ids(sample_ids)
+
+
+class TestSampleSynchronizer:
+ """Tests matching the messages of the evaluation inputs into the samples to evaluate."""
+
+ def test_matches_the_messages_of_a_sample_in_input_order(self):
+ """A sample is reported once every input has been received, in the order of the inputs."""
+ synchronizer, matched_samples = _synchronizer()
+
+ messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=list(reversed(_TOPICS)))
+
+ assert matched_samples == [tuple(messages[topic] for topic in _TOPICS)]
+ assert not synchronizer.incomplete_samples
+
+ def test_waits_for_the_missing_inputs_of_a_sample(self):
+ """A sample of which an input is missing is not reported yet."""
+ synchronizer, matched_samples = _synchronizer()
+
+ _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"])
+
+ assert matched_samples == []
+
+ def test_matches_samples_of_a_scene_recorded_before_the_previous_scene(self):
+ """The dataset continues with an older scene, whose samples are matched all the same."""
+ synchronizer, matched_samples = _synchronizer()
+
+ _publish_sample(synchronizer, _PREVIOUS_SCENE)
+ messages = _publish_sample(synchronizer, _NEXT_SCENE[0])
+
+ assert len(matched_samples) == 2
+ assert matched_samples[-1] == tuple(messages[topic] for topic in _TOPICS)
+
+ def test_keeps_a_waiting_sample_older_than_a_matched_one(self):
+ """A sample of a new, older scene is not dropped by a late sample of the previous scene."""
+ synchronizer, matched_samples = _synchronizer()
+ # the last sample of the previous scene still waits for the system under test, while the
+ # first sample of the next scene, recorded years earlier, is published already
+ pending = _publish_sample(synchronizer, _PREVIOUS_SCENE, topics=["label", "label_meta_info"])
+ messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"])
+
+ synchronizer.add("prediction", _message(_PREVIOUS_SCENE), _PREVIOUS_SCENE)
+ synchronizer.add("prediction", _message(_NEXT_SCENE[0]), _NEXT_SCENE[0])
+
+ assert len(matched_samples) == 2
+ assert matched_samples[0][1:] == (pending["label"], pending["label_meta_info"])
+ assert matched_samples[1][1:] == (messages["label"], messages["label_meta_info"])
+
+ def test_gives_up_on_the_sample_waiting_the_longest(self):
+ """Samples are dropped in the order they arrived, not by their stamp."""
+ synchronizer, matched_samples = _synchronizer(queue_size=2)
+
+ # a sample of the previous scene waits first, followed by two samples of the older scene
+ for stamp in [_PREVIOUS_SCENE, *_NEXT_SCENE[:2]]:
+ _publish_sample(synchronizer, stamp, topics=["label"])
+
+ # the queue size is exceeded, so the sample that has been waiting the longest is given up
+ # on, even though the samples kept for evaluation were recorded years before it
+ assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[:2]
+
+ _publish_sample(synchronizer, _NEXT_SCENE[2], topics=["label"])
+
+ assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[1:]
+ assert matched_samples == []
+
+ def test_reports_the_stamp_of_the_first_input_and_the_messages_by_input(self):
+ """The sample is identified by its message of the first input, whichever message arrived first."""
+ reported: list = []
+ synchronizer = SampleSynchronizer(
+ ["trajectory", "objects"], callback=lambda stamp, messages: reported.append((stamp, messages)), tolerance=0.05
+ )
+ objects, trajectory = _message(_NEXT_SCENE[0]), _message(_later(_NEXT_SCENE[0], 20))
+
+ synchronizer.add("objects", objects, _NEXT_SCENE[0])
+ synchronizer.add("trajectory", trajectory, _later(_NEXT_SCENE[0], 20))
+
+ assert reported == [(_later(_NEXT_SCENE[0], 20), {"trajectory": trajectory, "objects": objects})]
+
+ def test_a_single_input_reports_every_message_as_a_sample(self):
+ """An evaluation of a single topic evaluates each of its messages on its own."""
+ synchronizer, matched_samples = _synchronizer(topics=["ego_data"])
+
+ for stamp in _NEXT_SCENE:
+ _publish_sample(synchronizer, stamp, topics=["ego_data"])
+
+ assert len(matched_samples) == len(_NEXT_SCENE)
+
+ def test_matches_messages_of_different_stamps_within_the_tolerance(self):
+ """Topics a simulation publishes at slightly different times are matched within the tolerance."""
+ for tolerance, num_matched_samples in [(0.0, 0), (0.05, 1)]:
+ synchronizer, matched_samples = _synchronizer(topics=["ego_data", "objects"], tolerance=tolerance)
+
+ _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["ego_data"])
+ _publish_sample(synchronizer, _later(_NEXT_SCENE[0], 30), topics=["objects"])
+
+ assert len(matched_samples) == num_matched_samples
+
+ def test_does_not_match_messages_beyond_the_tolerance(self):
+ """Messages whose stamps differ by more than the tolerance belong to different samples."""
+ synchronizer, matched_samples = _synchronizer(topics=["ego_data", "objects"], tolerance=0.05)
+
+ _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["ego_data"])
+ _publish_sample(synchronizer, _later(_NEXT_SCENE[0], 60), topics=["objects"])
+
+ assert matched_samples == []
+ assert len(synchronizer.incomplete_samples) == 2
+
+ def test_matches_the_closest_message_of_a_topic_published_at_a_higher_rate(self):
+ """A message of a slower topic joins the message of a faster topic whose stamp is closest to its own."""
+ synchronizer, matched_samples = _synchronizer(topics=["ego_data", "objects"], tolerance=0.05)
+ ego_data = {
+ offset: _publish_sample(synchronizer, _later(_NEXT_SCENE[0], offset), topics=["ego_data"])
+ for offset in range(0, 100, 10)
+ }
+
+ objects = _publish_sample(synchronizer, _later(_NEXT_SCENE[0], 42), topics=["objects"])
+
+ assert matched_samples == [(ego_data[40]["ego_data"], objects["objects"])]
+
+
+class _FakeLogger:
+ """Collect the messages the node logs instead of publishing them to ROS."""
+
+ def __init__(self):
+ """Start with an empty log."""
+ self.messages: list = []
+
+ def info(self, message: str, **kwargs) -> None:
+ """Record a logged message, whatever its severity."""
+ self.messages.append(message)
+
+ debug = warn = error = info
+
+
+class _FakeEvaluationHandler:
+ """Stand in for the evaluation whose results the node aggregates and writes."""
+
+ def __init__(self):
+ """Start without finalized or written results."""
+ self.finalized_complete = None
+ self.written_results = None
+
+ def finalize(self, complete: bool = True) -> dict:
+ """Report results that are marked the way the node asked for."""
+ self.finalized_complete = complete
+ return {"num_samples": 2, "num_scenes": 1, "complete": complete, "aggregated_metrics": {}}
+
+ def save_results(self, output_path: str, results: dict = None) -> str:
+ """Keep the results instead of writing them to a file."""
+ self.written_results = results
+ return output_path
+
+
+class _RecordingEvaluationHandler:
+ """Stand in for the evaluation that records the samples the node evaluates."""
+
+ def __init__(self):
+ """Start without recorded samples."""
+ self.recorded_samples: list = []
+
+ def record_sample(self, sample_id: str, **messages) -> dict:
+ """Record a sample the way the evaluation stores it, still without a scene."""
+ entry = {"sample_id": sample_id, "scene_id": None, "metrics": {}}
+ self.recorded_samples.append(entry)
+ self.recorded_messages = messages
+ return entry
+
+
+def _evaluating_node(requests_samples: bool) -> SimpleNamespace:
+ """Stub the node state that evaluating a sample reads, counting the attempts to continue the evaluation."""
+ node = SimpleNamespace(
+ requests_samples=requests_samples,
+ evaluation_handler=_RecordingEvaluationHandler(),
+ num_evaluated_samples=0,
+ scenes_awaiting_sample=deque(),
+ samples_awaiting_scene=deque(),
+ evaluation_timeout=60.0,
+ evaluation_deadline=None,
+ results_path="",
+ visualization_publishers={},
+ num_advances=0,
+ get_logger=lambda logger=_FakeLogger(): logger,
+ )
+ node.advance_evaluation = lambda: setattr(node, "num_advances", node.num_advances + 1)
+ return node
+
+
+def _sample_messages(stamp: tuple[int, int]) -> dict:
+ """Fake the synchronized input messages of one sample, by input name."""
+ return {topic: _message(stamp) for topic in _TOPICS}
+
+
+class TestEvaluateSample:
+ """Tests evaluating a sample with the samples requested from the dataset or published by others."""
+
+ def test_passes_the_messages_by_input_name_and_identifies_the_sample_by_its_stamp(self):
+ """The evaluation receives each message by the name of its input."""
+ node = _evaluating_node(requests_samples=True)
+ messages = _sample_messages(_NEXT_SCENE[0])
+
+ AutonomyEvaluation.evaluate_sample(node, _NEXT_SCENE[0], messages)
+
+ assert node.evaluation_handler.recorded_messages == messages
+ assert node.evaluation_handler.recorded_samples[0]["sample_id"] == "1531885320.049418000"
+
+ def test_evaluation_requests_the_next_samples_after_evaluating_one(self):
+ """With the dataset as sample source, the evaluation continues with the next samples on its own."""
+ node = _evaluating_node(requests_samples=True)
+
+ AutonomyEvaluation.evaluate_sample(node, _NEXT_SCENE[0], _sample_messages(_NEXT_SCENE[0]))
+
+ assert node.num_advances == 1
+ # the dataset reports the scene of the sample with the response to the request
+ assert list(node.samples_awaiting_scene) == node.evaluation_handler.recorded_samples
+
+ def test_external_sample_source_leaves_publishing_samples_to_others(self):
+ """With an external sample source, samples are evaluated as they arrive, without requesting further ones."""
+ node = _evaluating_node(requests_samples=False)
+
+ for stamp in _NEXT_SCENE:
+ AutonomyEvaluation.evaluate_sample(node, stamp, _sample_messages(stamp))
+
+ assert node.num_evaluated_samples == len(_NEXT_SCENE)
+ assert node.num_advances == 0
+ # the scenes are only reported to whoever published the samples, so no sample waits for one
+ assert not node.samples_awaiting_scene
+
+
+class _FakeRequestClient:
+ """Stand in for the client of the sample request service of the dataset."""
+
+ def __init__(self, ready: bool):
+ """Start with the service available or not."""
+ self.ready = ready
+ self.removed_requests: list = []
+
+ def service_is_ready(self) -> bool:
+ """Report whether the dataset node offers the service."""
+ return self.ready
+
+ def remove_pending_request(self, future) -> None:
+ """Record a request that is given up on."""
+ self.removed_requests.append(future)
+
+
+def _advancing_node(service_ready: bool, dataset_available: bool = True, publishing_finished: bool = False):
+ """Stub the node state that advancing the evaluation reads, recording what the evaluation does next."""
+ node = SimpleNamespace(
+ evaluation_finished=False,
+ pending_request=None,
+ publishing_finished=publishing_finished,
+ dataset_available=dataset_available,
+ dataset_unavailable_since=None,
+ sample_request_client=_FakeRequestClient(ready=service_ready),
+ sample_request_service="/datasets/request_samples",
+ actions=[],
+ get_logger=lambda logger=_FakeLogger(): logger,
+ )
+ node.track_dataset_availability = lambda: AutonomyEvaluation.track_dataset_availability(node)
+ node.awaiting_evaluations = lambda: False
+ node.request_samples = lambda: node.actions.append("request")
+ node.finalize_evaluation = lambda: node.actions.append("finalize")
+ node.shutdown = lambda: node.actions.append("shutdown")
+ return node
+
+
+class TestAdvanceEvaluation:
+ """Tests requesting further samples and finishing the evaluation, with and without the dataset node."""
+
+ @pytest.fixture
+ def clock(self, monkeypatch) -> list:
+ """Control the steady clock the node measures how long the dataset has been gone with."""
+ now = [100.0]
+ monkeypatch.setattr(autonomy_evaluation.time, "monotonic", lambda: now[0])
+ return now
+
+ def test_requests_samples_while_the_dataset_is_available(self):
+ """The next samples are requested as soon as the previous ones have been evaluated."""
+ node = _advancing_node(service_ready=True, dataset_available=False)
+
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.actions == ["request"]
+ assert node.dataset_available
+
+ def test_waits_for_a_dataset_that_has_not_started_yet(self, clock):
+ """A dataset node that never offered its service is waited for, however long it takes."""
+ node = _advancing_node(service_ready=False, dataset_available=False)
+
+ clock[0] += 3600.0
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.actions == []
+ assert not node.publishing_finished
+
+ def test_finishes_after_the_dataset_reported_its_end_and_shut_down(self):
+ """The service is only needed to request samples, not to finish once publishing has ended."""
+ node = _advancing_node(service_ready=False, publishing_finished=True)
+
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.actions == ["finalize", "shutdown"]
+
+ def test_finishes_once_the_dataset_shut_down_after_its_last_sample(self, clock):
+ """A dataset gone for longer than the grace period publishes no further samples."""
+ node = _advancing_node(service_ready=False)
+
+ AutonomyEvaluation.advance_evaluation(node)
+ clock[0] += autonomy_evaluation._DATASET_SHUTDOWN_GRACE_PERIOD_S / 2
+ AutonomyEvaluation.advance_evaluation(node)
+
+ # within the grace period, a response the dataset sent before shutting down may still arrive
+ assert node.actions == []
+
+ clock[0] += autonomy_evaluation._DATASET_SHUTDOWN_GRACE_PERIOD_S
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.actions == ["finalize", "shutdown"]
+
+ def test_continues_with_a_dataset_that_is_back_within_the_grace_period(self, clock):
+ """A service that is briefly unavailable does not end the evaluation."""
+ node = _advancing_node(service_ready=False)
+ AutonomyEvaluation.advance_evaluation(node)
+ clock[0] += autonomy_evaluation._DATASET_SHUTDOWN_GRACE_PERIOD_S / 2
+
+ node.sample_request_client.ready = True
+ AutonomyEvaluation.advance_evaluation(node)
+ clock[0] += autonomy_evaluation._DATASET_SHUTDOWN_GRACE_PERIOD_S
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.actions == ["request", "request"]
+ assert not node.publishing_finished
+
+ def test_gives_up_on_a_request_the_shut_down_dataset_did_not_answer(self, clock):
+ """A request pending when the dataset shut down would never be answered."""
+ node = _advancing_node(service_ready=False)
+ request = node.pending_request = object()
+
+ AutonomyEvaluation.advance_evaluation(node)
+ clock[0] += 2 * autonomy_evaluation._DATASET_SHUTDOWN_GRACE_PERIOD_S
+ AutonomyEvaluation.advance_evaluation(node)
+
+ assert node.sample_request_client.removed_requests == [request]
+ assert node.pending_request is None
+ assert node.actions == ["finalize", "shutdown"]
+
+
+class TestReceiveMessage:
+ """Tests stamping the received messages before they are matched into samples."""
+
+ @staticmethod
+ def _receiving_node(now: tuple[int, int]) -> tuple[SimpleNamespace, list]:
+ """Stub the node state that receiving a message reads, collecting the stamped messages."""
+ received: list = []
+ node = SimpleNamespace(
+ message_synchronizer=SimpleNamespace(add=lambda name, message, stamp: received.append((name, message, stamp))),
+ get_clock=lambda: SimpleNamespace(now=lambda: SimpleNamespace(seconds_nanoseconds=lambda: now)),
+ )
+ return node, received
+
+ def test_stamps_a_message_with_its_header_stamp(self):
+ """A message is matched by the stamp its publisher gave it."""
+ node, received = self._receiving_node(now=_PREVIOUS_SCENE)
+ message = _message(_NEXT_SCENE[0])
+
+ AutonomyEvaluation.receive_message(node, "prediction", message)
+
+ assert received == [("prediction", message, _NEXT_SCENE[0])]
+
+ def test_stamps_a_message_without_header_on_reception(self):
+ """A message without header, e.g. a std_msgs/Bool, is stamped with the time it is received at."""
+ node, received = self._receiving_node(now=_PREVIOUS_SCENE)
+ message = SimpleNamespace(data=True)
+
+ AutonomyEvaluation.receive_message(node, "collision", message)
+
+ assert received == [("collision", message, _PREVIOUS_SCENE)]
+
+
+class TestInputTopic:
+ """Tests the topics the inputs of an evaluation are subscribed on."""
+
+ _DERIVED_TOPICS = {"label_meta_info": ("label", "/meta_info")}
+
+ @staticmethod
+ def _node(remappings: dict) -> SimpleNamespace:
+ """Stub the topic resolution of a node started with the given remappings of its relative names."""
+
+ def resolve_topic_name(topic: str, only_expand: bool = False) -> str:
+ return f"/{topic}" if only_expand or topic not in remappings else remappings[topic]
+
+ return SimpleNamespace(resolve_topic_name=resolve_topic_name)
+
+ def test_subscribes_an_input_on_its_name(self):
+ """An input is subscribed on its node-relative name, which remappings redirect."""
+ node = self._node({"label": "/object_list/lidar_01"})
+
+ assert AutonomyEvaluation.input_topic(node, "label", self._DERIVED_TOPICS) == "label"
+
+ def test_derived_input_follows_the_topic_of_its_source(self):
+ """Meta information is subscribed next to the remapped topic of the object list it belongs to."""
+ node = self._node({"label": "/object_list/lidar_01"})
+
+ topic = AutonomyEvaluation.input_topic(node, "label_meta_info", self._DERIVED_TOPICS)
+
+ assert topic == "/object_list/lidar_01/meta_info"
+
+ def test_remapped_derived_input_keeps_its_own_topic(self):
+ """A derived input that is remapped itself is subscribed where it is remapped to."""
+ node = self._node({"label": "/object_list/lidar_01", "label_meta_info": "/meta_info"})
+
+ assert AutonomyEvaluation.input_topic(node, "label_meta_info", self._DERIVED_TOPICS) == "label_meta_info"
+
+
+def _node(num_evaluated_samples: int = 2, results_path: str = "/results/evaluation.json") -> SimpleNamespace:
+ """Stub the node state that finalizing an evaluation reads, without initializing ROS."""
+ return SimpleNamespace(
+ evaluation="counting",
+ evaluation_finished=False,
+ evaluation_handler=_FakeEvaluationHandler(),
+ num_evaluated_samples=num_evaluated_samples,
+ results_path=results_path,
+ request_timer=SimpleNamespace(cancel=lambda: None),
+ get_logger=lambda logger=_FakeLogger(): logger,
+ )
+
+
+class TestFinalizeEvaluation:
+ """Tests reporting the results of an evaluation that ran to its end or was interrupted."""
+
+ def test_finished_evaluation_writes_complete_results(self):
+ """An evaluation that evaluated all its samples reports complete results."""
+ node = _node()
+
+ AutonomyEvaluation.finalize_evaluation(node)
+
+ assert node.evaluation_handler.finalized_complete is True
+ assert node.evaluation_handler.written_results["complete"] is True
+
+ def test_interrupted_evaluation_writes_incomplete_results(self):
+ """An evaluation stopped before its last sample, e.g. with Ctrl-C, still writes its results."""
+ node = _node()
+
+ AutonomyEvaluation.finalize_evaluation(node, complete=False)
+
+ assert node.evaluation_handler.finalized_complete is False
+ assert node.evaluation_handler.written_results["complete"] is False
+
+ def test_finished_evaluation_is_not_finalized_again_on_shutdown(self):
+ """Shutting down after the last sample must not overwrite the results with incomplete ones."""
+ node = _node()
+ AutonomyEvaluation.finalize_evaluation(node)
+
+ AutonomyEvaluation.finalize_evaluation(node, complete=False)
+
+ assert node.evaluation_handler.finalized_complete is True
+ assert node.evaluation_handler.written_results["complete"] is True
+
+ def test_manual_playback_writes_its_results_when_stopped(self):
+ """With manual playback, which runs no request timer, the results are written once the node is stopped."""
+ node = _node()
+ node.request_timer = None
+
+ AutonomyEvaluation.finalize_evaluation(node, complete=False)
+
+ assert node.evaluation_handler.written_results["complete"] is False
+
+ def test_interrupted_evaluation_without_samples_writes_nothing(self):
+ """An evaluation interrupted before its first sample has no results to write."""
+ node = _node(num_evaluated_samples=0)
+
+ AutonomyEvaluation.finalize_evaluation(node, complete=False)
+
+ assert node.evaluation_handler.written_results is None
+
+ def test_results_are_only_logged_without_a_results_path(self):
+ """Without 'results_path' the interrupted results are logged instead of written."""
+ node = _node(results_path="")
+
+ AutonomyEvaluation.finalize_evaluation(node, complete=False)
+
+ assert node.evaluation_handler.finalized_complete is False
+ assert node.evaluation_handler.written_results is None
diff --git a/autonomy_evaluation/tests/test_launch.py b/autonomy_evaluation/tests/test_launch.py
new file mode 100644
index 0000000..d8de268
--- /dev/null
+++ b/autonomy_evaluation/tests/test_launch.py
@@ -0,0 +1,54 @@
+# Copyright Thinking Cars GmbH
+# SPDX-License-Identifier: Apache-2.0
+
+"""Tests for the remapping of the topics of the selected evaluation in the launch file.
+
+The launch file does not know the topics of an evaluation in advance, so it remaps whichever topics
+the selected evaluation reads onto the launch arguments of their names.
+"""
+
+from __future__ import annotations
+
+import importlib.util
+from pathlib import Path
+
+from autonomy_evaluation.evaluations import load_evaluation
+
+_LAUNCH_FILE = Path(__file__).resolve().parents[1] / "launch" / "autonomy_evaluation.launch.py"
+
+
+def _launch_module():
+ """Import the launch file, which is no module of the package."""
+ spec = importlib.util.spec_from_file_location("autonomy_evaluation_launch", _LAUNCH_FILE)
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module
+
+
+class TestInputRemappings:
+ """Tests remapping the topics of an evaluation onto the given launch arguments."""
+
+ def setup_method(self):
+ """Select the nuScenes evaluation, reading a prediction and a label with its meta information."""
+ self.input_remappings = _launch_module().input_remappings
+ self.evaluation = load_evaluation("nuscenes_lidar_object_detection")
+
+ def test_remaps_the_topics_given_as_launch_arguments(self):
+ """A topic given by a launch argument of its name is subscribed there."""
+ remappings = self.input_remappings(self.evaluation, {"prediction": "/object_list/prediction", "name": "evaluation"})
+
+ assert ("prediction", "/object_list/prediction") in remappings
+
+ def test_subscribes_topics_without_launch_argument_in_the_private_namespace(self):
+ """A topic that is not given is subscribed in the private namespace of the node."""
+ remappings = self.input_remappings(self.evaluation, {})
+
+ assert remappings == [("prediction", "~/prediction"), ("label", "~/label")]
+
+ def test_leaves_a_derived_topic_to_the_node_unless_it_is_given(self):
+ """The node derives the meta information topic from the label topic, unless it is given itself."""
+ assert "label_meta_info" not in dict(self.input_remappings(self.evaluation, {"label": "/object_list/lidar_01"}))
+
+ remappings = self.input_remappings(self.evaluation, {"label_meta_info": "/meta_info"})
+
+ assert ("label_meta_info", "/meta_info") in remappings
diff --git a/autonomy_benchmarks/autonomy_benchmarks/__init__.py b/autonomy_evaluation/tests/utils/__init__.py
similarity index 57%
rename from autonomy_benchmarks/autonomy_benchmarks/__init__.py
rename to autonomy_evaluation/tests/utils/__init__.py
index ebbf0e0..2f60957 100644
--- a/autonomy_benchmarks/autonomy_benchmarks/__init__.py
+++ b/autonomy_evaluation/tests/utils/__init__.py
@@ -1,4 +1,4 @@
# Copyright Thinking Cars GmbH
# SPDX-License-Identifier: Apache-2.0
-"""Autonomy benchmarks package."""
+"""Utility test modules for autonomy_evaluation."""
diff --git a/autonomy_benchmarks/tests/utils/test_ObjectDetectionUtils.py b/autonomy_evaluation/tests/utils/test_ObjectDetectionUtils.py
similarity index 99%
rename from autonomy_benchmarks/tests/utils/test_ObjectDetectionUtils.py
rename to autonomy_evaluation/tests/utils/test_ObjectDetectionUtils.py
index be2fcdd..aeb3b64 100644
--- a/autonomy_benchmarks/tests/utils/test_ObjectDetectionUtils.py
+++ b/autonomy_evaluation/tests/utils/test_ObjectDetectionUtils.py
@@ -7,7 +7,7 @@
import numpy as np
import pytest
-from autonomy_benchmarks.utils.ObjectDetectionUtils import ObjectDetectionUtils
+from autonomy_evaluation.utils.ObjectDetectionUtils import ObjectDetectionUtils
def _box2d(x1, y1, x2, y2):
diff --git a/deployment/compose/docker-compose.yml b/deployment/compose/docker-compose.yml
index 70c08a2..4eac091 100644
--- a/deployment/compose/docker-compose.yml
+++ b/deployment/compose/docker-compose.yml
@@ -1,18 +1,19 @@
services:
- autonomy-benchmarks:
- image: ghcr.io/thinking-cars/autonomy_benchmarks:v1.0.0
+ autonomy-evaluation:
+ image: ghcr.io/thinking-cars/autonomy_evaluation:v1.0.0
environment:
# --- name ------
NAMESPACE: /
- NAME: autonomy_benchmarks
+ NAME: autonomy_evaluation
# --- other -----
REQUEST_SAMPLES: ~/request_samples
LOG_LEVEL: ${LOG_LEVEL:-info}
USE_SIM_TIME: ${USE_SIM_TIME:-true}
- BENCHMARK: nuscenes_lidar_object_detection
+ EVALUATION: nuscenes_lidar_object_detection
VISUALIZE: false
- MANUAL_PLAYBACK: false
+ SAMPLE_SOURCE: dataset
+ SYNC_TOLERANCE: 0.0
SAMPLES_PER_REQUEST: 1
SAMPLE_IDS: ""
EVALUATION_TIMEOUT: 60.0
@@ -21,14 +22,15 @@ services:
- /bin/bash
- -ic
- |
- ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \
+ ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
namespace:=$${NAMESPACE} \
name:=$${NAME} \
log_level:=$${LOG_LEVEL} \
use_sim_time:=$${USE_SIM_TIME} \
- benchmark:=$${BENCHMARK} \
+ evaluation:=$${EVALUATION} \
visualize:=$${VISUALIZE} \
- manual_playback:=$${MANUAL_PLAYBACK} \
+ sample_source:=$${SAMPLE_SOURCE} \
+ sync_tolerance:=$${SYNC_TOLERANCE} \
samples_per_request:=$${SAMPLES_PER_REQUEST} \
$${SAMPLE_IDS:+sample_ids:=$${SAMPLE_IDS}} \
evaluation_timeout:=$${EVALUATION_TIMEOUT} \
diff --git a/deployment/helm/Chart.yaml b/deployment/helm/Chart.yaml
index db887d2..372bb10 100644
--- a/deployment/helm/Chart.yaml
+++ b/deployment/helm/Chart.yaml
@@ -1,8 +1,8 @@
apiVersion: v2
-name: autonomy-benchmarks
+name: autonomy-evaluation
version: 1.0.0
appVersion: 1.0.0
-description: Benchmarking suite for automated driving tasks
+description: Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments
dependencies:
- repository: oci://ghcr.io/openads-project/openads-helm
name: openadservice
diff --git a/deployment/helm/values.yaml b/deployment/helm/values.yaml
index f917b77..35f3927 100644
--- a/deployment/helm/values.yaml
+++ b/deployment/helm/values.yaml
@@ -1,17 +1,18 @@
openadservice:
- name: autonomy-benchmarks
- image: ghcr.io/thinking-cars/autonomy_benchmarks:v1.0.0
+ name: autonomy-evaluation
+ image: ghcr.io/thinking-cars/autonomy_evaluation:v1.0.0
env:
# --- name ------
NAMESPACE: /
- NAME: autonomy_benchmarks
+ NAME: autonomy_evaluation
# --- other -----
REQUEST_SAMPLES: ~/request_samples
LOG_LEVEL: info
USE_SIM_TIME: true
- BENCHMARK: nuscenes_lidar_object_detection
+ EVALUATION: nuscenes_lidar_object_detection
VISUALIZE: false
- MANUAL_PLAYBACK: false
+ SAMPLE_SOURCE: dataset
+ SYNC_TOLERANCE: 0.0
SAMPLES_PER_REQUEST: 1
SAMPLE_IDS:
EVALUATION_TIMEOUT: 60.0
@@ -20,14 +21,15 @@ openadservice:
- /bin/bash
- -ic
- |
- ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \
+ ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
namespace:=${NAMESPACE} \
name:=${NAME} \
log_level:=${LOG_LEVEL} \
use_sim_time:=${USE_SIM_TIME} \
- benchmark:=${BENCHMARK} \
+ evaluation:=${EVALUATION} \
visualize:=${VISUALIZE} \
- manual_playback:=${MANUAL_PLAYBACK} \
+ sample_source:=${SAMPLE_SOURCE} \
+ sync_tolerance:=${SYNC_TOLERANCE} \
samples_per_request:=${SAMPLES_PER_REQUEST} \
${SAMPLE_IDS:+sample_ids:=${SAMPLE_IDS}} \
evaluation_timeout:=${EVALUATION_TIMEOUT} \
diff --git a/docker-compose.dev.yml b/docker-compose.dev.yml
index 63075d6..2f54860 100644
--- a/docker-compose.dev.yml
+++ b/docker-compose.dev.yml
@@ -1,12 +1,12 @@
-name: ${USER}-autonomy-benchmarks
+name: ${USER}-autonomy-evaluation
services:
- autonomy_benchmarks:
+ autonomy_evaluation:
extends:
file: docker-compose.yml
- service: autonomy_benchmarks
- image: ghcr.io/thinking-cars/autonomy_benchmarks:latest-dev
+ service: autonomy_evaluation
+ image: ghcr.io/thinking-cars/autonomy_evaluation:latest-dev
command: sleep infinity
volumes:
- .:/docker-ros/ws/src/target
diff --git a/docker-compose.yml b/docker-compose.yml
index 38a5097..dc3392c 100644
--- a/docker-compose.yml
+++ b/docker-compose.yml
@@ -1,15 +1,15 @@
-name: ${USER}-autonomy-benchmarks
+name: ${USER}-autonomy-evaluation
services:
- autonomy_benchmarks:
- image: ghcr.io/thinking-cars/autonomy_benchmarks:latest
+ autonomy_evaluation:
+ image: ghcr.io/thinking-cars/autonomy_evaluation:latest
command:
- /bin/bash
- -ic
- |
- ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \
- benchmark:=nuscenes_lidar_object_detection \
+ ros2 launch autonomy_evaluation autonomy_evaluation.launch.py \
+ evaluation:=nuscenes_lidar_object_detection \
prediction:=/object_list/prediction \
label:=/object_list/lidar_01 \
request_samples:=/datasets/request_samples \
@@ -49,7 +49,7 @@ services:
- ./config/params_nuscenes.yml:/params_nuscenes.yml
- $DATASET_DIR:/datasets
depends_on:
- - autonomy_benchmarks
+ - autonomy_evaluation
autoware_lidar_centerpoint:
profiles:
diff --git a/docs/IMPLEMENTATION.md b/docs/IMPLEMENTATION.md
index 741c484..7108049 100644
--- a/docs/IMPLEMENTATION.md
+++ b/docs/IMPLEMENTATION.md
@@ -1,6 +1,6 @@
# Implementation Details
-This repository supports the following benchmarks for object detection in automated driving systems:
+This repository supports the following evaluations of object detection in automated driving systems:
- [nuScenes Challenge](#nuscenes-challenge): 3D lidar object detection
@@ -16,7 +16,7 @@ Supported Datasets:
- [nuScenes Dataset](https://github.com/thinking-cars/autonomy_datasets/blob/main/docs/IMPLEMENTATION.md#nuscenes-dataset)
-> This benchmark uses [nuScenes Dataset](https://github.com/thinking-cars/autonomy_datasets/blob/main/docs/IMPLEMENTATION.md#nuscenes-dataset) via the [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) ROS package.
+> This evaluation uses [nuScenes Dataset](https://github.com/thinking-cars/autonomy_datasets/blob/main/docs/IMPLEMENTATION.md#nuscenes-dataset) via the [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) ROS package.
[](https://www.nuscenes.org/object-detection) 
@@ -99,14 +99,42 @@ Metrics are computed based on the following assumptions:
-### Adding more Benchmarks
+### Adding more Evaluations
-To contribute a new benchmark for a dataset or evaluation protocol:
+An evaluation declares the topics it reads and computes metrics from their messages. It may read the topics of the system under test only, e.g. the ego state and the surrounding objects to evaluate a closed-loop planner by its time to collision, or compare them with ground truth, e.g. the predictions of a perception algorithm with the labels of a dataset. The node subscribes to all declared topics, matches their messages into samples by their stamp, and passes the messages of each sample by the names of their topics.
-1. Create a new benchmark class in [autonomy_benchmarks/benchmarks/](../autonomy_benchmarks/autonomy_benchmarks/benchmarks/) that inherits from `AutonomyBenchmark`.
-2. Implement the three abstract methods: `required_inputs()`, `compute_sample_metrics()`, and `compute_aggregated_metrics()`.
-3. Configure the benchmark in `__init__` (thresholds, per-class ranges, and metric rules as instance attributes); keep static lookup tables (e.g. category-to-class mappings) as module-level `_CONSTANT_NAME` constants.
-4. Register the benchmark in the node's handler dispatch in [autonomy_benchmarks.py](../autonomy_benchmarks/autonomy_benchmarks/autonomy_benchmarks.py) so it can be selected via the `benchmark:=` launch argument.
-5. Add comprehensive tests in [tests/benchmarks/](../autonomy_benchmarks/tests/benchmarks/) following existing test patterns.
-6. Update documentation with benchmark details, metrics table, and dataset requirements.
-7. Create a [Pull Request](https://github.com/thinking-cars/autonomy_benchmarks-internal/pulls) on GitHub and wait for maintainer feedback.
+To contribute a new evaluation for a dataset or evaluation protocol:
+
+1. Create a new evaluation class in [autonomy_evaluation/evaluations/](../autonomy_evaluation/autonomy_evaluation/evaluations/) that inherits from `Evaluation`.
+2. Declare the topics it reads, each by the name it is configured with, e.g. `prediction:=/topic`, mapped to its ROS message type:
+ - `required_inputs()`: the topics of the system under test that are evaluated (required).
+ - `required_ground_truth()`: the reference the inputs are compared with, e.g. dataset labels (optional, none by default).
+ - `derived_topics()`: topics published next to another one, e.g. meta information on `/meta_info`, which follow that topic instead of being configured on their own (optional).
+3. Implement `compute_sample_metrics()`, whose parameters are named after the declared topics, and `compute_aggregated_metrics()`. Optionally, implement `visualization_outputs()` and `visualize_sample()` to publish a per-sample visualization for RViz.
+4. Configure the evaluation in `__init__` (thresholds, per-class ranges, and metric rules as instance attributes); keep static lookup tables (e.g. category-to-class mappings) as module-level `_CONSTANT_NAME` constants.
+5. Register the evaluation in `EVALUATIONS` of [registry.py](../autonomy_evaluation/autonomy_evaluation/evaluations/registry.py) so it can be selected via the `evaluation:=` launch argument, add comprehensive tests in [tests/evaluations/](../autonomy_evaluation/tests/evaluations/) following existing test patterns, and update the documentation with evaluation details, metrics table, dataset requirements, and the topics it reads.
+6. Create a [Pull Request](https://github.com/thinking-cars/autonomy_evaluation/pulls) on GitHub and wait for maintainer feedback.
+
+An evaluation that is specific to a system under test can also stay in a package of its own. The node loads it by its module and class, e.g. `evaluation:=my_package.evaluations:TimeToCollision`, as long as the module is importable in the environment of the node:
+
+```python
+from typing import Any, Dict, List, Optional
+
+from autonomy_evaluation.evaluations import Evaluation
+from perception_msgs.msg import EgoData, ObjectList
+
+
+class TimeToCollision(Evaluation):
+ def __init__(self) -> None:
+ super().__init__(name="time_to_collision", description="time to collision of a closed-loop planner")
+
+ def required_inputs(self) -> Dict[str, Any]:
+ # inputs only: the metric needs no ground truth
+ return {"ego_data": EgoData, "objects": ObjectList}
+
+ def compute_sample_metrics(self, ego_data: EgoData, objects: ObjectList, sample_id: Optional[str] = None) -> Dict[str, Any]:
+ return {"ttc": ...}
+
+ def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Dict[str, Any]:
+ return {"min_ttc": min(entry["metrics"]["ttc"] for entry in sample_results)}
+```