diff --git a/.devcontainer/devcontainer.json b/.devcontainer/devcontainer.json index f1baedb..5a81e51 100644 --- a/.devcontainer/devcontainer.json +++ b/.devcontainer/devcontainer.json @@ -5,7 +5,8 @@ // attach to container from docker compose setup "dockerComposeFile": ["../docker-compose.dev.yml"], - "service": "autonomy_benchmarks", + "service": "autonomy_evaluation", + "shutdownAction": "none", // set up git pre-commit hooks "onCreateCommand": "pip install pre-commit && apt update && apt install locales -y && locale-gen en_US.UTF-8", diff --git a/.github/workflows/docker-ros.yml b/.github/workflows/docker-ros.yml index 82f4d8d..bd99cc8 100644 --- a/.github/workflows/docker-ros.yml +++ b/.github/workflows/docker-ros.yml @@ -22,7 +22,7 @@ jobs: target: dev,run base-image: rwthika/ros2:jazzy rmw-implementation: rmw_zenoh_cpp - command: ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py + command: ros2 launch autonomy_evaluation autonomy_evaluation.launch.py enable-slim: false env: TARGET_CMAKE_ARGS: -DCMAKE_EXPORT_COMPILE_COMMANDS=1 diff --git a/README.md b/README.md index a008294..c528cf5 100644 --- a/README.md +++ b/README.md @@ -1,36 +1,37 @@ -# autonomy_benchmarks +# autonomy_evaluation

- - + +
- - - - - + + + + +

-> This repository will be part of the **Autonomy.Hub Ecosystem** +> This repository is part of the **Autonomy.Benchmarks** suite of the **Autonomy.Hub Ecosystem** -As part of the Autonomy.Hub Ecosystem, **Autonomy.Benchmarks** enables the Automated Driving community to easily benchmark their automated driving building blocks across different tasks and datasets: +Within the Autonomy.Benchmarks suite, **Autonomy.Evaluation** generates the metrics-based evidence for benchmarking automated driving deployments. It evaluates arbitrary ROS systems under test, either from their own topics alone, e.g. a closed-loop planner by its time to collision, or against ground truth, e.g. a perception algorithm against the labels of a dataset replayed by [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets), and reports the resulting metrics per scene and over all evaluated samples: -- 🔄 **Unified ROS 2 Interface**: Work with multiple datasets using the benefits of the ROS 2 ecosystem -- 📊 **Comprehensive Benchmarks**: Use the provided benchmarks with [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) to benchmark building blocks across different automated driving tasks +- 🔄 **Unified ROS 2 Interface**: Evaluate any ROS system under test, on datasets, in simulation or live, using the benefits of the ROS 2 ecosystem +- 🧩 **Pluggable Evaluations**: Select an evaluation by name, or bring your own from another package, reading any topics as inputs and, where needed, ground truth +- 📊 **Established Metrics**: Use the provided evaluations, which follow the protocols of established challenges, with [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) across different automated driving tasks - ⚡ **Efficient Data Pipeline**: Works seamlessly with preprocessed Rosbag files from [Autonomy.Datasets](https://github.com/thinking-cars/autonomy_datasets) for fast execution during development - 🐳 **Dockerized Environment**: Reproducible setup with all dependencies included - 🔌 **Modular Architecture**: Easy integration with other ROS 2 packages -## Supported Benchmarks +## Supported Evaluations -This repository supports various automated driving evaluation benchmarks. +This repository supports evaluations of various automated driving tasks. Detailed metric definitions and computation notes are documented in [docs/IMPLEMENTATION.md](docs/IMPLEMENTATION.md). -> [**Contributions**](docs/IMPLEMENTATION.md#adding-more-benchmarks) adding more benchmarks are welcome +> [**Contributions**](docs/IMPLEMENTATION.md#adding-more-evaluations) adding more evaluations are welcome -| Benchmark | Challenge | Dataset | Task | +| Evaluation | Challenge | Dataset | Task | | --------- | --------- | ------- | ---- | | [**nuScenes 3D Lidar Object Detection**](docs/IMPLEMENTATION.md#3d-lidar-object-detection) | [![3D Object Detection Challenge](https://img.shields.io/badge/origin-3D_Object_Detection_Challenge-green)](https://www.nuscenes.org/object-detection) | [nuScenes](https://github.com/thinking-cars/autonomy_datasets) | 3D bounding box detection from lidar | @@ -43,7 +44,7 @@ Detailed metric definitions and computation notes are documented in [docs/IMPLEM Clone [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) and follow its setup instructions to prepare your dataset. -Use the provided [docker-compose.yml](docker-compose.yml) to start the full pipeline — dataset publisher, system-under-test, and benchmark node: +Use the provided [docker-compose.yml](docker-compose.yml) to start the full pipeline — dataset publisher, system under test, and evaluation node: ```bash # enable GUI output from Docker container @@ -57,13 +58,13 @@ docker compose up -d docker compose down ``` -Configure the benchmark task and dataset via ROS launch arguments in [docker-compose.yml](docker-compose.yml): +Configure the evaluation and dataset via ROS launch arguments in [docker-compose.yml](docker-compose.yml): ```yaml -command: ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py benchmark:=nuscenes_lidar_object_detection prediction:=$your_prediction_topic label:=$your_label_topic request_samples:=/datasets/request_samples visualize:=true +command: ros2 launch autonomy_evaluation autonomy_evaluation.launch.py evaluation:=nuscenes_lidar_object_detection prediction:=$your_prediction_topic label:=$your_label_topic request_samples:=/datasets/request_samples visualize:=true ``` -The benchmark node requests the samples it evaluates from the dataset node via its `request_samples` service, which publishes them and responds once they have been published. The dataset therefore publishes the next sample only once the system under test has processed the current one. As soon as all samples have been published, the benchmark aggregates its metrics per sample, per scene of the dataset and over the whole benchmark. See the [node documentation](autonomy_benchmarks/README.md#autonomy_benchmarks) for the sample request settings and the results. +The evaluation node requests the samples it evaluates from the dataset node via its `request_samples` service, which publishes them and responds once they have been published. The dataset therefore publishes the next sample only once the system under test has processed the current one. As soon as all samples have been published, the node aggregates its metrics per scene of the dataset and over all evaluated samples. To evaluate samples published by others instead, e.g. by a closed-loop simulation, set `sample_source:=external`. See the [node documentation](autonomy_evaluation/README.md#autonomy_evaluation) for the topics of the evaluations, the sample settings and the results. ## 💻 Development @@ -71,11 +72,11 @@ The benchmark node requests the samples it evaluates from the dataset node via i 1. Clone the repository. ```bash - git clone https://github.com/thinking-cars/autonomy_benchmarks.git + git clone https://github.com/thinking-cars/autonomy_evaluation.git ``` 1. Initialize the [`.openads-dev-environment`](https://github.com/openads-project/openads-dev-environment) submodule containing development environment configuration. ```bash - cd autonomy_benchmarks + cd autonomy_evaluation git submodule update --init --recursive ``` 1. Open the repository in [Visual Studio Code](https://code.visualstudio.com). @@ -108,11 +109,11 @@ colcon test-result --verbose ## 📝 Documentation -Package and node interfaces are documented in the respective package READMEs listed below. Implementation details are found in the [Source Code Documentation](https://thinking-cars.github.io/autonomy_benchmarks). +Package and node interfaces are documented in the respective package READMEs listed below. Implementation details are found in the [Source Code Documentation](https://thinking-cars.github.io/autonomy_evaluation). | Package | Description | | --- | --- | -| [autonomy_benchmarks](autonomy_benchmarks/README.md) | Benchmarking suite for automated driving tasks | +| [autonomy_evaluation](autonomy_evaluation/README.md) | Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments | ## ⚖️ Licensing diff --git a/autonomy_benchmarks/README.md b/autonomy_benchmarks/README.md deleted file mode 100644 index 33ac320..0000000 --- a/autonomy_benchmarks/README.md +++ /dev/null @@ -1,86 +0,0 @@ -# `autonomy_benchmarks` - -Benchmarking suite for automated driving tasks - -## Nodes - -### `autonomy_benchmarks` - -The node requests the samples it evaluates from the dataset, using the `request_samples` service -of [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets), which publishes them -and responds once they have been published. By default one sample is requested at a time, so the -dataset only publishes the next sample once the system under test has delivered its output for -the current one and the benchmark has evaluated it. Increase `samples_per_request` to publish -samples in batches, set it to `0` to publish the whole dataset with a single request, or list the -IDs of individual samples in `sample_ids` to evaluate only those. - -Once the dataset reports that all requested samples have been published, the per-sample metrics -are aggregated and reported on three levels: for the whole benchmark (`aggregated_metrics`), for -the samples of each scene the dataset published them from (`scene_results`, matched with the -samples via the `published_scene_ids` of the responses), and for every single sample -(`sample_results`). The dataset metrics are logged, and the results of all three levels are -written to a JSON file if `results_path` is set. Samples that are not evaluated within -`evaluation_timeout` seconds of being published, e.g. because the system under test skipped them, -are left out. - -```bash -ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \ - prediction:=/object_list/prediction \ - label:=/object_list/lidar_01 \ - request_samples:=/datasets/request_samples \ - results_path:=/results/nuscenes_lidar_object_detection.json -``` - -To look at the samples one by one, set `manual_playback` together with `visualize`. The node then requests no samples itself, and the samples are published with the [playback panel](https://github.com/thinking-cars/autonomy_datasets/blob/main/autonomy_datasets_rviz_plugins/README.md) in RViz instead, whose _Service_ field has to name the `request_samples` service of the dataset node (`/datasets/request_samples` by default). Every sample that arrives is evaluated and shown in RViz, and `samples_per_request`, `sample_ids` and `evaluation_timeout` have no effect. As the responses of the dataset only reach the panel, the node neither learns the scenes of the samples, so `scene_results` stays empty, nor when the dataset has ended: the results of the evaluated samples are reported once the node is stopped, e.g. with Ctrl-C, and are marked as incomplete. - -```bash -ros2 launch autonomy_benchmarks autonomy_benchmarks.launch.py \ - prediction:=/object_list/prediction \ - label:=/object_list/lidar_01 \ - visualize:=true \ - manual_playback:=true -``` - -```mermaid -flowchart LR - NODE("autonomy_benchmarks") - NODE o--o|~/request_samples| SC0:::hidden - classDef hidden display: none; -``` - -#### Service Clients - -| Service | Type | Description | -| --- | --- | --- | -| `~/request_samples` | `autonomy_datasets_msgs/srv/RequestSamples` | request samples from dataset | - -#### Parameters - -| Parameter | Type | Default | Description | -| --- | --- | --- | --- | -| `benchmark` | `string` | `nuscenes_lidar_object_detection` | benchmark name | -| `visualize` | `bool` | `false` | publish the per-sample true positives, false positives and false negatives for RViz | -| `manual_playback` | `bool` | `false` | leave requesting the samples of the dataset to the user | -| `samples_per_request` | `int` | `1` | number of samples to request from the dataset at a time; 0 requests all remaining samples at once, 1 evaluates every sample before the next one is published | -| `sample_ids` | `string` | - | comma-separated IDs of the dataset samples to evaluate (e.g. '0,10,20'); if empty, all samples of the dataset are evaluated | -| `evaluation_timeout` | `float` | `60.0` | seconds to wait for a published sample to be evaluated before continuing without it | -| `results_path` | `string` | - | path of the JSON file the benchmark results are written to; results are only logged if empty | - -## Launch Files - -### [`autonomy_benchmarks.launch.py`](launch/autonomy_benchmarks.launch.py) - -| Argument | Default | Description | -| --- | --- | --- | -| `request_samples` | `"~/request_samples"` | service of the dataset node used to request the samples to evaluate | -| `benchmark` | `"nuscenes_lidar_object_detection"` | benchmark to run | -| `name` | `"autonomy_benchmarks"` | node name | -| `namespace` | `""` | node namespace | -| `log_level` | `"info"` | ros logging level | -| `use_sim_time` | `"true"` | use sim time | -| `visualize` | `"false"` | publish the per-sample true positives, false positives and false negatives and open RViz | -| `manual_playback` | `"false"` | request the samples to evaluate via the playback panel in RViz instead of one after another, and open RViz on it | -| `samples_per_request` | `"1"` | number of samples to request from the dataset at a time (0 requests all remaining samples at once) | -| `sample_ids` | `""` | comma-separated IDs of the dataset samples to evaluate (all samples if empty) | -| `evaluation_timeout` | `"60.0"` | seconds to wait for a published sample to be evaluated before continuing without it | -| `results_path` | `""` | path of the JSON file the benchmark results are written to (results are only logged if empty) | diff --git a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py b/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py deleted file mode 100644 index bdb62a9..0000000 --- a/autonomy_benchmarks/autonomy_benchmarks/benchmarks/__init__.py +++ /dev/null @@ -1,8 +0,0 @@ -# Copyright Thinking Cars GmbH -# SPDX-License-Identifier: Apache-2.0 - -"""Benchmark implementations for AutonomyHub.""" - -from autonomy_benchmarks.benchmarks.AutonomyBenchmark import AutonomyBenchmark - -__all__ = ["AutonomyBenchmark"] diff --git a/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py b/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py deleted file mode 100644 index 9a567d6..0000000 --- a/autonomy_benchmarks/launch/autonomy_benchmarks.launch.py +++ /dev/null @@ -1,144 +0,0 @@ -#!/usr/bin/env python3 - -# Copyright Thinking Cars GmbH -# SPDX-License-Identifier: Apache-2.0 - -from launch import LaunchDescription -from launch.actions import DeclareLaunchArgument -from launch.conditions import IfCondition -from launch.substitutions import LaunchConfiguration, PathJoinSubstitution -from launch_ros.actions import Node, SetParameter -from launch_ros.parameter_descriptions import ParameterValue -from launch_ros.substitutions import FindPackageShare - -# Names of the benchmark's data inputs. Each name must match a key returned by -# the benchmark's ``required_inputs()`` (and thus a ``compute_sample_metrics`` -# parameter). This single list is the source of truth: it drives both the -# per-input launch arguments and the topic remappings below, so adding a new -# input (e.g. "map", "radar") is a one-line change here. -_INPUTS = ["prediction", "label"] - -# Inputs whose topic is derived from another input's topic instead of getting a -# launch argument of its own, as ``input name -> (source input, topic suffix)``. -# The dataset publishes the meta information of an object list next to it, on -# "/meta_info". -_DERIVED_INPUTS = {"label_meta_info": ("label", "/meta_info")} - - -def generate_launch_description(): - """Create and return the launch description for the autonomy_benchmarks node.""" - - # Service of the dataset node the benchmark requests the samples to evaluate from, remapped - # onto the topic its argument resolves to just like the benchmark's data inputs. - remappable_topics = [ - DeclareLaunchArgument( - "request_samples", - default_value="~/request_samples", - description="service of the dataset node used to request the samples to evaluate", - ), - ] - - args = [ - *remappable_topics, - DeclareLaunchArgument( - "benchmark", - default_value="nuscenes_lidar_object_detection", - description="benchmark name", - choices=["nuscenes_lidar_object_detection"], - ), - DeclareLaunchArgument("name", default_value="autonomy_benchmarks", description="node name"), - DeclareLaunchArgument("namespace", default_value="", description="node namespace"), - DeclareLaunchArgument( - "log_level", default_value="info", description="ROS logging level (debug, info, warn, error, fatal)" - ), - DeclareLaunchArgument("use_sim_time", default_value="true", description="use simulation clock"), - DeclareLaunchArgument( - "visualize", - default_value="false", - choices=["true", "false"], - description="publish the per-sample true positives, false positives and false negatives and open RViz on them", - ), - DeclareLaunchArgument( - "manual_playback", - default_value="false", - choices=["true", "false"], - description="request the samples to evaluate via the playback panel in RViz instead of one after another, " - "and open RViz on it", - ), - DeclareLaunchArgument( - "samples_per_request", - default_value="1", - description="number of samples to request from the dataset at a time (0 requests all remaining samples at once)", - ), - DeclareLaunchArgument( - "sample_ids", - default_value="", - description="comma-separated IDs of the dataset samples to evaluate (all samples if empty)", - ), - DeclareLaunchArgument( - "evaluation_timeout", - default_value="60.0", - description="seconds to wait for a published sample to be evaluated before continuing without it", - ), - DeclareLaunchArgument( - "results_path", - default_value="", - description="path of the JSON file the benchmark results are written to (results are only logged if empty)", - ), - # One argument per benchmark input; defaults to the node-relative name so - # an unset input is a no-op remap. Override with e.g. prediction:=/real/topic. - *[ - DeclareLaunchArgument( - name, - default_value=f"~/{name}", - description=f"real ROS topic feeding the benchmark's '{name}' input", - ) - for name in _INPUTS - ], - ] - - # Remap each node-relative input name onto the topic its argument resolves to, - # then onto the topic derived from another input for the derived inputs. - remappings = [(name, LaunchConfiguration(name)) for name in _INPUTS] - remappings += [(name, [LaunchConfiguration(source), suffix]) for name, (source, suffix) in _DERIVED_INPUTS.items()] - remappings += [(la.default_value[0].text, LaunchConfiguration(la.name)) for la in remappable_topics] - - node = Node( - package="autonomy_benchmarks", - executable="autonomy_benchmarks", - namespace=LaunchConfiguration("namespace"), - name=LaunchConfiguration("name"), - parameters=[ - {"benchmark": LaunchConfiguration("benchmark")}, - {"visualize": ParameterValue(LaunchConfiguration("visualize"), value_type=bool)}, - {"manual_playback": ParameterValue(LaunchConfiguration("manual_playback"), value_type=bool)}, - {"samples_per_request": ParameterValue(LaunchConfiguration("samples_per_request"), value_type=int)}, - {"sample_ids": ParameterValue(LaunchConfiguration("sample_ids"), value_type=str)}, - {"evaluation_timeout": ParameterValue(LaunchConfiguration("evaluation_timeout"), value_type=float)}, - {"results_path": ParameterValue(LaunchConfiguration("results_path"), value_type=str)}, - ], - arguments=["--ros-args", "--log-level", LaunchConfiguration("log_level")], - remappings=remappings, - output="screen", - emulate_tty=True, - ) - rviz = Node( - package="rviz2", - executable="rviz2", - name="rviz2", - arguments=[ - "-d", - PathJoinSubstitution([FindPackageShare("autonomy_benchmarks"), "config", "conf.rviz"]), - ], - condition=IfCondition(LaunchConfiguration("visualize")), - output="screen", - ) - - return LaunchDescription( - [ - *args, - SetParameter("use_sim_time", LaunchConfiguration("use_sim_time")), - node, - rviz, - ] - ) diff --git a/autonomy_benchmarks/setup.cfg b/autonomy_benchmarks/setup.cfg deleted file mode 100644 index 7e58c70..0000000 --- a/autonomy_benchmarks/setup.cfg +++ /dev/null @@ -1,4 +0,0 @@ -[develop] -script_dir=$base/lib/autonomy_benchmarks -[install] -install_scripts=$base/lib/autonomy_benchmarks diff --git a/autonomy_benchmarks/tests/benchmarks/__init__.py b/autonomy_benchmarks/tests/benchmarks/__init__.py deleted file mode 100644 index 5b0f7f7..0000000 --- a/autonomy_benchmarks/tests/benchmarks/__init__.py +++ /dev/null @@ -1,4 +0,0 @@ -# Copyright Thinking Cars GmbH -# SPDX-License-Identifier: Apache-2.0 - -"""Benchmark-oriented test modules for autonomy_benchmarks.""" diff --git a/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py b/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py deleted file mode 100644 index 4bc831e..0000000 --- a/autonomy_benchmarks/tests/benchmarks/test_AutonomyBenchmark.py +++ /dev/null @@ -1,194 +0,0 @@ -# Copyright Thinking Cars GmbH -# SPDX-License-Identifier: Apache-2.0 - -"""Tests for the result store of the AutonomyBenchmark base class. - -Metrics are reported on three levels: for every single sample, aggregated over the samples of each -scene of the dataset, and aggregated over all evaluated samples. A minimal benchmark whose metrics -are trivial to predict is used, so that the tests cover the grouping and not a metric definition. -""" - -from __future__ import annotations - -import json -from typing import Any, Dict, List - -from autonomy_benchmarks.benchmarks.AutonomyBenchmark import AutonomyBenchmark - - -class CountingBenchmark(AutonomyBenchmark): - """Minimal benchmark counting the objects of a sample and summing them up when aggregating.""" - - def __init__(self) -> None: - """Name the benchmark.""" - super().__init__(name="counting", description="counts objects") - - def required_inputs(self) -> Dict[str, Any]: - """Declare the inputs, which this benchmark does not read from ROS messages.""" - return {"prediction": object, "label": object} - - def compute_sample_metrics(self, prediction: Any, label: Any, sample_id: str = None) -> Dict[str, Any]: - """Report the given prediction and label counts of a single sample.""" - return {"num_predictions": prediction, "num_labels": label} - - def compute_aggregated_metrics(self, sample_results: List[Dict[str, Any]]) -> Dict[str, Any]: - """Sum the counts of the given samples.""" - return { - "num_predictions": sum(entry["metrics"]["num_predictions"] for entry in sample_results), - "num_labels": sum(entry["metrics"]["num_labels"] for entry in sample_results), - } - - -def _benchmark_of(samples) -> CountingBenchmark: - """Record ``(sample_id, scene_id, num_predictions, num_labels)`` samples in a benchmark.""" - benchmark = CountingBenchmark() - for sample_id, scene_id, num_predictions, num_labels in samples: - benchmark.record_sample(prediction=num_predictions, label=num_labels, sample_id=sample_id, scene_id=scene_id) - return benchmark - - -class TestSampleResults: - """Tests recording samples with the scene of the dataset they belong to.""" - - # The recorded samples themselves are no longer reported alongside the aggregated - # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize() - # def test_records_sample_with_its_scene(self): - # """A recorded sample keeps its ID, its scene and its metrics.""" - # benchmark = _benchmark_of([("0", "scene_a", 2, 3)]) - # - # assert benchmark.finalize()["sample_results"] == [ - # {"sample_id": "0", "scene_id": "scene_a", "metrics": {"num_predictions": 2, "num_labels": 3}} - # ] - - def test_scene_can_be_set_after_the_sample_was_recorded(self): - """An evaluation loop that learns the scene late sets it on the returned entry.""" - benchmark = CountingBenchmark() - - entry = benchmark.record_sample(prediction=1, label=1, sample_id="0") - assert entry["scene_id"] is None - entry["scene_id"] = "scene_a" - - assert benchmark.sample_results_by_scene() == {"scene_a": [entry]} - - -class TestFinalize: - """Tests aggregating the recorded samples per scene and over the whole benchmark.""" - - def test_aggregates_per_sample_scene_and_benchmark(self): - """Metrics are reported for every sample, every scene and all samples together.""" - results = _benchmark_of( - [ - ("0", "scene_a", 1, 1), - ("1", "scene_a", 2, 3), - ("2", "scene_b", 4, 5), - ] - ).finalize() - - assert results["num_samples"] == 3 - assert results["num_scenes"] == 2 - assert results["aggregated_metrics"] == {"num_predictions": 7, "num_labels": 9} - assert results["scene_results"]["scene_a"] == { - "num_samples": 2, - "sample_ids": ["0", "1"], - "aggregated_metrics": {"num_predictions": 3, "num_labels": 4}, - } - assert results["scene_results"]["scene_b"]["aggregated_metrics"] == {"num_predictions": 4, "num_labels": 5} - # The metrics of the single samples are no longer reported alongside the aggregated - # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize() - # assert [entry["metrics"] for entry in results["sample_results"]] == [ - # {"num_predictions": 1, "num_labels": 1}, - # {"num_predictions": 2, "num_labels": 3}, - # {"num_predictions": 4, "num_labels": 5}, - # ] - - def test_groups_samples_of_a_scene_that_are_not_recorded_consecutively(self): - """Samples are grouped by their scene, not by the order they were recorded in.""" - results = _benchmark_of( - [ - ("0", "scene_a", 1, 0), - ("1", "scene_b", 2, 0), - ("2", "scene_a", 4, 0), - ] - ).finalize() - - assert results["scene_results"]["scene_a"]["sample_ids"] == ["0", "2"] - assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 5 - assert results["scene_results"]["scene_b"]["sample_ids"] == ["1"] - - def test_samples_without_a_scene_are_only_aggregated_over_the_benchmark(self): - """A sample that cannot be attributed to a scene still counts for the whole benchmark.""" - results = _benchmark_of([("0", "scene_a", 1, 0), ("1", None, 2, 0)]).finalize() - - assert results["num_samples"] == 2 - assert results["num_scenes"] == 1 - assert results["aggregated_metrics"]["num_predictions"] == 3 - assert results["scene_results"]["scene_a"]["aggregated_metrics"]["num_predictions"] == 1 - - def test_reports_no_scene_without_recorded_scenes(self): - """Samples recorded without a scene aggregate to no scene results at all.""" - results = _benchmark_of([("0", None, 1, 0)]).finalize() - - assert results["num_scenes"] == 0 - assert results["scene_results"] == {} - - -class TestSaveResults: - """Tests writing the results of all three levels to a JSON file.""" - - def test_writes_sample_scene_and_benchmark_metrics(self, tmp_path): - """The stored results hold the metrics of every sample, every scene and the benchmark.""" - benchmark = _benchmark_of([("0", "scene_a", 1, 1), ("1", "scene_b", 2, 2)]) - - output_path = benchmark.save_results(str(tmp_path / "results" / "counting.json")) - - stored = json.loads(open(output_path).read()) - assert stored["aggregated_metrics"] == {"num_predictions": 3, "num_labels": 3} - assert sorted(stored["scene_results"]) == ["scene_a", "scene_b"] - # The single samples are no longer written alongside the aggregated - # results, while 'sample_results' is commented out in AutonomyBenchmark.finalize() - # assert [entry["sample_id"] for entry in stored["sample_results"]] == ["0", "1"] - - def test_writes_previously_computed_results(self, tmp_path): - """Results that have already been computed are written as they are.""" - benchmark = _benchmark_of([("0", "scene_a", 1, 1)]) - results = benchmark.finalize() - - output_path = benchmark.save_results(str(tmp_path / "counting.json"), results=results) - - assert json.loads(open(output_path).read())["aggregated_metrics"] == results["aggregated_metrics"] - - -class TestIncompleteResults: - """Tests marking the results of an evaluation that did not process all samples.""" - - def test_results_of_all_samples_are_complete(self): - """Results aggregated after the last sample of the benchmark are marked as complete.""" - assert _benchmark_of([("0", "scene_a", 1, 1)]).finalize()["complete"] is True - - def test_interrupted_results_are_marked_incomplete(self): - """Results aggregated before the last sample, e.g. after Ctrl-C, are marked as incomplete.""" - results = _benchmark_of([("0", "scene_a", 1, 1)]).finalize(complete=False) - - assert results["complete"] is False - # the samples that were evaluated are still reported - assert results["num_samples"] == 1 - assert results["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1} - - def test_writes_incomplete_results_to_the_results_file(self, tmp_path): - """The results file of an interrupted evaluation marks the results it holds as incomplete.""" - benchmark = _benchmark_of([("0", "scene_a", 1, 1)]) - - output_path = benchmark.save_results(str(tmp_path / "counting.json"), complete=False) - - stored = json.loads(open(output_path).read()) - assert stored["complete"] is False - assert stored["aggregated_metrics"] == {"num_predictions": 1, "num_labels": 1} - - def test_written_results_keep_the_flag_they_were_finalized_with(self, tmp_path): - """A given payload is written as it is, marked the way it was finalized.""" - benchmark = _benchmark_of([("0", "scene_a", 1, 1)]) - results = benchmark.finalize(complete=False) - - output_path = benchmark.save_results(str(tmp_path / "counting.json"), results=results) - - assert json.loads(open(output_path).read())["complete"] is False diff --git a/autonomy_benchmarks/tests/test_autonomy_benchmarks.py b/autonomy_benchmarks/tests/test_autonomy_benchmarks.py deleted file mode 100644 index 651eed7..0000000 --- a/autonomy_benchmarks/tests/test_autonomy_benchmarks.py +++ /dev/null @@ -1,307 +0,0 @@ -# Copyright Thinking Cars GmbH -# SPDX-License-Identifier: Apache-2.0 - -"""Tests for the helpers of the autonomy_benchmarks node. - -The node itself drives the evaluation via the ``request_samples`` service of the dataset and is -covered by running it against a dataset. Tested here are the parsing of the samples to evaluate, -as an unparsable value stops the node, the matching of the received input messages into the -samples to evaluate, which has to hold up when the dataset continues with a scene that was -recorded before the scene played before it, the evaluation of a sample, which requests the next -samples unless the user controls the playback manually, and the finalization of the results, -which reports the samples of an interrupted evaluation as incomplete. -""" - -from __future__ import annotations - -from collections import deque -from types import SimpleNamespace - -import pytest -from autonomy_benchmarks.autonomy_benchmarks import AutonomyBenchmarks, parse_sample_ids, SampleSynchronizer - -_TOPICS = ["prediction", "label", "label_meta_info"] - -# stamps of the last sample of a scene and of the first samples of the scene the dataset continues -# with, which nuScenes recorded years earlier -_PREVIOUS_SCENE = (1537853053, 397270000) -_NEXT_SCENE = [(1531885320, 49418000), (1531885320, 548742000), (1531885321, 48634000)] - - -def _message(stamp: tuple[int, int]) -> SimpleNamespace: - """Fake an input message stamped with the recording time of its sample.""" - return SimpleNamespace(header=SimpleNamespace(stamp=SimpleNamespace(sec=stamp[0], nanosec=stamp[1]))) - - -def _synchronizer(queue_size: int = 10) -> tuple[SampleSynchronizer, list]: - """Create a synchronizer of the benchmark inputs next to the list of the samples it matched.""" - matched_samples: list = [] - synchronizer = SampleSynchronizer(_TOPICS, callback=lambda *msgs: matched_samples.append(msgs), queue_size=queue_size) - return synchronizer, matched_samples - - -def _publish_sample(synchronizer: SampleSynchronizer, stamp: tuple[int, int], topics=None) -> dict: - """Add one message per given input (all of them by default), all stamped with the same time.""" - messages = {topic: _message(stamp) for topic in topics or _TOPICS} - for topic, message in messages.items(): - synchronizer.add(topic, message) - return messages - - -class TestParseSampleIds: - """Tests parsing the 'sample_ids' parameter into the sample IDs to request.""" - - def test_parses_comma_separated_ids(self): - """Sample IDs are parsed in the given order.""" - assert parse_sample_ids("0,10,20") == [0, 10, 20] - - def test_parses_single_id(self): - """A single ID is a valid request.""" - assert parse_sample_ids("7") == [7] - - def test_ignores_surrounding_whitespace(self): - """Sample IDs separated by ', ' are parsed like IDs separated by ','.""" - assert parse_sample_ids(" 1, 2 ,3 ") == [1, 2, 3] - - @pytest.mark.parametrize("sample_ids", ["", " ", ","]) - def test_no_ids_evaluate_the_whole_dataset(self, sample_ids): - """An empty value requests no specific samples, so the whole dataset is evaluated.""" - assert parse_sample_ids(sample_ids) == [] - - @pytest.mark.parametrize("sample_ids", ["1;2", "first", "1.5", "1-2"]) - def test_rejects_values_that_are_no_ids(self, sample_ids): - """A value that is no comma-separated list of IDs is rejected.""" - with pytest.raises(ValueError): - parse_sample_ids(sample_ids) - - -class TestSampleSynchronizer: - """Tests matching the messages of the benchmark inputs into the samples to evaluate.""" - - def test_matches_the_messages_of_a_sample_in_input_order(self): - """A sample is reported once every input has been received, in the order of the inputs.""" - synchronizer, matched_samples = _synchronizer() - - messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=list(reversed(_TOPICS))) - - assert matched_samples == [tuple(messages[topic] for topic in _TOPICS)] - assert not synchronizer.incomplete_samples - - def test_waits_for_the_missing_inputs_of_a_sample(self): - """A sample of which an input is missing is not reported yet.""" - synchronizer, matched_samples = _synchronizer() - - _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"]) - - assert matched_samples == [] - - def test_matches_samples_of_a_scene_recorded_before_the_previous_scene(self): - """The dataset continues with an older scene, whose samples are matched all the same.""" - synchronizer, matched_samples = _synchronizer() - - _publish_sample(synchronizer, _PREVIOUS_SCENE) - messages = _publish_sample(synchronizer, _NEXT_SCENE[0]) - - assert len(matched_samples) == 2 - assert matched_samples[-1] == tuple(messages[topic] for topic in _TOPICS) - - def test_keeps_a_waiting_sample_older_than_a_matched_one(self): - """A sample of a new, older scene is not dropped by a late sample of the previous scene.""" - synchronizer, matched_samples = _synchronizer() - # the last sample of the previous scene still waits for the system under test, while the - # first sample of the next scene, recorded years earlier, is published already - pending = _publish_sample(synchronizer, _PREVIOUS_SCENE, topics=["label", "label_meta_info"]) - messages = _publish_sample(synchronizer, _NEXT_SCENE[0], topics=["label", "label_meta_info"]) - - synchronizer.add("prediction", _message(_PREVIOUS_SCENE)) - synchronizer.add("prediction", _message(_NEXT_SCENE[0])) - - assert len(matched_samples) == 2 - assert matched_samples[0][1:] == (pending["label"], pending["label_meta_info"]) - assert matched_samples[1][1:] == (messages["label"], messages["label_meta_info"]) - - def test_gives_up_on_the_sample_waiting_the_longest(self): - """Samples are dropped in the order they arrived, not by their stamp.""" - synchronizer, matched_samples = _synchronizer(queue_size=2) - - # a sample of the previous scene waits first, followed by two samples of the older scene - for stamp in [_PREVIOUS_SCENE, *_NEXT_SCENE[:2]]: - _publish_sample(synchronizer, stamp, topics=["label"]) - - # the queue size is exceeded, so the sample that has been waiting the longest is given up - # on, even though the samples kept for evaluation were recorded years before it - assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[:2] - - _publish_sample(synchronizer, _NEXT_SCENE[2], topics=["label"]) - - assert list(synchronizer.incomplete_samples) == _NEXT_SCENE[1:] - assert matched_samples == [] - - -class _FakeLogger: - """Collect the messages the node logs instead of publishing them to ROS.""" - - def __init__(self): - """Start with an empty log.""" - self.messages: list = [] - - def info(self, message: str, **kwargs) -> None: - """Record a logged message, whatever its severity.""" - self.messages.append(message) - - debug = warn = error = info - - -class _FakeBenchmarkHandler: - """Stand in for the benchmark whose results the node aggregates and writes.""" - - def __init__(self): - """Start without finalized or written results.""" - self.finalized_complete = None - self.written_results = None - - def finalize(self, complete: bool = True) -> dict: - """Report results that are marked the way the node asked for.""" - self.finalized_complete = complete - return {"num_samples": 2, "num_scenes": 1, "complete": complete, "aggregated_metrics": {}} - - def save_results(self, output_path: str, results: dict = None) -> str: - """Keep the results instead of writing them to a file.""" - self.written_results = results - return output_path - - -class _RecordingBenchmarkHandler: - """Stand in for the benchmark that records the samples the node evaluates.""" - - def __init__(self): - """Start without recorded samples.""" - self.recorded_samples: list = [] - - def record_sample(self, sample_id: str, **messages) -> dict: - """Record a sample the way the benchmark stores it, still without a scene.""" - entry = {"sample_id": sample_id, "scene_id": None, "metrics": {}} - self.recorded_samples.append(entry) - return entry - - -def _evaluating_node(manual_playback: bool) -> SimpleNamespace: - """Stub the node state that evaluating a sample reads, counting the attempts to continue the benchmark.""" - node = SimpleNamespace( - manual_playback=manual_playback, - input_topics=_TOPICS, - benchmark_handler=_RecordingBenchmarkHandler(), - num_evaluated_samples=0, - scenes_awaiting_sample=deque(), - samples_awaiting_scene=deque(), - evaluation_timeout=60.0, - evaluation_deadline=None, - results_path="", - visualization_publishers={}, - num_advances=0, - get_logger=lambda logger=_FakeLogger(): logger, - ) - node.advance_benchmark = lambda: setattr(node, "num_advances", node.num_advances + 1) - return node - - -def _sample_messages(stamp: tuple[int, int]) -> tuple: - """Fake the synchronized input messages of one sample, in the order of the inputs.""" - return tuple(_message(stamp) for _ in _TOPICS) - - -class TestEvaluateSample: - """Tests evaluating a sample with the playback driven by the benchmark or by the user in RViz.""" - - def test_benchmark_requests_the_next_samples_after_evaluating_one(self): - """Without manual playback, the benchmark continues with the next samples on its own.""" - node = _evaluating_node(manual_playback=False) - - AutonomyBenchmarks.evaluate_sample(node, *_sample_messages(_NEXT_SCENE[0])) - - assert node.num_advances == 1 - # the dataset reports the scene of the sample with the response to the request - assert list(node.samples_awaiting_scene) == node.benchmark_handler.recorded_samples - - def test_manual_playback_leaves_requesting_samples_to_the_user(self): - """With manual playback, samples are evaluated as they arrive, without requesting further ones.""" - node = _evaluating_node(manual_playback=True) - - for stamp in _NEXT_SCENE: - AutonomyBenchmarks.evaluate_sample(node, *_sample_messages(stamp)) - - assert node.num_evaluated_samples == len(_NEXT_SCENE) - assert node.num_advances == 0 - # the scenes are only reported to the playback panel, so no sample waits for one - assert not node.samples_awaiting_scene - - -def _node(num_evaluated_samples: int = 2, results_path: str = "/results/benchmark.json") -> SimpleNamespace: - """Stub the node state that finalizing a benchmark reads, without initializing ROS.""" - return SimpleNamespace( - benchmark="counting", - benchmark_finished=False, - benchmark_handler=_FakeBenchmarkHandler(), - num_evaluated_samples=num_evaluated_samples, - results_path=results_path, - request_timer=SimpleNamespace(cancel=lambda: None), - get_logger=lambda logger=_FakeLogger(): logger, - ) - - -class TestFinalizeBenchmark: - """Tests reporting the results of a benchmark that ran to its end or was interrupted.""" - - def test_finished_benchmark_writes_complete_results(self): - """A benchmark that evaluated all its samples reports complete results.""" - node = _node() - - AutonomyBenchmarks.finalize_benchmark(node) - - assert node.benchmark_handler.finalized_complete is True - assert node.benchmark_handler.written_results["complete"] is True - - def test_interrupted_benchmark_writes_incomplete_results(self): - """An evaluation stopped before its last sample, e.g. with Ctrl-C, still writes its results.""" - node = _node() - - AutonomyBenchmarks.finalize_benchmark(node, complete=False) - - assert node.benchmark_handler.finalized_complete is False - assert node.benchmark_handler.written_results["complete"] is False - - def test_finished_benchmark_is_not_finalized_again_on_shutdown(self): - """Shutting down after the last sample must not overwrite the results with incomplete ones.""" - node = _node() - AutonomyBenchmarks.finalize_benchmark(node) - - AutonomyBenchmarks.finalize_benchmark(node, complete=False) - - assert node.benchmark_handler.finalized_complete is True - assert node.benchmark_handler.written_results["complete"] is True - - def test_manual_playback_writes_its_results_when_stopped(self): - """With manual playback, which runs no request timer, the results are written once the node is stopped.""" - node = _node() - node.request_timer = None - - AutonomyBenchmarks.finalize_benchmark(node, complete=False) - - assert node.benchmark_handler.written_results["complete"] is False - - def test_interrupted_benchmark_without_samples_writes_nothing(self): - """An evaluation interrupted before its first sample has no results to write.""" - node = _node(num_evaluated_samples=0) - - AutonomyBenchmarks.finalize_benchmark(node, complete=False) - - assert node.benchmark_handler.written_results is None - - def test_results_are_only_logged_without_a_results_path(self): - """Without 'results_path' the interrupted results are logged instead of written.""" - node = _node(results_path="") - - AutonomyBenchmarks.finalize_benchmark(node, complete=False) - - assert node.benchmark_handler.finalized_complete is False - assert node.benchmark_handler.written_results is None diff --git a/autonomy_evaluation/README.md b/autonomy_evaluation/README.md new file mode 100644 index 0000000..5ea37a4 --- /dev/null +++ b/autonomy_evaluation/README.md @@ -0,0 +1,98 @@ +# `autonomy_evaluation` + +Metrics-based evaluation of automated driving modules, generating the evidence for benchmarking automated driving deployments + +`autonomy_evaluation` is the part of the **Autonomy.Benchmarks** suite that turns the output of a system under test into +metrics. It evaluates the samples that [autonomy_datasets](https://github.com/thinking-cars/autonomy_datasets) replays +against the labels of the dataset, and reports the metrics per scene and over all evaluated samples, as one part of the evidence +an automated driving deployment is benchmarked on. + +## Nodes + +### `autonomy_evaluation` + +The node runs the evaluation selected by `evaluation`, either one of this package by its name or one implemented in another package as `:` (see [Adding more Evaluations](../docs/IMPLEMENTATION.md#adding-more-evaluations)). Each evaluation declares the topics it reads: its _inputs_ from the system under test, and, if it compares them with a reference, its _ground truth_. The node subscribes to all of them, each on its node-relative name, which the launch file remaps onto the topic given by the launch argument of the same name, e.g. `prediction:=/object_list/prediction`. A topic without such an argument is subscribed in the private namespace of the node. The node logs which topic it evaluates as input and which as ground truth. + +| Evaluation | Inputs | Ground truth | +| --- | --- | --- | +| `nuscenes_lidar_object_detection` | `prediction` (`perception_msgs/ObjectList`) | `label` (`perception_msgs/ObjectList`), `label_meta_info` (`autonomy_datasets_msgs/ObjectListMetaInfo`, subscribed next to `label` on `