Running the sim backend for a five-drone show, the simulation advances about
seven times slower than wall clock. I profiled it and the cause is not the
physics — it is per-call dispatch overhead in rowan and numpy, invoked on
3- and 4-element vectors 2000 times a second per drone.
Measurements
Core loop only (no ROS): mellinger controller, np backend, 6 s of flight
(takeoff then go_to), on f1e0995:
| drones |
real-time factor |
| 1 |
0.70x |
| 5 |
0.15x |
| 16 |
0.05x |
cProfile over 1000 steps with 5 drones (7.36 s total):
backend/np.py: Quadrotor.step — 4.83 s cumulative, 66%. Spread across
rowan.rotate, rowan.normalize, rowan.calculus.integrate, np.cross,
np.dot, np.polyval.
crazyflie_sil.py: _fwsetpoint_to_sim_data_types_state — 1.51 s, 20%.
Called via getSetpoint() on every step, but its return value is only read
by the none backend and the pdf/record_states/blender visualizations.
setState — 0.30 s, 4%. (Already much cheaper since the
rowan -> transforms3d change.)
Through ROS the full launch was 0.07x, with /clock and /tf both published
on every 0.5 ms step adding to the load.
Note on dt
The obvious fix — raising dt from 0.0005 — does not work. At dt = 0.001 the
drone falls during go_to; at 0.002 it stops tracking. executeController is
called once per physics step and the firmware controller assumes a fixed rate,
so dt is load-bearing. The steps have to get cheaper, not fewer.
Proposed fix
Two PRs, linked below. The first is a pure optimization with no behaviour
change (~12x, verified against the previous implementation to 1e-12); the
second makes the /tf, /clock and executor rates configurable. Together the
full launch with 5 drones and a gui goes from 0.07x to 1.0x.
Happy to adjust the scope — in particular, vectorizing the physics across all
N drones instead of optimizing the per-drone path may be the better long-term
direction for large swarms, and I would rather follow your preference there.
Running the sim backend for a five-drone show, the simulation advances about
seven times slower than wall clock. I profiled it and the cause is not the
physics — it is per-call dispatch overhead in
rowanandnumpy, invoked on3- and 4-element vectors 2000 times a second per drone.
Measurements
Core loop only (no ROS):
mellingercontroller,npbackend, 6 s of flight(takeoff then go_to), on
f1e0995:cProfileover 1000 steps with 5 drones (7.36 s total):backend/np.py: Quadrotor.step— 4.83 s cumulative, 66%. Spread acrossrowan.rotate,rowan.normalize,rowan.calculus.integrate,np.cross,np.dot,np.polyval.crazyflie_sil.py: _fwsetpoint_to_sim_data_types_state— 1.51 s, 20%.Called via
getSetpoint()on every step, but its return value is only readby the
nonebackend and the pdf/record_states/blender visualizations.setState— 0.30 s, 4%. (Already much cheaper since therowan -> transforms3d change.)
Through ROS the full launch was 0.07x, with
/clockand/tfboth publishedon every 0.5 ms step adding to the load.
Note on
dtThe obvious fix — raising
dtfrom 0.0005 — does not work. Atdt = 0.001thedrone falls during
go_to; at 0.002 it stops tracking.executeControlleriscalled once per physics step and the firmware controller assumes a fixed rate,
so
dtis load-bearing. The steps have to get cheaper, not fewer.Proposed fix
Two PRs, linked below. The first is a pure optimization with no behaviour
change (~12x, verified against the previous implementation to 1e-12); the
second makes the
/tf,/clockand executor rates configurable. Together thefull launch with 5 drones and a gui goes from 0.07x to 1.0x.
Happy to adjust the scope — in particular, vectorizing the physics across all
N drones instead of optimizing the per-drone path may be the better long-term
direction for large swarms, and I would rather follow your preference there.