The accelerator you own is not always the one that helps
A machine doing continuous video analysis was running hot and pinning its processor. The obvious answer was sitting right there: the chip has an integrated graphics unit, doing nothing. Move the hard work onto it. That is what accelerators are for.
It did not work, and the reason it did not work was not a configuration problem I could grind my way through. The detection model I was running had no path onto that particular graphics unit. The vendor's own toolkit was the supported route, and that toolkit expects a different family of hardware. The one generic path that did exist, I measured, and it was slower than the processor I was trying to relieve.
So I had spent an evening on a dead end, and the machine was still hot. Then I did the thing I should have done first: I measured which stage was actually expensive.
A kitchen running late is not always slow at cooking. Sometimes everything is waiting on one person washing up. Buying a better oven does not clear the sink.
Analysing a video stream is at least two jobs wearing one name. First the compressed stream has to be turned back into pictures. Then something has to look at the pictures and decide what is in them. I had assumed the looking was the expensive part, because the looking is the clever part, and clever feels costly.
It was the other one. Turning compressed video back into frames, continuously, for several streams, was most of the heat. And that job did have a hardware path on the graphics unit I already owned, because video decode is a fixed function block that exists specifically to do this. Enabling it moved the load off the processor and the temperature came down.
The detection stayed on the processor. It is still there. It was never the problem, so making it faster would have bought me nothing, and I would have been able to write a confident log about the accelerator that fixed everything while being wrong about why.
The lesson is not "use hardware decode". The lesson is that the obvious knob is often the wrong knob, and the only way to know is to measure which stage is hot before you spend money or an evening.
How to tell decode cost from inference cost, and how to test an accelerator path honestly before committing to it. The measurement approach generalises to any pipeline where you are about to optimise the stage that feels expensive.
1. Separate the stages before you measure anything
Write the pipeline out as distinct jobs. You cannot attribute cost to a stage you have not named.
# pipeline.md
# A. receive the stream (network, cheap)
# B. decode to frames (CPU or fixed-function hardware)
# C. scale / convert format (often forgotten, often expensive)
# D. run the model on a frame (inference)
# E. act on the result (cheap)
Stage C is the one that catches people. A conversion sitting between decode and model can cost more than either, and it never appears in the diagram you drew in your head.
2. Measure the baseline properly, not with a glance
Instantaneous load, over a window, while the real workload runs. A single reading taken when you happened to look is not a measurement.
# per-process CPU, sampled, NOT a lifetime average
top -b -n 5 -d 2 | grep <your-process>
# temperature over the same window
sensors # Linux, lm-sensors
# and record it, because you will need to compare later
Worth knowing: the CPU percentage in a simple process listing is an average over the process's whole life, so a job that was busy an hour ago and idle now still reads busy. Sample repeatedly and use the later samples.
3. Isolate decode by running it with no model at all
This is the test that answered my question, and it takes two minutes. Decode the same streams, throw the frames away, and look at the load.
# decode only, discard output, software path
ffmpeg -i <stream> -f null -
# watch CPU during that. If this alone is most of your baseline,
# your cost is decode and the model is not your problem.
4. Isolate the model by running it on frames you already have
The mirror image. Feed the model still images from disk so no decoding happens, and time it.
# time N inferences on pre-extracted frames
# if this is fast and decode was slow, stop optimising the model
# if this is slow, now you have a real reason to look at accelerators
Now you have two numbers that add up to roughly your baseline, and the larger one tells you where to spend your effort. Everything before this step was guessing.
5. Check whether a hardware path exists for your hardware, before trying to use it
Ask the system what it supports rather than assuming a graphics unit implies a path for your workload. These are different questions and they often have different answers.
# Linux: is there a video acceleration device, and what can it do?
ls /dev/dri/
vainfo # lists supported decode/encode profiles
# does your tool have the accelerated path compiled in?
ffmpeg -hwaccels
ffmpeg -decoders | grep -i <your-codec>
Read the profile list against the codec you actually receive. Support for one codec says nothing about another, and this is where an afternoon disappears.
6. Test the accelerated path and compare the same numbers
Same streams, same window, same measurement. A change you cannot compare is a change you are taking on faith.
# decode only, hardware path
ffmpeg -hwaccel <method> -i <stream> -f null -
# compare against the software number from step three:
# CPU during the run
# temperature at the end of the window
# whether frames were actually dropped
That last one matters. A hardware path that lowers CPU by silently dropping frames has not helped you, it has changed what you are measuring. Check the output count, not just the load.
7. Accept the result even when it is boring
Mine was: leave the model where it is, move the decode, done. No new hardware, no exotic toolkit, one flag. The satisfying answer and the correct answer are frequently different, and preferring the satisfying one is how people end up with accelerators that accelerate nothing.
# results.md, written even when the answer is dull
# decode, software : ............ CPU, ............ temp
# decode, hardware : ............ CPU, ............ temp
# model alone : ............ per frame
# decision : moved decode only. model left on CPU deliberately.
The honest catch: hardware decode is not free. It shifts load onto a block with its own limits, and stacking enough streams will saturate it too. You have moved the ceiling, not removed it, so write down where the new one is.
Related: measure before you delete is the same instinct applied to disk space instead of heat.