Insights
VSAT × AI

Edge Video Analytics Over Satellite: Sizing the Return Link

Eight 1080p cameras on a remote pumping station will typically want around 20 Mbit/s of upstream video. The satellite service feeding that station gives you far less than that going up — so the design question is what you let cross the link, not whether to put a model on the site.

28 August 2026 9 min read Vizocom Editorial
Eight cameras feed an on-site inference node. Continuous video never crosses the satellite return link; event clips and event records do, reaching the satellite and the network operations centre 8 cameras Edge inference 7–25 W on site satellite return link GEO round-trip floor ≈ 478 ms 24 Mbit/s frames never crosses 167 kbit/s clips 1.4 kbit/s events NOC

What you'll take away

  • Calculate your site’s upstream video budget before a vendor quotes you one.
  • Choose between continuous streaming, event-clip upload and metadata-only on numbers rather than instinct.
  • Size the burst — the moment an incident generates a dozen clips at once — not only the average.
  • Ask a video-analytics vendor four questions that separate a product from a demonstration.

Why the return link is the constraint, not the camera count

24 Mbit/s
Sustained upstream for eight 1080p cameras, streamed continuously
478 ms
GEO round-trip floor at zenith, propagation only
144×
Less upstream when the site sends event clips instead of frames
0.33
Median analyst precision at an 86% false-alarm rate, against 0.80 at 50%

A satellite network is asymmetric at the architectural level. In the DVB-RCS2 system specification the forward link is a continuous DVB-S2 carrier from a teleport, while the return link is MF-TDMA: each terminal transmits in assigned time slots drawn from a shared pool under demand assignment, with EIRP power control specified for the terminal (ETSI TS 101 545-1 V1.2.1, 2014).

Downstream capacity comes from one big, well-powered transmitter; upstream capacity is rationed among every terminal in the beam, each a small dish working against its own link budget. LEO and multi-orbit services shift the ratio and cut the latency, but do not reverse the direction.

Latency compounds it. For a geostationary satellite at 35,786 km, propagation alone gives a round-trip time — up, down, and back again — of about 478 ms at zenith, rising to roughly 542 ms at a 10-degree elevation angle as the slant range stretches to about 40,600 km. A single ground-to-ground traverse is half that. These are floors, before any modulation, scheduling or hub processing, and a protocol that acknowledges every chunk of an upload feels all of them.

The arithmetic you can do before you buy anything

Start with an honest per-camera figure, because there is no universal correct one. Axis's paper on bitrate control shows why: under average-bitrate control, a stream working to a 500 kbit/s budget was allowed to rise to roughly 4,000 kbit/s during a short period of movement in an otherwise quiet scene (Axis, Bitrate control for IP video). Motion and scene complexity, not resolution alone, set the number.

For a worked example, take 2.5 Mbit/s average per camera — defensible for 1080p30 H.265 on a moderately active outdoor scene, and a figure to replace with your own. On the codec: an NTIA/ITS subjective study using 20 scenes, 25 viewers and ITU-T P.913 absolute category rating found H.265 statistically equivalent to H.264 at half the bitrate, but also found mean opinion scores for H.265 at 4 Mbit/s spanning 38% of the subjective scale across scenes (Catellier and Pinson, IEEE ISM, 2015). The 50% saving is an average across content, not a guarantee for your car park at dusk.

Now compare three architectures for the same eight cameras. Assume 40 detection events per camera per day, a 15-second clip per event, a 600-byte structured event message and a 40 KB thumbnail. All three rows below carry the same 20% transport overhead, so they compare like with like.

Worked example. Bitrate assumption 2.5 Mbit/s per camera; replace it with a measurement from your own site.
ArchitectureWhat crosses the linkSustained upstreamPer dayVersus streaming
Continuous streamingEvery frame, always24 Mbit/s259 GBbaseline
Event clips15 s clip per detection~167 kbit/s average1.8 GB~144× less
Metadata + thumbnailEvent record plus one still~1.4 kbit/s average15.6 MB~16,600× less
Sustained upstream for the same eight cameras on a logarithmic scale: continuous streaming 24 Mbit per second, event clips 167 kilobit per second, metadata 1.4 kilobit per second Sustained upstream, same eight cameras logarithmic scale · 20% transport overhead applied to all three Continuous 24 Mbit/s Event clips 167 kbit/s · 144× less Metadata 1.4 kbit/s · 16,600× less assumes 2.5 Mbit/s per camera — measure your own
The same eight cameras, three architectures, on a logarithmic scale — a linear axis cannot hold a 16,600× range. Source: computed from the assumptions in the table above.

The gap between rows two and three is the decision most sites get wrong. Streaming to event clips buys two orders of magnitude; clips to metadata buys another two, and costs the ability to see what happened without asking for it.

The burst is what breaks the design

Average rate is not the number that hurts you. One incident produces a cluster of events across several cameras as a vehicle or person crosses overlapping fields of view. Twelve 15-second clips is about 56 MB of video, or roughly 68 MB on the wire with the same 20% overhead, and that queue drains at whatever the return link actually delivers.

Time to upload twelve event clips, about 68 megabytes on the wire: 17.6 minutes at 512 kilobit per second, 9 minutes at 1 megabit, 4.5 minutes at 2 megabit, 1.8 minutes at 5 megabit — every case exceeds a ninety-second response window Clearing one incident: 12 clips, ≈68 MB on the wire 90-second response window 512 kbit/s 17.6 min 1 Mbit/s 9.0 min 2 Mbit/s 4.5 min 5 Mbit/s 1.8 min not one row delivers the evidence inside the window
Drain time for one incident's clip queue against four return-link rates, with a 90-second response window marked. Source: computed; 68 MB on the wire at the stated link rate.

An intrusion resolves in about ninety seconds. Not one of those rows delivers the evidence inside that window — even a 5 Mbit/s return link is still uploading when the event is over. The fix is priority, not bandwidth: structured event first, thumbnail second, clip third, full-resolution footage only when a human asks.

That layering is standardised, which matters for procurement. ONVIF Profile M (v1.0, June 2021) mandates analytics metadata streaming over RTSP and defines the metadata types a device may carry — object class, geolocation, human body and face, vehicle and licence plate (ONVIF Profile M Specification). MQTT and MQTTS event publishing sits in the specification as a conditional feature, which is the detail that catches buyers out: two Profile M cameras need not share an MQTT event path. Specify Profile M and MQTT event handling explicitly, and you can change camera vendor without rewriting the event pipeline.

What actually runs at the site

Inference hardware is no longer the limiting factor. NVIDIA rates the Jetson Orin Nano at up to 40 TOPS within a 7–15 W envelope and the Orin NX at up to 100 TOPS within 10–25 W (NVIDIA Jetson Orin) — budgets a solar-and-battery cabinet can carry. Treat TOPS as a headline, not a throughput commitment: what determines your design is frames per second per stream, measured on your board inside your enclosure. Both figures assume power envelopes a sealed outdoor box in a Gulf summer will struggle to dissipate, so assume thermal throttling until you have measured otherwise.

A large parabolic ground-station antenna at the Alaska Satellite Facility, Fairbanks
Near Space Network antenna, Alaska Satellite Facility NASA image · public domain · self-host before publishing
The ground segment is where the return link is rationed. Image: NASA, “Alaska Satellite Facility” — NASA imagery is generally not subject to copyright in the United States; NASA credited as source.

False positives are a bandwidth problem and a staffing problem

Every false detection is also an upload, and then a human interruption. A controlled experiment by Layman and Roden put 51 participants in front of intrusion-alert queues at two false-alarm rates. The group facing an 86% rate reached a median precision of 0.33 on their escalations against 0.80 for the 50% group, and took roughly 40% longer on the task — a median difference of 5.3 minutes (Layman and Roden, arXiv:2307.07023, July 2023). Detection sensitivity did not differ significantly between the groups.

Carry that carefully. The participants were students recruited at UNC Wilmington, not practising analysts, and the alerts were network intrusion alerts rather than video detections on an industrial site. Treat the direction as instructive and the magnitude as indicative only. The mechanism is general, though: precision degrades the analyst, not only the link. At 10% precision your return link carries nine wasted clips for every real one, and the person watching stops believing the tenth.

What this looks like on a real site

For an oil and gas facility in the Rub' al Khali, a mining camp in the Sahel or a humanitarian compound in the Horn of Africa, three things follow.

Retain locally, upload selectively. Delete-at-edge forecloses options you may be contractually required to keep open, so establish the retention period your insurer and jurisdiction require before sizing storage, and treat the satellite link as a request channel rather than a replication channel.

Plan for models you cannot reach. Sites change — new vehicles, fencing and lighting — and if retraining means shipping a model over a metered link, budget the megabytes and schedule the window.

Assume the classifier carries assumptions about the data it was trained on. Regional plate formats, Arabic signage, regional dress and livestock crossing a perimeter at dusk are all plausible sources of degraded accuracy here, and no published benchmark will settle any of them for you. Ask the vendor for regional validation data, and accept “we don't have it” as an honest answer worth pricing in.

Practical takeaway: the return-link worksheet

Sustained upstream: R = N × B × (1 + h)N cameras, B measured average Mbit/s per camera, h transport overhead (use 0.2 for RTP/UDP over an encapsulated satellite link unless you have measured otherwise).

Burst drain time: T = (E × C × 8) / R_up seconds — E events in the burst, C clip size in MB, R_up return rate in Mbit/s. If T exceeds your required response time, you need priority tiers, not a bigger plan.

The site checklist

  • Measure per-camera bitrate on site over 24 hours, not from a datasheet.
  • Use the committed return rate in your contract, not the burst or “up to” figure.
  • Baseline events per camera per day for two weeks before sizing clip traffic.
  • Define four priority tiers: event record, thumbnail, clip, footage on demand.
  • Set local retention to your evidentiary requirement; verify it survives a power cycle.
  • Test the burst deliberately — walk a perimeter and time the queue.
  • Measure inference frames per second in the enclosure at peak ambient temperature.
  • Instrument false-positive rate per camera per week, and review it monthly.

Four questions for the vendor

  1. What is the measured precision and recall for my object classes, in a climate and scene like mine, and who measured it?
  2. Does the device implement ONVIF Profile M event publishing over MQTT, and can I see the event schema?
  3. What is sustained frames-per-second per stream at my resolution, at 55 °C, enclosure closed?
  4. How is a model updated over a 1 Mbit/s metered link, how large is it, and what happens if it is interrupted?

What we don't know yet

The evidence base thins quickly beyond the physics and the standards

  • No independent, published benchmark of video-analytics precision and recall under desert industrial conditions could be verified for this article. Vendor accuracy figures are self-reported and rarely state test conditions.
  • The false-alarm study above used students classifying network intrusion alerts. No equivalent controlled experiment for video analytics appears in the public literature.
  • Per-camera bitrate has no authoritative universal value; it depends on scene, motion and encoder settings. Every bitrate figure in the worked example is an assumption you should replace with a measurement.
  • Model drift rates on unattended remote sites over 6 to 12 months are not publicly measured in any source that could be confirmed.
  • 3GPP Release 19, which reached fully implementable specifications at the end of December 2025, adds store-and-forward and regenerative payload work to non-terrestrial networks. Its effect on metadata-only backhaul is a roadmap question, not a deployed capability.

Conclusion

The physics of a satellite return link has not changed and will not. What changed is that a 15-watt board at the site can now decide what deserves to cross it, which turns a bandwidth problem into an information-design problem: choose the tiers, measure the real bitrates, test the burst rather than the average. A site that sends events instead of footage sends the part someone will act on.

Frequently asked

How much upstream bandwidth do eight 1080p cameras need over satellite?

Streamed continuously at an assumed 2.5 Mbit/s each, about 24 Mbit/s sustained once 20% transport overhead is included — roughly 259 GB a day. Uploading a 15-second clip per detection instead drops that to around 167 kbit/s average, and sending only structured event records with a thumbnail drops it to about 1.4 kbit/s. Measure your own cameras before trusting any of those figures.

Why is a satellite return link so much narrower than the download?

It is architectural, not incidental. In DVB-RCS2 the forward link is a continuous carrier from a teleport with a large antenna and a high-power amplifier, while the return link is MF-TDMA: every terminal transmits in assigned time slots from a shared pool, and each terminal is a small dish working against its own link budget. LEO and multi-orbit services shift the ratio but not the direction.

What latency should I plan for on a GEO link?

Propagation alone gives a round-trip time of about 478 ms with the satellite at zenith, rising to roughly 542 ms at a 10-degree elevation angle. Those are floors, before modulation, scheduling or hub processing, so any protocol that acknowledges each chunk of an upload will feel all of them.

Does ONVIF Profile M guarantee I get MQTT events?

No. Profile M mandates analytics metadata streaming over RTSP, but MQTT and MQTTS event publishing sit in the specification as conditional features. Two Profile M cameras need not share an MQTT event path, so specify Profile M and MQTT event handling explicitly in the tender.

When is edge video analytics the wrong answer?

When precision is too low to trust. Every false detection is an upload and a human interruption, and a queue of wasted clips teaches the operator to ignore the real one. If a vendor cannot show measured precision and recall for your object classes in conditions like yours, treat the deployment as a trial with a measured false-positive rate, not a finished system.

Vizocom sizes VSAT links and edge AI deployments in the same conversation, because in these environments the bandwidth budget and the inference budget are one problem.

Sources

  1. ETSI, TS 101 545-1 V1.2.1, Second Generation DVB Interactive Satellite System (DVB-RCS2), Part 1: Overview and System Level Specification, April 2014. https://www.etsi.org/deliver/etsi_ts/101500_101599/10154501/01.02.01_60/ts_10154501v010201p.pdf
  2. Axis Communications, Bitrate control for IP video: average bitrate (ABR), variable bitrate (VBR), and maximum bitrate (MBR), white paper. https://whitepapers.axis.com/en-us/bitrate-control-for-ip-video
  3. A. Catellier and M. Pinson, NTIA/ITS, Characterization of the HEVC Coding Efficiency Advance Using 20 Scenes, IEEE International Symposium on Multimedia, December 2015. https://its.ntia.gov/publications/download/2015_IEEE_ISM_catellier.pdf
  4. ONVIF, Profile M Specification v1.0, June 2021. https://www.onvif.org/wp-content/uploads/2021/06/onvif-profile-m-specification-v1-0.pdf
  5. NVIDIA, Jetson Orin module specifications, accessed 28 August 2026. https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/
  6. L. Layman and W. Roden, A Controlled Experiment on the Impact of Intrusion Detection False Alarm Rate on Analyst Performance, arXiv:2307.07023, submitted 13 July 2023. https://arxiv.org/abs/2307.07023
  7. 3GPP, Non-Terrestrial Networks (NTN) overview, accessed 28 August 2026. https://www.3gpp.org/technologies/ntn-overview
  8. 3GPP, Release 19, accessed 28 August 2026. https://www.3gpp.org/specifications-technologies/releases/release-19
  9. NASA, Alaska Satellite Facility — Near Space Network antenna (image), accessed 28 August 2026. https://www.nasa.gov/image-detail/fairbanks-27/