Skip to main content
Captia Technology

Protocol supported by Captia Connect

MQTT in Captia Connect: how data published to a broker is ingested

Lightweight pub/sub messaging for telemetry. Connect subscribes to the relevant topics and normalises the payloads. Connect acts as a client subscribed to the broker topics, normalises the payloads into time series and applies local buffering if the broker or the network uplink drops.

What it is and the role it plays

What is MQTT and how does Captia Connect use it?

MQTT is a lightweight pub/sub messaging protocol over TCP designed for telemetry, with a central broker, hierarchical topics and QoS levels from 0 to 2. In Captia Connect it is an acquisition route: Connect acts as a client subscribed to the relevant topics, normalises the payloads into time series and applies local buffering if the broker or the network uplink drops.

How Captia Connect speaks MQTT

In MQTT nobody talks to anybody: everybody talks to the broker. So the first question in an integration is not which equipment publishes, but where the broker is and who owns it. Connect takes one of two positions depending on the answer, and they are worth keeping apart because the work, the risks and the conversation with the client change entirely.

Connect as a subscriber to a broker that already exists. This is the common retrofit case. The plant already has a broker set up by the sensor integrator, by the machine builder or by the IT team, and Connect joins as one more client: it opens the session, subscribes to the topics that concern it and publishes nothing. Here the work is agreement rather than deployment: which credentials, which subscription filters, and what happens when somebody changes the hierarchy without warning.

Connect alongside a broker at the edge. When there is no broker, or when the existing one sits in a vendor cloud and the plant does not want to depend on that uplink, the broker is brought up on the edge node itself, in a Docker container like the rest of the acquisition layer. Equipment publishes to a local network address, telemetry traffic never leaves the plant, and what crosses towards the platform is the already normalised series. What to demand of that broker, how to size it and which high availability model makes sense in industry is the subject of the industrial MQTT broker guide, which is where it is developed.

Session, port and transport

Connect opens an outbound TCP connection to the broker and negotiates the MQTT session over it. The port is 1883 when traffic runs in the clear and 8883 when it runs over TLS, which is the registered port for secure MQTT and the only one to consider on a new installation. Where the broker is only reachable over WebSocket, which happens with vendor platforms behind a corporate proxy, the session travels encapsulated over HTTPS on 443. The connection is always outbound from the node: MQTT requires no inbound port on the edge or on the publishing equipment.

Four things are decided in the CONNECT packet, and they later explain half the incidents. The client identifier must be stable and unique: if two clients present the same identifier the broker evicts the previous one, and the classic symptom is a pair of clients reconnecting in a loop every few seconds with data arriving in fragments. The keep alive sets how often a heartbeat is sent and therefore how quickly the broker notices the client has vanished: the timeout is one and a half times that interval, so a keep alive of 60 seconds means an unclean disconnection takes up to 90 seconds to show. The persistent session determines whether the broker keeps subscriptions and pending QoS 1 and 2 messages while the client is away instead of starting from scratch on every reconnection. And the credentials identify the client to the broker authentication service.

Connect uses a persistent session on the subscriptions carrying QoS 1, because that is what keeps a thirty second reconnection from leaving a thirty second hole in the history. In exchange, that decision forces the previous one: a stable client identifier, because the session is recovered by identifier. If the identifier is generated at random on every start, the persistent session is worth nothing.

Subscription: filters, wildcards and volume

Connect does not subscribe signal by signal. It subscribes with topic filters, which accept a single level wildcard, +, and a multi level wildcard, #, which may only appear at the end. A filter such as plant1/+/energy/# picks up energy from every line without enumerating them, and that is precisely the property that stops a new piece of equipment from requiring a change to the acquisition: if it publishes in the right place in the hierarchy, it simply appears.

The same property is a hazard when misused. Subscribing to # at the root picks up everything crossing the broker, including control topics, third party applications and broker diagnostics, and turns ingestion into a landfill somebody will have to clean downstream. The starting criterion is the opposite: filters as specific as the hierarchy allows, with the wildcard only at the level where the fleet is genuinely expected to grow.

When the volume on a busy topic exceeds what one client should process, a modern broker offers shared subscriptions (filters prefixed with $share/), which distribute messages across several subscribers in the same group rather than delivering a copy to each. That is the correct way to scale ingestion without duplicating data, and it depends on the broker, not on the client.

QoS: what each level guarantees and what it costs

The quality of service level is negotiated twice, and this is what causes most confusion on site: the publisher chooses the QoS with which it hands the message to the broker, and the subscriber chooses the maximum QoS at which it wants to receive it. The effective level of a delivery is the lower of the two. A sensor publishing at QoS 1 into a client subscribed at QoS 0 produces QoS 0 delivery, and the guarantee is lost on the leg that matters without anything looking misconfigured.

Connect does not apply one level across the whole installation. It decides by the nature of the data, on a short rule: if the next sample replaces the previous one, QoS 0 is enough; if the sample is an event that does not repeat, QoS 1; QoS 2 only when a duplicate would have accounting or legal consequences.

The three MQTT QoS levels and the Captia Connect criterion for each
LevelWhat it guaranteesReal costWhen Connect uses it
QoS 0At most once delivery. The message is sent and never acknowledged: if TCP breaks midway that value is lost and nobody noticesOne packet per message and no session memory in the brokerContinuous quantities sampled at a short period, where the next reading replaces the previous one: temperature, pressure, instantaneous power
QoS 1At least once delivery. The sender retries until acknowledged, so the message may arrive duplicated but is not lostTwo packets per message and a pending queue in the broker while the client has not acknowledgedData that tolerates no gaps: energy meter readings, machine states, starts and stops, process alarms
QoS 2Exactly once delivery. A four step exchange discards the duplicate before handing it to the applicationFour packets per message and per message state at both ends, with noticeably higher latencyOnly where a duplicate has consequences: billing events, batch records, traceability with legal weight

One practical consequence is worth saying out loud: QoS 1 and QoS 2 are transport guarantees, not capture guarantees. If the sensor is not publishing because its battery is flat, no QoS level invents the data. What QoS 1 guarantees is that what was published arrived; what the last will guarantees is that the silence is noticed. They are different things and both are needed.

Retained and last will: the state MQTT does not have by default

MQTT is a protocol of flow, not of state. A client that subscribes to a topic receives nothing until somebody publishes again, and in a plant that means a machine setpoint published only on change can take days to appear after a node restart. The mechanism that solves it is the retained flag: the broker keeps the last message published with that flag on each topic and delivers it immediately to anyone subscribing later.

Connect leans on retained messages for the cold start. On reconnection, the first wave of retained messages gives it the current value of everything modelled as state (setpoints, active recipe, operating mode) without waiting for the next change. The trade off to watch is that a retained message never expires: if a piece of equipment is removed without clearing its topic, that value is still there months later and looks current. Which is why retained is kept for what genuinely is state, and never used as a substitute for a history.

The last will is the other half. It is a message the client hands to the broker when connecting, with instructions to publish it should the session end uncleanly, that is without an orderly DISCONNECT. When a sensor hangs, loses power or drops off the network, the broker publishes that message and every subscriber learns the point is no longer alive. Combined with a retained device status topic, it is the standard way to tell a stable value from a frozen one, which is the difference between a quiet plant and a plant with a dead sensor nobody has noticed.

Connect treats those notices as first class signals: a publisher disconnection is recorded in the history as an event, with its instant, so that a gap in a series has an explanation rather than remaining a mystery. The full device birth and death convention, with sequence numbers and session state, is what Sparkplug B formalises.

TLS and credentials

MQTT carries no security inside the protocol: it delegates it to TLS underneath and to the broker authentication service above. That is its structural difference from OPC UA, and the reason a misconfigured broker is one of the most common findings in an industrial network audit: port 1883 open, no user and anonymous access allowed is the shipping configuration of more than one improvised deployment.

The Connect criterion has three layers. Channel: TLS to port 8883 with verification of the broker certificate, hostname included, because a TLS session that accepts any certificate protects against a passive observer but not against an impersonator. Identity: a dedicated credential for the Connect client, never the integrator one and never one shared with the SCADA, and where the broker supports it, mutual authentication with a client certificate instead of username and password. Authorisation: an access control list granting that credential subscribe permission on its branch of the hierarchy and nothing else. Connect acquires; it needs no publish permission on process topics and should not have it.

Local buffering: why a broker outage is not a data loss

This is where the position of Connect in the architecture changes the answer, and three kinds of outage that people tend to lump together are worth separating.

Outage between the Connect node and the platform. This is the benign case and the one local buffering covers. Acquisition does not stop: Connect keeps receiving from the broker, normalises and persists to the time series database on the node itself, then synchronises once connectivity returns. Because every sample keeps its timestamp, the delayed block lands where it belongs in the history instead of collapsing onto the arrival instant.

Outage between Connect and the broker. Connect reconnects and, with a persistent session and QoS 1, the broker hands over what was queued during the absence. Whatever travelled at QoS 0 in that interval is not recovered, because it was never stored: that was the decision taken when choosing the level, and it is why QoS 0 is reserved for quantities that are resampled shortly afterwards.

Outage between the publisher and the broker. This is the one nobody covers from the subscriber side. If the sensor cannot reach the broker, the data depends on the device having an internal queue, and many field devices do not. It is the strongest technical argument for putting the broker at the edge, in the same network domain as the publishers: it shortens the fragile leg until it sits inside the plant, where the link is cable and does not depend on a mobile operator.

How Captia Connect behaves under each kind of outage in an MQTT architecture
Where it breaksWhat happens to acquisitionWhat is recovered afterwards
Connect node to the platformIt does not stop. Connect keeps receiving from the broker, normalises and persists locallyThe whole interval, with the original timestamp of every sample
Connect to the brokerReception stops. Connect retries the connection continuouslyWhatever was published at QoS 1 or 2 on persistent session subscriptions. QoS 0 traffic from that interval does not exist
Publisher to the brokerThe topic stops updating. The last will of the device announces the dropOnly what the device itself managed to queue. The drop is recorded as an event
Broker process failureAll MQTT ingestion on that branch stops, whatever the QoSRetained messages and persistent sessions if the broker keeps its store on disk rather than in memory only

What publishes over MQTT in an installation

MQTT does not appear in a plant because somebody chose it on a drawing: it appears because equipment that arrived after 2015 came speaking it. It shows up in four families, and the integration work differs in each.

The first is retrofitted sensing: temperature and humidity probes, vibration monitors, flow meters and pulse counters fitted to equipment that had no instrumentation. These are usually battery powered devices reaching a gateway over radio, and it is the gateway that speaks MQTT. They publish on a long period and, where the manufacturer allows it, their topic is set; where it does not, they publish wherever the firmware decides and the work moves to the mapping.

The second is field gateways and concentrators, the most frequent case in mixed installations. A gateway reads the power analysers, drives and meters in the cabinet over Modbus and publishes the result over MQTT. From the Connect side the origin is MQTT, but it is worth knowing that polling sits behind it: the real period of the signal is set by the gateway, not by the rate at which messages arrive, and that distinction prevents believing in a resolution that is not there.

The third is OEM equipment with a native connector: recent process machines, compressors, chillers, photovoltaic inverters and chargers. They publish a closed set of variables in the hierarchy the manufacturer decided, almost always with the serial number at the root. There is no room to rename at source, so normalisation happens entirely at the edge.

The fourth is the SCADA or supervisory system itself, once an MQTT publishing module has been added to it. It is a convenient tapping point because signals arrive already aggregated and named in operations language, with the usual trade off: data comes at the granularity the SCADA applies, not that of the process.

And a field warning that saves meetings: MQTT on the datasheet does not mean the equipment can be pointed at the plant broker. A good share of connected industrial equipment publishes only to its manufacturer cloud, with the broker address fixed in firmware. In those cases the route is not MQTT but the API that cloud exposes, which comes in through the standard integrations.

From MQTT message to time series

An MQTT message is two things: a topic and a payload of bytes. The protocol says nothing about what is inside, declares no types, declares no units and does not require a timestamp. That indifference is what makes it lightweight and what moves to the edge all the work that OPC UA resolves in its information model. In MQTT, normalisation is not a finishing touch: it is half the project.

The topic is the identity, which is why the hierarchy mortgages the platform

In MQTT the identity of a signal lives in its topic. It is the only stable clue the message carries, which is why a badly designed hierarchy is paid for over years. The three failure modes seen again and again are always the same.

Putting what changes into the topic. A root based on the serial number of the equipment, or on the gateway address, turns replacing a failed sensor into a break in the historical series: the new unit publishes somewhere else and what was one continuous signal splits in two. The topic should identify the measurement point by its place in the plant, not the hardware occupying it today.

Flattening the hierarchy. When everything hangs off a single level with compound names such as plant1_line2_press_temp, wildcards stop working: there is no way to subscribe to the whole of line 2 or to collect every temperature in the plant without listing them one by one. Each new device then demands a configuration change in the acquisition, which is exactly what MQTT was there to avoid.

Mixing data with everything else. Telemetry, commands, configuration and diagnostics under the same branch make it impossible to grant Connect a clean read only permission, and force filtering by content what should be filtered by topic.

The pattern Connect proposes where it can influence the design runs from stable to volatile, with place first and the nature of the data last: company/plant/area/line/asset/type/signal. With that shape, an area filter picks up everything in that area, a type filter picks up all the energy in the plant, and swapping a sensor breaks nothing. Where the broker belongs to somebody else and the hierarchy is imposed, the translation happens at the edge: Connect maintains the correspondence between source topic and the canonical identity of the signal, so vendor naming never propagates into the history. Full hierarchy design, with its governance criteria, is covered in the industrial MQTT broker guide and in the idea of a unified namespace.

The payload and the timestamp

Three payload shapes turn up in the field and all three need an answer. A plain text scalar, one topic per signal with a number inside, is what the simplest devices publish: no type, no unit and no instant. A JSON object with several signals is the norm on gateways and OEM equipment: it groups measurements and often includes its own time field, in whatever format and time zone the firmware decided. And compact binary, typical of radio links with a limited payload, requires the vendor structure to unpack it.

The Connect rule on time is the same as in every other protocol, and it is what makes buffering work: if the payload carries a credible source timestamp, that one wins. Only when it does not is the sample stamped on reception at the edge. The difference shows exactly when it matters: as an outage clears, a batch of queued messages arrives in a burst, and stamping by arrival time would flatten half an hour of process into two seconds of chart. On the word credible there is no flexibility: a device without a synchronised clock that restarts in 1970 after every power cut delivers worse timestamps than the edge, and that check belongs to commissioning, not to later.

What an MQTT message delivers and what Captia Connect does with each part
What arrives in the messageWhat Connect does with itWhich problem it avoids
Publication topicTranslated into the canonical identity of the signal, tied to asset, line and areaVendor naming or a serial number ending up as the name of the signal in the history
Payload with no declared typeType and unit are fixed per signal in the ingestion configuration, onceCounters treated as text, booleans arriving as zero and one, and aggregates that do not add up
Timestamp inside the payloadUsed as the sample instant where the device has a reliable clock, with the time zone normalisedA burst of queued messages piling up at the reception instant and misrepresenting the dynamics
Message with no timestampStamped at the edge on reception, recording that the source provided noneAssuming a temporal precision the device never gave
JSON object with several measurementsSplit into one series per measurement, each with its own unit and stable nameAn opaque history of documents, impossible to aggregate or compare across equipment
Last will notice from the publisherRecorded as a disconnection event at the instant the broker emits itUnexplained gaps in a series and frozen values that look alive

Once that is done, the signal enters the same model as those acquired by any other route and the rest of the platform need not know where it came from. The general framing of that work is in the data acquisition system guide.

When MQTT is not the right route

MQTT solves one problem very well, and precisely for that reason it is often asked for what it cannot give. These are the six scenarios where insisting costs money.

The equipment does not publish and has to be asked. This is the structural limit: MQTT is pub/sub and needs a publisher. An older controller, a power analyser or a meter with a serial port will not push anything anywhere. Reading them means polling them, and that is Modbus, over TCP or over RTU on RS-485. Adding MQTT there means inserting a gateway that polls and republishes: worthwhile if that gateway already exists or will concentrate many devices, not worthwhile as a new box between Connect and equipment Connect already reads directly.

The information model is needed, not just the value. Where the aim is to discover what a machine exposes, with readable names, declared types and per reading quality, MQTT does not have it: the broker moves bytes and does not know what it carries. The ecosystem answer to that gap is Sparkplug B, which imposes type, sequence and device state on top of MQTT. If the equipment already ships an OPC UA server that model is already solved, and building an MQTT layer on top just to obtain it is duplicated work. Which architecture suits a new design is discussed in the OPC UA versus MQTT resource.

The loop is control and has a timing requirement. MQTT runs over TCP and its latency depends on the broker, the queue and the network; neither the protocol nor the broker offers a deadline guarantee. For an interlock, an emergency stop or any closed loop, the route is the plant automation system with its deterministic bus. Connect acquires for decisions and for the history; it closes no loops and must not sit in the critical path of a safety function.

The past has to be retrieved. A broker is not a historian. It keeps the last retained message per topic and the pending queues of live sessions, and nothing else. Whoever connects today cannot ask for last week, because last week is not there. That is exactly the job of time series storage at the edge and on the platform, and confusing the two layers leads to the awkward conversation of discovering, after an incident, that there is no data to reconstruct it with.

The data lives outside the plant. Production orders in the ERP, laboratory results, hourly energy prices or maintenance records are not going to appear on any broker. The route is the standard integrations over REST API, webhooks or CSV.

Nobody governs the topic hierarchy. This is an organisational limit and the one that spoils the most projects. MQTT imposes no structure, so if three integrators publish as they please for two years the broker ends up a set of incompatible topics with the same quantity written four different ways. The protocol will not fix it, and translating at the edge hides the symptom rather than the cause. That work is about context and data governance, and it is the subject of the industrial data contextualisation guide.

Coexistence with the rest of the installation

A normal plant does not speak one protocol, it speaks the ones it accumulated. Connect acquires in parallel over every available route and lifts them into the same time series model, so that in the history a vibration published over MQTT from a gateway, a process signal read from the PLC over OPC UA and a consumption taken from a meter over IEC 870-5-102 all sit at the same level. Nobody needs to know which protocol a value came from in order to use it.

The split that recurs in mixed plants is simple: MQTT covers what was added and OPC UA covers what was already there. Retrofitted sensing, gateways and connected equipment publish; machinery with a controller and the SCADA are read as a client. They do not compete and there is no need to choose. Which one suits as the starting architecture of a new design, and on what criteria, is the discussion in the OPC UA versus MQTT resource, which is where it is developed.

The migration actually seen in the field is not one of protocol but of broker ownership. Many installations start by publishing to the sensor manufacturer cloud because that is what came in the box, and at some point the plant wants the data at home, wants to add equipment from another supplier, or wants to stop depending on an internet link to see its own temperature. The move then is to bring the broker to the edge and repoint the publishers, which is done device by device and without stopping production: the new broker coexists with the old one for the duration, because a repointed publisher stops sending to one and starts sending to the other with nothing else noticing. Equipment with the address fixed in firmware is the exception and stays on its cloud API.

As the MQTT fleet grows and several applications start consuming, the natural evolution is to give the hierarchy semantics: a single topic convention, declared types in the payload and explicit state for every device. That destination has a name, unified namespace, and a specification that grounds it, Sparkplug B. Connect does not require getting there to work, but the tidier the hierarchy, the less translation has to be maintained at the edge.

The platform modules already work on that base, and the engineering project that deploys it is MQTT integration.

Frequently asked questions

Questions about MQTT in Captia Connect

Is a dedicated MQTT broker required, or can Captia Connect use the existing one?
Both are possible. If the plant already has a broker, Connect joins as one more client: it subscribes to the agreed topics with a read only credential and publishes nothing. If there is none, or the existing one sits in a vendor cloud and the plant does not want to depend on that uplink, the broker is brought up on the edge node itself in a Docker container and equipment publishes to a local network address.
Which QoS level should be used for plant data that cannot be lost?
QoS 1, which guarantees at least once delivery and retries until acknowledged. QoS 0 is reserved for continuous quantities that are resampled shortly afterwards, since a lost reading is replaced by the next one. QoS 2 is only justified where a duplicate message has accounting or legal consequences, because it quadruples the packets and adds latency. Note that the effective level is the lower of the one used by the publisher and the one requested by the subscriber.
Is data lost if the broker or the connection to the platform goes down?
It depends where the break is. If the link between the Connect node and the platform drops, nothing is lost: Connect keeps receiving from the broker, normalises and persists locally, then synchronises when the network returns, preserving the timestamp of each sample. If the connection between Connect and the broker drops, reconnection recovers what was published at QoS 1 or 2 on a persistent session, but not QoS 0 traffic. If it is the publisher that cannot reach the broker, only what that device managed to queue on its own is recovered.
Why does the design of the topic hierarchy matter so much?
Because in MQTT the topic is the identity of the signal: it is the only stable clue the message carries. A hierarchy with the serial number at the root breaks the historical series every time a sensor is replaced. A flat hierarchy makes wildcards useless and forces a reconfiguration of the acquisition with every new device. The recommended pattern runs from stable to volatile, with location first and the nature of the signal last.
What are retained messages and the last will for in a plant?
Retained solves the cold start: the broker keeps the last message flagged as retained on each topic and delivers it immediately to anyone subscribing later, so after a restart the current state is known without waiting for the next change. The last will solves silence: it is the message the broker publishes when a client disappears without disconnecting cleanly, and it is what allows a stable value to be told apart from one frozen by a failed sensor.
Do ports have to be opened to the internet to integrate MQTT?
No. The client connection to the broker is always outbound, and when the broker sits at the edge all telemetry traffic stays inside the plant network. The starting criterion is TLS to port 8883 with verification of the broker certificate, a dedicated credential for Connect, and an access control list granting it subscribe permission on its own branch and nothing else.

Related links

Continue from here

This page describes how Captia Connect speaks the protocol. The definition of the term, the topic guide and the engineering service that deploys it live elsewhere.

The other Connect protocols