Skip to main content
Captia Technology

Protocol supported by Captia Connect

REST API, webhooks and CSV in Captia Connect: the three IT integration routes

For ERP, MES, cloud applications and closed systems: endpoint polling, event reception and file ingestion. Connect polls endpoints at the configured cadence, exposes validated webhook receivers and ingests files on a schedule, and lifts all three into the same time series model.

What it is and the role it plays

What is Standard integrations: REST API, webhooks and CSV and how does Captia Connect use it?

REST API, webhooks and CSV are the three routes Captia Connect uses to integrate systems that do not speak an industrial protocol. Connect polls REST endpoints at the configured cadence, exposes webhook receivers that verify the sender signature and ingests CSV files on a schedule, validating the header against a declared schema, lifting all three into the same time series model as plant acquisition.

How Captia Connect speaks REST API, webhooks and CSV

These three routes cover everything that enters or leaves the platform without passing through a fieldbus: the ERP, the MES, the CRM, the laboratory, the weighbridge, the machine builder cloud portal and the legacy application nobody wants to touch. Connect treats them as acquisition, exactly as it treats OPC UA or Modbus: each one is a connector running on the edge node, inside its Docker container, writing to the local time series database and synchronising to the platform over the outbound Tailscale link. What changes between them is not the destination, it is who starts the conversation and what guarantees the data carries when it lands.

The underlying difference with a plant protocol is that there is no variable being sampled continuously here. There are business records and events with a key, an instant and a set of fields, and most of the problems in this kind of integration are not network problems: they are problems of identity, time zone and schema.

REST API: Connect as the client that asks

With REST, Connect is the one asking. The connector is configured against a base URL over HTTPS on port 443, with certificate verification, and runs a polling cycle whose cadence is a business decision rather than a technical one: if the ERP closes production orders once per shift, polling every fifteen seconds only loads the ERP and brings no decision forward. The useful band runs from one minute to one hour, and longer still for master data that barely changes, such as part numbers, cost centres and shift calendars.

Authentication is settled in one of three ways, in order of how often they actually turn up on industrial sites. An API key in a header, the most common and the most fragile, because it does not expire and usually carries more permissions than the job needs. OAuth 2.0 with the client credentials flow, where Connect requests a token, holds it in memory and renews it before expiry rather than waiting for the first 401. And basic authentication over TLS, which appears on older systems and calls for particular care with the account. The starting criterion never changes: a dedicated user, read permissions, and scope limited to the resources the integration needs.

Pagination is where data goes missing without anyone noticing. A response that returns two hundred records when there are a thousand raises no error at all. Connect walks the pages until the set is exhausted, and the mechanism depends on what the source offers: a cursor or continuation token, which is the reliable one because it does not shift when new records arrive mid traversal, or offset and limit, which does shift and therefore requires ordering by a stable field. A page ceiling per cycle is also set, so that an initial backfill of years of history does not monopolise the connector.

The query is incremental. Connect keeps a watermark with the instant of the last record retrieved and, on the next cycle, asks only for what came after it, with a deliberate overlap reaching back. That overlap exists because business systems write late: a delivery note dated 10:00 may only appear in the database at 10:04, and a window with no overlap would leave it out for good. The price of the overlap is repeated records, which is why duplicate control is not optional: each record is identified by its natural key in the source system combined with its instant, and reingesting that same key overwrites rather than appends. Without that rule, every overlap inflates the totals and the first production report comes out wrong.

On failure the connector discriminates. A 429 with a retry header is honoured as given, a 5xx is retried with exponential backoff and a degree of randomness so that connectors do not synchronise against the same server, and a 4xx that is not a 429 is never retried: it is a configuration or permission error, and the right response is to raise it rather than insist. Meanwhile plant acquisition carries on, because these connectors are independent of one another.

Webhooks: Connect as an event receiver

Webhooks reverse the direction: the business system pushes the event as it happens and Connect exposes an HTTPS receiver that takes it. You gain latency, seconds instead of a polling cycle, and you lose control, because it is the sender that decides whether the data arrives. All the engineering work consists of giving that control back to the receiver.

Origin verification. A webhook URL is public by definition, so anyone who knows it can post a body to it. Connect validates the HMAC signature the sender computes over the raw body with a shared secret, comparing it in constant time and against the bytes as received before deserialising, because re-serialising the object changes the key order and breaks the signature. The signature includes a timestamp, and requests outside a short window are discarded so that a captured legitimate request cannot be replayed later. Where the sender supports it, mutual TLS or an allowlist of source addresses is added. What is never done is accepting an unverified webhook because the sender offers no signing: in that case the correct route is REST polling.

Retries and idempotency. Serious senders retry when they do not receive a 2xx, and that behaviour has a direct consequence for receiver design: Connect acknowledges as soon as it has validated and persisted the event, and processes afterwards. If the receiver processes inline and takes longer than the sender timeout, the sender treats the delivery as failed and repeats it, so the same event is counted twice even though the first attempt finished cleanly. Each event is therefore recorded with the identifier the sender supplies, and identifiers already seen are retained for a window wide enough to cover the whole retry policy of the source. Delivery is at least once; receiver idempotency is what turns it into exactly once as far as the history is concerned.

Where the receiver lives. No inbound port is published on the plant network to receive webhooks. The receiver is exposed on the platform side and the event comes down to the edge node over the same Tailscale link the synchronisation already uses, or else the receiver runs on the node and is only reachable inside the private network of that link. The industrial network gains no internet facing surface from integrating a webhook.

CSV: scheduled file ingestion

CSV survives because it is the only output many laboratory systems, many weighing indicators and nearly all legacy programmes offer. Connect collects it from an agreed location, usually an SFTP directory or a network folder, on a schedule, and treats it as a batch rather than a stream.

Before a single row is read, five things about the format have to be pinned down, and none of them can be guessed reliably. The delimiter, which in spreadsheet exports produced on a machine with Spanish regional settings is a semicolon rather than a comma. The encoding, where UTF-8 coexists with UTF-8 carrying a byte order mark (which, if not stripped, turns the first column name into something that no longer matches the schema) and with legacy code pages that mangle accented characters. The decimal separator, comma or point, which combined with the field separator produces the classic file where 1.234,5 and 1,234.5 mean opposite things. The date format, where 03/04 is ambiguous until somebody states whether it is March or April, and where a missing time zone is the most expensive error of the lot. And the line ending together with the quoting rules, because a text field containing a line break with badly escaped quotes throws out the rest of the file.

The file that one day changes its columns. This is the real failure mode of CSV and it deserves separate treatment, because it does not fail loudly, it fails silently. Somebody updates the source report, inserts a new column in third position, and from that batch onwards the column read as weight starts carrying the lot number. A connector that relies on position keeps ingesting without complaint and the history is contaminated until somebody looks at a chart and the numbers do not add up. Connect avoids this by relying on the declared header rather than the column index: the batch schema is compared against the first row of the file before anything is processed. If a new, unused column appears, it is ignored and logged. If a column in use disappears or is renamed, the entire batch is rejected and raised; a file is never ingested half way. That policy, and the conversation with the owner of the source system that goes with it, is what the industrial data contracts guide develops in general terms.

Two operational details remain, and they prevent most incidents. The first is atomicity: a file still being written must not be read, so the source writes under a temporary name and renames on completion, or drops a marker file when it closes, and the connector only picks up what is complete. The second is batch idempotency: every processed file is identified by name and by a fingerprint of its contents, so reprocessing the same export, which happens every time somebody drops yesterday's file again, does not duplicate a single row.

The three standard integration routes of Captia Connect compared by initiative, latency and guarantees
RouteWho starts itUsual latencyMain riskWhen it is the right one
REST APIConnect queries the source systemThe polling cadence, one minute to one hourLoss through mishandled pagination and duplicates from window overlapThe source has an API and the data must be complete and reconcilable
WebhooksThe source system pushes the eventSeconds from the event occurringDelivery neither guaranteed nor unique, and unverified originThe event is discrete and its value decays fast, such as a state change or an alarm
CSVConnect collects the deposited fileThe batch cycle, hourly to dailyColumn changes, encoding and ambiguous date formatsThe source offers no API and a file export is all there is

Which systems are integrated over these routes

There is no terminal block or screened cable here. There is software, and it is worth sorting it into families because the integration work differs sharply between them.

Management systems are the majority case: ERP, MES, CRM and corporate suites such as SAP. They supply the context axis the plant does not have: the production order, the part being made, the batch, the shift, the customer and the cost. Almost all of them offer a REST API, and the more modern ones also offer webhooks for state changes. This is the integration that turns a consumption curve into consumption per unit produced, and the service that deploys it is ERP integration.

Laboratory and quality systems deliver test results several hours after the process that produced them. Many export CSV to a folder and little else. The data is low frequency and high value, and its integration is nearly always by file.

Instruments with file or line output: weighbridges and weighing systems, titrators, bench analysers and test equipment. They emit a ticket or a line per measurement, often over a serial port that a concentrator dumps to a file. They are integrated as CSV, with the particularity that the file grows at the end and the connector has to track how far it read.

Legacy systems with no industrial protocol: the warehouse application written twenty years ago, the maintenance programme whose support ended long since. They usually have a database behind them and, with luck, a view or a procedure somebody can expose as a thin API or dump to a file on a schedule. If they have none of that, see the limits section.

External services over API: the electricity retailer and hourly energy prices, the weather forecast that feeds photovoltaic generation forecasting, machine builder cloud portals publishing telemetry from their own equipment. Here Connect is a client of a third party and controls neither availability nor rate limits, which makes retry handling rather more than a detail.

Outside these families sits what genuinely is plant, and the other routes exist for that: PLCs and controllers over OPC UA, drives, power analysers and controllers over Modbus, sensors and gateways that publish on their own over MQTT.

From business record to time series

A PLC node delivers a value with its type and its instant. An ERP record delivers a row with twenty fields, of which Connect needs to identify three things: which one is the instant, which one is the identity and which ones are magnitudes with a unit. That is the whole mapping, and getting any of the three wrong produces a history that looks correct and is not.

The instant. The date field declared by the source system is used, not the moment Connect received the record, for precisely the same reason as in a plant protocol: a late batch or a retried webhook would otherwise be bunched at arrival time and misrepresent the dynamics. The field is normalised to an instant with an explicit time zone. When the source delivers local time with no zone, which is the norm in a CSV, the zone is declared in the connector configuration: without that declaration, the October clock change produces an hour that appears twice and the March one an hour that does not exist, and both wreck consumption aggregates precisely on the night shift.

The identity. Every record needs a stable key in the source system: the order number, the event identifier, the batch code. That key is what makes ingestion idempotent and it is also what allows the join with the plant. If the ERP says order 48213 was made between 06:12 and 09:40, that key is what turns the power curve of that line into the consumption of that order.

Magnitudes and their unit. Neither JSON nor CSV declares units. A field called energy may arrive in Wh or in kWh, and the difference is three orders of magnitude that nobody spots when the chart scale adjusts itself. The unit is fixed in the mapping, documented, and validated against a plausible range, so that an impossible value raises an alert instead of entering the series.

There is also a difference of nature worth making explicit. Much of this data is not a sample of a continuous signal, it is a state with validity: the order in progress, the active shift, the part being produced. These are modelled as steps holding until the next change, never interpolated, and that distinction is what prevents ramps appearing between two events that were never a ramp. The general work of giving a signal its context is covered in the industrial data contextualisation guide, and the general capture framework in the data acquisition system guide.

What an API, webhook or CSV record delivers and what it becomes inside the platform
What the record carriesWhat is done with itWhich problem it prevents
Date field from the source systemNormalised to an instant with explicit time zone and used as the sample timestampA late batch or a retry being bunched at arrival time
Natural key of the recordUsed as the idempotency key and as the link to the asset, line or orderDuplicates from window overlap or from a webhook sender retry
Numeric field with no declared unitGiven a unit and scale in the mapping and validated against a plausible rangeFactor of a thousand errors between Wh and kWh that nobody sees until month end
Column header of the fileCompared with the declared schema before processing and used to accept or reject the batchAn inserted column shifting values and silently contaminating the history
Business state with validityModelled as a step holding until the next change, without interpolationInvented ramps between two events that were only two state changes

When these routes are not the right ones

Standard integrations are convenient, which is why they get overused. These are the cases where the answer has to be no.

The data is a process signal from plant equipment. This is the most frequent and the most expensive mistake. If what is wanted is the power draw of a line, the temperature of a furnace or the state of a machine, the route is not an intermediate API but the equipment itself, over OPC UA, Modbus or MQTT. Polling an API every minute is not equivalent to a subscription: everything happening between calls is lost, the aggregation applied by the intermediate system is inherited, and the whole thing depends on that system staying alive. A two second start-up peak does not exist in an API polled once a minute.

Seconds of latency are needed and only REST is available. Lowering the polling cadence has a ceiling imposed by the source system, which starts returning 429s or simply slows down for everybody. If the source offers a webhook, that is the route. If it does not and latency is genuinely a requirement, the honest conclusion is that the data has to be taken somewhere else.

Guaranteed delivery is needed and only a webhook is available. A webhook that never arrives leaves no trace anywhere: if the sender does not retry, or exhausts its retries during a maintenance window, the event simply did not happen. For any data that will later be billed, audited or declared, a webhook cannot be the only source; it is paired with periodic REST reconciliation that compares what was received against what the source says it emitted.

The CSV is produced by a person by hand. A manual export gets forgotten during holidays, saved under a different name, opened in a spreadsheet that reformats the dates on saving, and eventually changes its columns. If the source system has an API or read access to its database, that path costs more at the start and less every month. An automatic CSV deposited by a scheduled process is a different thing and is a solid route.

The full file keeps growing. A CSV that resends the entire history on every cycle forces a comparison against everything already ingested and ends up dominating connector processing time. As volume grows, either the source must provide incremental exports or the integration moves to an API.

The source system offers neither API nor export. Closed applications with no data interface and no access to their database. There is no standard integration here over any of the three routes, and presenting one as if there were misleads the client. What applies is bespoke development against that system, which is OT and IT integration work rather than connector configuration.

And the limit that is not technical, the one that actually breaks these projects: without an agreement on the schema there is no stable integration. Somebody on the source system side has to own the duty of warning before a field, a unit or a date format changes. Where that agreement does not exist, the integration does not fail on day one, it fails the day somebody updates a report, and the only available defence is the one set out in industrial data contracts: validate the schema on every batch and reject rather than ingest blind.

Coexistence with the rest of the installation

These three routes do not compete with plant protocols, they intersect with them. The plant supplies the time axis, dense and high frequency; the business systems supply the context axis, sparse and heavy with meaning. Connect acquires from both sides in parallel and lands them in the same time series model, which is what allows a question such as how much energy it cost to make this part to have an answer without anybody exporting anything by hand. That intersection is the subject of OT and IT convergence.

Within one installation it is normal for the three routes to coexist, and not through disorder: each system offers what it offers. The corporate ERP is polled over REST every fifteen minutes, the warehouse management system pushes a webhook when a dispatch is closed and the laboratory leaves a CSV every night. Connect runs the three connectors independently, so a source that goes down does not drag the others with it or halt plant acquisition.

The usual maturity path walks the three in this order. It starts with CSV, because it can be set up in an afternoon and proves the value without asking anything of anybody. It moves to REST as soon as the manual file starts failing or the volume becomes a nuisance, and that step usually coincides with the IT department joining the table. And webhooks are added last, only where latency justifies them, keeping REST polling underneath as a reconciliation net. Replacing polling with the webhook instead of layering them is the shortcut paid for with the first lost event.

The platform modules already work on that base, the full catalogue of routes sits in the Connect protocol index and the engineering project that connects plant and management is ERP integration.

Frequently asked questions

Questions about Standard integrations: REST API, webhooks and CSV in Captia Connect

How often does Captia Connect poll a REST API on an ERP?
The cadence is set by the decision the data supports, not by what the server can take. The usual band runs from one minute to one hour, and longer for master data that barely changes, such as part numbers or shift calendars. Connect polls incrementally, asking only for what follows the last watermark with an overlap reaching back, so that records the source system wrote late are still picked up.
How is the same record prevented from being loaded twice?
With an idempotency key. Each record is identified by its natural key in the source system combined with its instant, and reingesting that same key overwrites rather than appends. This covers both the deliberate overlap of the REST query window and the retries of a webhook sender, which are the two real sources of duplicates.
Does a port have to be opened on the plant network to receive webhooks?
No. The receiver is exposed on the platform side and the event comes down to the edge node over the outbound Tailscale link, or else the receiver runs on the node and is only reachable inside that private network. The industrial network gains no internet facing surface from integrating a webhook.
What happens when the CSV changes its columns overnight?
The batch is rejected and raised rather than ingested shifted. Connect validates the file header against the declared schema before processing any row: a new column that is not used is ignored and logged, but if a column in use disappears or is renamed nothing is ingested. It is the only way to stop an inserted column silently contaminating the history.
Can a plant protocol be replaced with a manufacturer REST API?
Not for process signals. Polling an API every minute is not equivalent to subscribing to a variable: everything happening between calls is lost, the aggregation applied by the intermediate system is inherited, and availability depends on that system. For power, temperature or machine state the route is OPC UA, Modbus or MQTT against the equipment itself.
What is needed on the client side to start an integration of this kind?
Three concrete things: a read only user with its authentication method, documentation of the resource or of the file columns including units and time zone, and an owner on the source system who warns before a field or a format changes. The last one decides whether the integration still works a year from now.

Related links

Continue from here

This page describes how Captia Connect speaks the protocol. The definition of the term, the topic guide and the engineering service that deploys it live elsewhere.

The other Connect protocols