Skip to content
transcodelyproduct updateingest rulesS3SupabaseCloudflare R2

Ingest Rules: Drop a File in Your Bucket, Get a Ladder Back

9 min read Dimitar Todorov

An ingest rule is a standing instruction on one of your storage origins: when an object matching these filters lands, create this job. S3, Google Cloud Storage, Supabase Storage and Cloudflare R2 post their own notifications to the rule's endpoint, and no server of yours sits in the path. Here is what it does about duplicates, secrets, paused rules and re-uploads.

The most common video pipeline in the world is a bucket, a function, and a transcoding call. A file lands in S3, an event fires, a Lambda wakes up, reads the key, and submits a job somewhere. Search for how to do it and you will find a decade of tutorials, most of them still pointing at Elastic Transcoder, which AWS shut down last November.

The function in the middle is the part everyone rewrites and nobody wants to own. It needs credentials, retries, deduplication, logging, and a place to run. It is the reason “just transcode what lands in the bucket” is a project rather than a setting.

As of API 5.19.0, it is a setting. An ingest rule lives on one of your storage origins and says: when an object matching these filters appears, create this job for it. Your storage provider posts its own notification straight to the rule’s endpoint. We authenticate it, record it, deduplicate it, and turn it into a job. Nothing of yours runs in between.

object lands in your bucket
        │
        ▼
provider notification  ──POST──►  /ingest/{ing_id}
                                      │  authenticate, parse, deduplicate
                                      ▼
                                 storage event          received
                                      │
                                      ▼
                            filters, then CreateJob
                                      │
                                      ▼
                                 storage event          created · skipped · failed

Creating one

A rule belongs to an app and watches exactly one origin. The origin has to be active with read permission, and that is checked when the rule is created rather than at first delivery, because a rule on a write-only origin would accept events forever and never fetch a byte.

curl -X POST https://api.transcodely.com/transcodely.v1.IngestRuleService/Create \
  -H "Authorization: Bearer $TRANSCODELY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "origin_id": "ori_a1b2c3d4e5f6",
    "name": "Inbox uploads",
    "filters": {
      "prefix": "uploads/",
      "suffixes": [".mp4", ".mov"],
      "min_bytes": 1024
    },
    "action": {
      "outputs": [{ "preset": "pst_x9y8z7w6v5" }],
      "managed": true,
      "output_path_template": "{input_dir}/{input_name}/{job_id}"
    }
  }'

The action is the part of a job that can be decided ahead of time: the outputs, where they go (managed for Transcodely hosting, or output_origin_id for a write-enabled bucket of your own), the path template, thumbnails, priority and metadata. It reuses the same output spec a hand-written job uses, so anything you can ask for on a job you can ask for on a rule, and the conversion from a stored action to job parameters is literally the same function the CreateJob RPC runs. The two cannot drift.

The input is what the event supplies: the rule’s origin, and the key the event names.

Three extra path variables let outputs mirror the source layout. For uploads/2026/holiday.mp4, {input_key} is the whole key, {input_name} is holiday, and {input_dir} is uploads/2026. Provider-supplied keys are sanitized before substitution: a leading slash and any .. segment are stripped, so a crafted object key cannot write outside the prefix you configured.

The response carries the rule’s inbound secret in full, once. We keep only enough of it to verify a delivery and to show you its first and last characters. If you lose it, rotate the rule; the previous secret keeps verifying for 24 hours so you can update the sender without dropping an event.

Four providers, three ways to prove a delivery is yours

The endpoint accepts three proofs, because the senders are not equally capable, and which one you get is decided by the transport rather than by taste.

Sender Proof Why
Cloudflare R2, via a queue-consumer Worker Signature You write the request, so you can sign it
Supabase Storage, via a database webhook Bearer header The hook can set one static header
Amazon S3, via an SNS HTTPS subscription ?token= in the URL SNS attaches no caller-controlled header
Google Cloud Storage, via a Pub/Sub push subscription ?token= in the URL Push sets Authorization itself, to its own OIDC token

The signature is HMAC-SHA256 of "<timestamp>.<raw body>" with the secret as the key, carried as Transcodely-Signature: t=…,v1=… with a five-minute window. It is the same header, the same scheme and the same function we use to sign outbound webhooks, so a verifier you already wrote works unchanged in the other direction.

The query token is the floor, and its cost is stated plainly in the docs: a URL travels through more places than a header. Proxies log request lines, consoles display it back to whoever can see the page, and it is what you paste into someone else’s form. We never log it, never echo it in an error and never return it, but everything upstream of us is yours to account for. The public ingest-templates repository ships the wiring for all four providers, including relay variants for S3 and GCS that keep the secret out of the URL when that matters in your environment.

The payload shape is detected from the body. You never declare it. Amazon percent-encodes object keys and writes a space as +; we decode them, so the key the job reads is the real one. For GCS only OBJECT_FINALIZE is acted on. For Supabase the trigger must be INSERT only, and the next section is why.

Duplicates, and the one case that is not a duplicate

Every one of the four sources delivers at least once, and Supabase’s own documentation warns that a database webhook can fire more than once for the same upload. So an object is identified by the tuple (rule, bucket, key, etag), and that tuple is a unique index. A redelivery loses the insert race and is answered 202 with the id of the event that already won. The job carries a second, independent guard: its idempotency key is a hash of the same tuple. Together they mean the processor can crash mid-create, be retried, and still produce exactly one job.

A new version of the same key is a different object and gets its own job, when the source reports a version identity. S3, R2 and GCS always do, as an eTag or a generation. Supabase’s INSERT often carries no eTag at all, and the re-upload itself arrives as an UPDATE, which the rule ignores by design, because an upload also produces an UPDATE once its metadata is written and acting on both would give one upload two jobs. The consequence is that on Supabase, overwriting a key produces no second job. When you want the new bytes transcoded, replay the original event and the rule re-runs against whatever now lives at that key. The Supabase guide walks through it.

Every event is on the record

Each delivery becomes a storage event with a status you can read from the API or from the origin’s page in the dashboard:

Status Meaning
received Stored, not processed yet
matched Claimed for processing; job creation in flight
skipped Deliberately no job. reason says why
created job_id names the job
failed Job creation was refused and will not be retried

A skipped reason is a filter that did not match (filter_prefix, filter_suffix, filter_content_type, filter_size), a bucket_mismatch (almost always a subscription wired to the wrong rule), rule_disabled, or duplicate. A failed reason is the same API error code a CreateJob call would have returned: limit_exceeded, billing_past_due, and so on. Definitive refusals are not retried, because an app over its monthly spend cap stays over it until the cap or the month changes, and retrying would only fill the log. Transient failures on our side back off and retry up to five times.

Two behaviours here are deliberate and worth knowing before you rely on them.

A paused rule still answers 202 and records the event as skipped, so your provider’s subscription does not start failing and backing off while ingestion is switched off. Those objects are not transcoded when you turn the rule back on; nothing retries them on its own. Re-enabling returns events_skipped_while_disabled, the count of objects waiting for you, and ReplayEvent is how you act on them.

Updates merge. Narrowing a rule to a different folder is {"filters": {"prefix": "raw/"}} and nothing else; the suffix, content-type and size filters keep their values. This is not a nicety. A rule is a standing instruction to spend money, and its filters are the only thing between “the objects I meant” and “everything that lands in the bucket”. An update that replaced the filter set wholesale would turn a one-line rename into a rule that transcodes the whole bucket, and the way you would find out is the invoice. Removing something needs an explicit clear_filters or clear_action flag.

Test it before you wire anything

IngestRuleService.Test runs the same matching the live path runs, against a key you supply, and creates nothing. A matched: false answer names the filter that rejected the key; a matched: true answer returns the exact job request the rule would submit. It is the first thing to run when a rule is not producing jobs, and the second is ListEvents, which is the full delivery log.

What it costs

Nothing beyond the jobs it creates. Rules, events, replays and tests are free. A job created by a rule is priced exactly like the same job submitted by hand, and it carries ingest_rule_id and ingest_object_key in its metadata so the job.succeeded webhook tells you which upload it was for.

Where it fits

If your users upload to Supabase, R2, S3 or GCS and you have a function whose only job is to call a transcoding API when a file lands, an ingest rule replaces that function. If the outputs should go back into a bucket you own, set output_origin_id; the R2 guide covers zero-egress delivery from there. If they should be playable straight away, set managed: true and each upload becomes a hosted video with a CDN URL.

Ingest rules are in the dashboard under each origin, and in the API reference. They are not in the SDKs yet; the RPCs are plain Connect calls, and the docs show every one as curl.

Try it on your own footage

Upload a clip, pick a ladder, and read the manifest that comes back. Test encodes up to 30 seconds are free.

Topics

transcodelyproduct updateingest rulesS3SupabaseCloudflare R2

Share this article