Finding and Removing Duplicate Signals with the Clay CLI
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Finding and Removing Duplicate Signals with the Clay CLI
Because clay signals create does not deduplicate, a workspace that has been scripted against for a while can hold many copies of the same signal. Grouping on what a signal targets finds candidates cheaply, but identical targeting does not mean identical configuration, so the candidates need a second check before anything is deleted.
What you will build
A script, find-duplicates.sh, that finds candidate groups from clay signals list, fingerprints every candidate with clay signals get, and prints which signal to keep and which to delete for each true duplicate set. Deletion is a separate, explicit step by id.
clay signals list ──► group by (type, entityType, sorted segmentIds) candidates
clay signals get ──► fingerprint (inputs, schedule, filter) true duplicates
──► keep oldest, print the extras, delete by id
AI Prompt
Using the Clay CLI, find duplicate audience signals and delete the extra copies. Requirements: - `clay signals list` returns every signal in one response with no filters or pagination. Group audience signals by `.signal.type`, `.input.audiences.entityType`, and the sorted `.input.audiences.segmentIds`. - Rows from `list` do not carry `schedule`, `filter`, or `signal.inputs`. Two signals with the same targeting can still differ in those, so read each candidate with `clay signals get <id>` and fingerprint `signal.inputs`, `schedule.periodUnit`, and `filter` together. - Only signals with the same targeting and the same fingerprint are duplicates. Keep the oldest by `createdAt`. - Identify signals by their trigger definition id (td_...), never by matching on a name or on the text of a segment id. The literal "ALL" in segmentIds is a sentinel for a whole audience, not a saved segment. - `clay signals delete` is not idempotent: a second delete of the same id returns not_found (exit 6). Read each target back with `clay signals get` before deleting it. - Run the verification step below before finishing.
Prerequisites
- The
clayCLI on PATH, authenticated viaclay login(an OAuth session, not a Public API key) jq,xargs, and bash
Note: JSON samples below are trimmed to the fields relevant to each step. Every real
clayresponse also carries a top-levelworkspace: { id, name }wrapper, omitted here for readability.
1. The script
#!/usr/bin/env bash
# find-duplicates.sh
set -euo pipefail
clay signals list > /tmp/dup-list.json
# Step 1: candidate groups from list alone (cheap, one call)
jq -r '[.data[] | select(.input.audiences)
| {id, type: .signal.type, entity: .input.audiences.entityType,
segs: (.input.audiences.segmentIds | sort | join(",")), createdAt}]
| group_by([.type, .entity, .segs])
| map(select(length > 1))
| map({type: .[0].type, entity: .[0].entity, segs: .[0].segs, members: map(.id)})' /tmp/dup-list.json > /tmp/dup-groups.json
echo "candidate groups: $(jq length /tmp/dup-groups.json); signals in them: $(jq '[.[].members|length]|add // 0' /tmp/dup-groups.json)"
# Step 2: fingerprint every member with `get` (8 in parallel) so variants are not mistaken for copies
mkdir -p /tmp/dup-get && rm -f /tmp/dup-get/*
jq -r '.[].members[]' /tmp/dup-groups.json | xargs -P 8 -I{} sh -c 'clay signals get {} > /tmp/dup-get/{}.json'
jq -s '[.[] | {id, createdAt, fp: ({inputs: .signal.inputs, sched: .schedule.periodUnit, filter: .filter} | tostring)}]' /tmp/dup-get/*.json > /tmp/dup-fp.json
# Step 3: true duplicates = same targeting AND same fingerprint; keep the oldest, list the rest
jq -r --slurpfile fp /tmp/dup-fp.json '
($fp[0] | map({(.id): .}) | add) as $F
| .[] | .members | map($F[.]) | group_by(.fp) | map(select(length > 1))[]
| sort_by(.createdAt) | {keep: .[0].id, extra: (.[1:] | map(.id))}
| "\(.keep)\t\(.extra | join(" "))"' /tmp/dup-groups.json
2. Run it against a workspace
./find-duplicates.sh
Real output on a workspace of 145 Paused signals created over several earlier test sessions (first line, then the first three result rows; each row is keep<TAB>extras):
candidate groups: 18; signals in them: 134 td_0tl867wdCvqFYHygukT td_0tl89ulvYiVNHk3iKHh td_0tl8a967bZh3GotoGeQ td_0tm3k46fNUjjuzDWCNE td_0tm3kv6a8YmFZHfRsfC td_0tlqczkwx9iKcPz6dsj td_0tlqd0dHzvxwPEHmTWK td_0tlqdsdsgmYcS2Zg7yZ td_0tm396lqebxBtvZ8d4n
Of the 18 candidate groups on that workspace, 14 contained signals with differing configurations. Grouping on targeting alone would have treated those differing signals as copies of each other.
3. Try it on a controlled set
Three identical JobChange signals and one variant with a weekly schedule were created against the same segment:
clay signals create --type JobChange --name "TPC19 dupdemo" --input '{"entityType":"CONTACT","segmentIds":["audseg_0tm3lyrDBZzZgPaCsWx"]}'
clay signals create --type JobChange --name "TPC19 dupdemo weekly variant" --schedule weekly --input '{"entityType":"CONTACT","segmentIds":["audseg_0tm3lyrDBZzZgPaCsWx"]}'
After the three identical creates and the one variant, the script's output for that set was:
td_0tmb344mRrYTYpKW4ES td_0tmb346oERyntRFGZ7Q td_0tmb348MCA66vREmRBZ
The first id is the oldest and is kept. The weekly variant, td_0tmb34agZySz9B8oyiY, does not appear: it shares the targeting but not the schedule.
4. Read each extra back, then delete it by id
clay signals get td_0tmb346oERyntRFGZ7Q | jq -c '{id,name,type:.signal.type,input,runStatus}'
clay signals delete td_0tmb346oERyntRFGZ7Q
clay signals delete td_0tmb348MCA66vREmRBZ
Real output of the read-back for the first extra, then each delete:
{"id":"td_0tmb346oERyntRFGZ7Q","name":"TPC19 dupdemo","type":"JobChange","input":{"table":null,"audiences":{"segmentIds":["audseg_0tm3lyrDBZzZgPaCsWx"],"entityType":"CONTACT"}},"runStatus":"Paused"}
{"ok":true}
{"ok":true}
5. Re-run to confirm
After the two deletes, the script was re-run and grep -c for the kept id and the variant id in its output returned 0. The signals themselves are still in the workspace:
clay signals list | jq -r '.data[]|select(.name|startswith("TPC19 dupdemo"))|"\(.id) \(.name)"'
td_0tmb34agZySz9B8oyiY TPC19 dupdemo weekly variant td_0tmb344mRrYTYpKW4ES TPC19 dupdemo
Verify the result
After deleting a set's extras, re-run the script and confirm the kept id no longer appears in any result row:
./find-duplicates.sh | grep -c "td_0tmb344mRrYTYpKW4ES" || true
Expected: the count is 0, because the set has one member left and a one-member set is not a duplicate.
How it works
signals list is a single call but returns reduced rows: it omits schedule, filter, lastRunAt, error, and signal.inputs. That is enough to group by targeting, which makes it a cheap first filter, and not enough to prove two signals are the same. The fingerprint step spends one signals get per candidate (run 8 at a time with xargs -P 8) to compare the full configuration. Trigger definition ids are unique per signal, so deleting by id removes exactly the signal that was read back.
Common issues
Identical targeting with different settings
Signals that watch the same segment can still differ in schedule, filter, or per-type inputs such as look-back or topic providers. Those are variants, not copies, and the fingerprint step keeps them apart.
The "ALL" sentinel in segment ids
A signal watching ["ALL"] watches the whole audience of its entity type, and a segment whose name happens to contain the text "ALL" is a different thing. Group and compare on ids, never on name text.
A second delete returns not_found
Deleting an id that was already deleted returns not_found (exit 6), the same error as an id that never existed. Re-run the audit instead of retrying a delete.
Table-watching signals are not covered
The script selects signals with .input.audiences. Signals that watch a table view carry .input.table and are skipped.
Next steps
- See the reconcile doc in this batch to stop duplicates from being created in the first place.
- Use
clay signals pauseinstead ofdeletewhen a signal might be needed again.
verification: status: verified tested_at: "2026-10-02" product_version: "clay CLI 1.8.0+71cb1bf09c7e" command: "./find-duplicates.sh | grep -c \"<kept-signal-id>\" || true" expected_result: "After a set's extras are deleted, the kept id no longer appears in any result row, so the count is 0."