Two editors tag the same interview on the same afternoon. One writes “press conf.” The other writes “Presser, Q3 earnings.” A third asset from the same event gets no tags at all because the ingest was rushed.
None of that is a technology failure. All three tools worked exactly as designed.
This is what a missing metadata management process looks like in practice. Not a system that broke, but a system nobody defined. And the cost compounds quietly, because every untagged or inconsistently tagged asset stays wrong until someone goes back and fixes it, which almost nobody does.
Why does metadata quality decay over time?
Metadata does not stay accurate on its own. It drifts.
New shows launch with new tag conventions. A rights team adds fields the archive team never sees. Someone changes a controlled vocabulary and 40,000 legacy assets quietly stop matching it. Staff turn over and take undocumented conventions with them.
The result is an archive where recall depends on which era an asset was ingested in. Editors learn which searches work and stop trusting the rest, which is the point where a library becomes storage again.
A process exists to slow that decay and catch it when it happens.

What is a metadata management process?
It is the set of rules, roles, and checkpoints that govern how descriptive information about your media is created, validated, stored, updated, and retired.
Four parts make it real rather than aspirational.
A schema. The fields you capture, which are mandatory, and what format each one takes.
A controlled vocabulary. The approved list of values for those fields, so “presser” and “press conference” cannot both exist.
Ownership. A named person accountable for each stage, not a department in general.
Measurement. A recurring score that tells you whether the first three are actually being followed.
Automation accelerates this. It does not replace it. Automated metadata tagging with governance scoring applies your rules consistently at volume, but the rules still have to exist first.
What are the stages of the metadata lifecycle?
| Stage | What happens | Common failure |
|---|---|---|
| Plan | Define schema, vocabulary, and mandatory fields | Copied from a vendor template, never adapted |
| Capture | Metadata created at ingest, manually or automatically | Rushed ingest leaves fields empty |
| Enrich | AI adds transcripts, topics, recognition, compliance tags | Enrichment lives outside the MAM |
| Validate | Completeness and quality checks before the asset moves on | No gate, so errors travel downstream |
| Use | Search, clip, publish, license, report | Editors work around bad data instead of reporting it |
| Maintain | Audits, vocabulary updates, backfill, retirement | Nobody owns this stage at all |
The last row is where most programmes fail. Teams design capture carefully and then treat maintenance as optional, which guarantees the decay described above.
Why should you design the taxonomy before choosing a tool?
Because the tool will happily automate whatever structure you give it, including a bad one.
Start from retrieval. List the ten searches your team runs most often and the five reports leadership asks for. Those queries define which fields are mandatory. Everything else is optional or does not belong.
Keep the mandatory set small. A schema with 60 required fields gets ignored at 2 am during a breaking news cycle. A schema with eight gets filled.
Then write the definitions down. An undefined field is a field that means something different to every person who touches it.
Which metadata standards should you build on?
Do not start from a blank page. Established standards give you a structure that other systems and partners already understand.
| Standard | Best suited to |
|---|---|
| EBUCore | Broadcast production and exchange, built as an extension of Dublin Core |
| PBCore | Public media archives and cultural collections |
| SMPTE metadata dictionary | Technical and production metadata in professional workflows |
| IPTC Video Metadata Hub | Cross-platform news and video description |
| IAB Content Taxonomy | Advertising, brand safety, and contextual classification |
| Dublin Core | A minimal descriptive baseline when nothing heavier is needed |
Most media organisations end up with a hybrid. A recognised base standard for interoperability, plus a documented local extension for the fields only your operation cares about.
Who actually owns metadata?
This is where most processes quietly collapse, because metadata sits between departments and belongs to none of them.
A workable model assigns four roles. A data steward owns the schema and vocabulary. Ingest operators own capture at the point of entry. Editorial and compliance leads own accuracy for their own content types. Engineering owns the integration and the write-back path.
Name individuals, not teams. “Media operations owns metadata” is how a field goes unmaintained for three years.
How do you measure metadata quality?
Score four dimensions and review them on a fixed cadence.
Completeness. The percentage of assets carrying every mandatory field. This is the fastest signal and the easiest to automate.
Consistency. How often values fall inside the controlled vocabulary rather than free text.
Accuracy. Sampled verification against the actual content, since a filled field is not automatically a correct one.
Timeliness. How long after ingest an asset becomes fully described and searchable.
Governance dashboards make these visible rather than anecdotal, which changes the conversation from opinion to evidence. The same discipline applies to any structured data programme, which is why data labeling, wrangling, and validation practices map closely onto media metadata work.
What do you do about the legacy archive?
Backfilling everything at once is rarely affordable or necessary. Prioritise instead.
- Assets with active licensing or resale potential
- Content tied to recurring formats, anniversaries, or returning talent
- Anything under a compliance retention or accessibility obligation
- Material your team has searched for and failed to find in the last year
- Everything else, batched and processed when capacity allows
That last search-failure list is the most useful and the most overlooked. Log failed searches for a month and the archive will tell you exactly which decade to enrich first. Batch enrichment through managed media enrichment services is usually more economical than pulling editorial staff onto retroactive tagging.
How often should you audit?
Set three cadences and hold them.
Weekly, review completeness scores for new ingest. Monthly, sample accuracy across content types and check for vocabulary drift. Quarterly, review the schema itself against new formats, new platforms, and new regulatory requirements.
Anything left to “when we get time” does not happen. Put the quarterly review in the calendar with a named owner before the programme launches.
What about the usual objections?
“Our editors know where everything is.” Institutional memory is not a retrieval system. It leaves when people do, and it does not scale to a library nobody has personally worked on.
“AI tagging means we don’t need governance.” Automation applies rules faster. It also propagates a bad taxonomy faster. Consistent output against the wrong schema is still the wrong schema.
“This is a big project.” It does not have to start big. One content type, eight mandatory fields, one named owner, and a monthly score. Expand once that holds for a quarter.
Frequently asked questions
What is the difference between metadata management and asset management? Asset management stores and moves files. Metadata management governs the descriptive information that makes those files findable, compliant, and reusable.
Can metadata management be fully automated? Capture and enrichment can be largely automated. Schema design, vocabulary decisions, and quality accountability remain human responsibilities.
How many fields should be mandatory? Fewer than most teams expect. Anchor the mandatory set to your most common searches and reports, then leave everything else optional.
How does this connect to compliance? Content-level tagging supports discovery and governance. Broadcast compliance logging and monitoring proves what went to air. Metadata governance strengthens the first and feeds evidence into the second.
Does a metadata process require a new platform? No. It requires defined rules, named owners, and enrichment that writes into the systems you already run.
Where Digital Nirvana fits into this
Digital Nirvana approaches metadata as a governed process rather than a tagging feature. MetadataIQ handles automated tagging, quality scoring, rules-based validation, and native integration with Avid MediaCentral and Grass Valley, so the validate stage of the lifecycle becomes a real gate rather than an intention.
Teams needing individual capabilities can call AI microservices for recognition, OCR, and scene description through APIs. Transcript-driven fields come from transcription, captioning, and localization workflows, and ongoing review capacity for accuracy sampling is supported by human-in-the-loop AI operations.
Why the domain expertise matters here
Metadata governance fails on specifics that generic data frameworks never address. Frame-accurate timecoding. Rights windows. Sponsor reporting fields. The difference between a compliance flag and an editorial note. Which field a topic tag belongs in so it survives a MAM migration.
Digital Nirvana has run these workflows inside large broadcast environments for years, including networks such as CBS, Fox, and Sinclair, across media and broadcasting AI operations. The customer success stories repeat the same lesson: the organisations that win are the ones with a defined process behind the automation.
Conclusion
A metadata management process is not paperwork. It is the difference between an archive that answers questions and one that only stores them.
The organisations retrieving content in seconds are not the ones with the best models. They are the ones who decided what a good record looks like, named someone to own it, and checked the score every month.
Start with one content type. Define eight fields. Measure completeness for 30 days. Expand from there.
Want to pressure-test your current schema against a live workflow? Book a 15 minute governance review and bring your ten most common searches.
Key takeaways
- Metadata quality decays by default. A process exists to slow the decay and catch it early.
- Design the taxonomy from your most frequent searches and reports, not from a vendor template.
- Build on a recognised standard such as EBUCore, PBCore, or IPTC, then document your local extension.
- Keep the mandatory field set small. Short schemas get filled, long ones get skipped.
- Name individuals as owners for schema, capture, accuracy, and integration. Departments do not maintain fields.
- Score completeness, consistency, accuracy, and timeliness on a fixed weekly, monthly, and quarterly cadence.
- Prioritise legacy backfill using logged search failures, licensing value, and compliance obligations.