Verilyze Research

Publication and event: how to separate an information stream from a reputational incident

Reputation analysis often starts with a feed: how many items, where they ran, what share was negative. That frame is bad at telling distribution noise from the event itself. This note records how Verilyze separates a publication, a claim and an event — and why that distinction must not be hidden inside a dashboard.

Executive summary

A publication is a document: text, URL, source, time. A claim is a statement inside the document that can be supported or rejected. An event is an analytical assembly of several documents around one incident, statement or decision.

If every publication is counted as a separate event, reposts and syndication look like escalation. If clustering is too aggressive, distinct incidents collapse into one. The right unit of analysis for a reputation desk is an event with evidence, not a row in a feed.

This piece does not measure Russian media, Telegram or a given industry. It has no document count, source share or September 2026 time series, because no such corpus was released for publication. It is the frame Verilyze uses in the product; empirical cuts will be published separately when they can be described honestly.

Key findings

  • A publication is not an event: one incident produces many documents, and one document can mention several events.
  • A mention counter without deduplication and event assembly inflates scale and confuses distribution with what happened.
  • An event needs a type, a date or period, related entities and a list of evidence; without evidence it is a hypothesis, not a fact.
  • An official release, a repost and an analytical write-up play different roles in an event even when the wording looks similar.
  • An LLM summary of a document cluster does not replace the primary source and must not be the only layer an analyst sees.
  • A missing publication in the available corpus does not prove that the event did not happen.

Scope

Period
not applicable: this is not a corpus cut for a calendar interval
Document count
not applicable; no corpus was exported
Source count
not applicable
Language
Russian and English editions of the note; the object of analysis is the method, not a language corpus
Region
not bounded; examples are general and not tied to an unmeasured “whole market”
Objects
the notions of publication, claim, event and evidence
Export date
not applicable
Material type
Verilyze Research methodological note

Method for this note

The note follows the Verilyze loop described in the methodology: collection, normalisation, deduplication, entities, events, claims and evidence. It is not an experimental benchmark and not a blind vendor comparison.

Examples in the text illustrate document roles. They are not hidden client cases, and they are not product demo storylines presented as something that happened.

Three layers: document, claim, event

A document answers “what was published?”. It has a carrier, a URL, a time and a text. A claim answers “what is being said?”: that company X received an order, that person Y resigned, that an incident happened at site Z. An event answers “what happened in the world these documents are writing about?”.

Mixing the layers produces typical monitoring errors. Ten reprints of one release look like ten events. Two different incidents with similar headlines look like one. A negative quote inside a neutral report looks like the outlet’s own stance.

Publication roles inside one event

Even when documents describe one incident, they do different jobs. At minimum it is useful to distinguish:

  • Primary source: a public statement by a participant, regulator, court, company or witness.
  • Distribution: a repost, syndication or short recap with no new facts.
  • Interpretation: an analytical write-up, a column, an assessment of consequences.
  • Reaction: a comment by a related entity, a denial, a clarification.

A desk needs all four roles, but they must not share one counter. A rise in distribution without new facts is channel dynamics, not proof that another event occurred.

How an event is assembled

Assembly starts with documents that share entities, time and incident type. Obvious duplicates are dropped. The remaining documents are grouped around an event hypothesis. The hypothesis becomes working only when there is evidence: links to original materials, not only a model-generated cluster title.

An event should have a type, a date or period, related entities and a list of evidence documents. If source dates disagree, the event gets an interval, not an “average date”. If entities are ambiguous, the merge must not pretend to be finished.

Where automatic clustering fails

  • Identically named companies and people land in one cluster because the string matches.
  • Serial events (“another inspection”, “another outage”) collapse into one although they are distinct episodes.
  • A denial and the original allegation look like confirmation because the entity set matches.
  • An old incident reappears in a new piece and inherits the reprint date instead of the incident date.
  • A model summary adds details that appear in none of the sources.

That is why material clustering is a human-review zone. The system can propose a cluster; the analyst decides whether it is one event by opening the primary sources.

What counts as evidence

Evidence is not “the model is confident”. It is material that can be checked: URL, text, date, source. Several independent primary sources are stronger than many reposts of one text. A missing URL or a deleted page lowers confidence even if the summary still reads coherently.

Where it is legally and technically permissible, an event card should lead to public primary sources. If a source has disappeared, that is a corpus limit, not proof that the claim is false or true.

Limitations

This study reflects the Verilyze methodological frame as of the publication date. It is not a complete picture of every publication on the internet. It has no corpus export, no channel table and no measurement of signal speed.

  • The note must not be used as mention, tone or risk statistics for an industry.
  • Document-role illustrations must not be mapped onto specific clients or product demo objects.
  • The frame depends on available open sources; closed channels are out of scope.
  • NER, deduplication and LLM errors remain and are described in the methodology.

Publication metadata

Published
16 September 2026
Updated
16 September 2026
Author / desk
Verilyze Research
Methodology version
1.0
Open demoContact

Request an invite

Leave a request — we will get in touch and open a workspace so you can evaluate Verilyze on your own tasks.

Request sent

We will contact you at the email you provided.

Write to us

Describe the question — we will reply to the email you provide.

Message sent

We will reply to the email you provided.