Pipeline at a glance
- 1Collect
- 2Normalise
- 3Deduplicate
- 4Extract entities
- 5Assemble events
- 6Build links
- 7Analyse tone
- 8Find anomalies
- 9Score risk
- 10Conclusions and evidence
Steps may run partially, iteratively and with different completeness depending on the task and data availability. A skipped step is not filled with an invented result.
Data collection
The input is a document from an open source or a user-supplied item. When a field is available, the system stores the source, collection time, URL, raw text, available metadata and a document identifier. The primary source is kept insofar as that is applicable and lawful: so a conclusion can be checked, not only read as a paraphrase.
Which venues are in scope, and where coverage ends, is described on the sources page. A missing document in the corpus does not mean the event did not happen.
Normalisation and deduplication
Identical and near-identical publications, reposts and syndicated reprints should not look like independent events. Normalisation brings dates, source names and technical noise into a comparable shape; deduplication links repeats to a canonical document when that can be done with enough confidence.
A precise deduplication success rate is not published: without ongoing measurement that number would be advertising. Errors run both ways: a duplicate may remain, and distinct items may be merged too aggressively.
Entities
The system extracts the objects the corpus talks about. Types include company, person, organisation, location, product, topic and other entities when they are recognised stably in the text.
A matching name does not always mean a single entity. Namesakes, identically named companies, abbreviations and translations have to be resolved by context, links and aliases — not by the name string alone.
Entity resolution in Verilyze is the mapping of mentions to an object card. If confidence is not high enough, a mention should not be treated as finally merged with the object. NER errors and name ambiguity are known limits, not exceptions.
Events
A publication is not an event. Several documents can describe one incident, statement, inspection or communication. An event is an analytical assembly: type, date or period, related entities, evidence from sources and, when used, a significance score.
How to separate a publication stream from a reputational incident is discussed in “Publication and event”. That piece is a methodological note, not a measurement of a specific market.
Claims and evidence
The core Verilyze requirement is that a conclusion is tied to sources. Evidence lets you open the original materials and check what a claim stands on. Missing evidence lowers confidence: that conclusion must not be treated as an established fact.
An LLM summary is not primary evidence. A model digest can help navigate the corpus, but it does not replace the publication text, date, URL and context. If the summary disagrees with the source, the source wins.
Tone
The base tone scheme is positive, negative and neutral. The score is built on a publication or a span, not on a single marker word. Context beats a lexicon: an opponent’s quote, irony, negation and a recap of someone else’s complaint often look “negative” lexically without being the author’s stance.
Irony, dense wording and mixed texts can be misclassified. Tone is an analytical feature, not a moral judgement of the object and not a legal qualification of the publication.
Risk
Three layers need to stay distinct. A fact is what follows directly from a source and can be checked. An analytical feature is tone, an anomaly, a link, a spike. A risk score is an aggregated indicator built from several signals.
The risk score aggregates several analytical signals. The concrete model and weights may change as the product develops; the resulting score should be read as an analytical indicator, not as an objective probability of an event.
The score formula is not published while it is still changing and not locked by measurement. A fictional formula would be worse than an honest refusal to disclose it.
Anomalies
An anomaly is a departure from an object’s usual background, not any mention. In the current loop, anomalies include a spike in publication volume, an unusual source, a sharp tone shift, an atypical entity pairing and a deviation from the object’s baseline.
An anomaly is a signal to investigate, not proof of an incident. A spike may be a repost wave, a calendar event, a collection error or a real escalation; an analyst has to tell those apart using evidence.
Forecasts
Forecast components estimate possible continuations of the current information dynamics from available data. That is a model estimate, not a guarantee of a future event.
Verilyze does not predict the future as an established fact. The forecast horizon is days and weeks of information dynamics, not a legal, financial or operational outcome. Sparse data does not become more accurate because the wording sounds confident.
Confidence and uncertainty
There is not yet a single numeric confidence metric for every system output. Automatic results vary in reliability: it depends on corpus completeness, source quality, entity unambiguity, available evidence and how hard the wording is.
If the system shows a confidence level for a particular conclusion, read it as an internal indicator, not a statistical guarantee. A missing indicator does not mean the conclusion is reliable.
Human verification
An analyst should check material conclusions against the primary source: date, context, related objects and whether sources are independent. An automatic rating, tone, risk score or forecast does not replace professional review when a management, legal, HR or reputation decision hangs on the result.
Material decisions must not be taken on AI output alone. Verilyze is an informational and analytical loop, not an automatic adjudicator of facts.
Limitations
- Incomplete source coverage: the corpus reflects materials available to the system, not every publication on the internet.
- Deleted and edited publications may disappear from the primary source after collection.
- Language nuance, irony and mixed texts raise classification error.
- NER errors, entity merges and false links are possible.
- Tone errors and LLM summaries are possible and should be checked against the source text.
- Incomplete context: the system does not see closed discussions or materials outside the corpus.
- Missing data ≠ missing event.
Data and privacy
User-supplied materials are used to complete the task, operate the service and for purposes set out in the terms. Uploaded files, documents, texts and links are not used to train Verilyze’s own models and are not used to independently fine-tune third-party models.
This page does not claim a broader training policy than the legal documents. See the privacy policy and terms of use.
Methodology versioning
This page describes methodology version 1.0. Material changes to the loop will raise the version and the updated date. Wording fixes that do not change the model may update the date only.