Source types
The types below are ones Verilyze can work with when the materials are publicly available and technically reachable for a given task.
- Media
- Public items from news and trade outlets, news sites, publications and available metadata. Coverage of a given outlet depends on the site, language and whether the item can be obtained lawfully.
- Telegram
- Publicly available channels and materials, when they are technically reachable. Verilyze does not claim access to private channels, closed chats or personal correspondence.
- VK and social networks
- Public materials available through official platform mechanisms or other lawful ways of obtaining public data. Fields and completeness depend on the API, the object’s publicity and platform limits.
- Web
- Company sites, official publications, topical resources and other public web pages that can be collected for a task. Availability depends on the resource, robots limits and URL stability.
- Blogs, forums and communities
- Public discussions and authored posts when they are available and connected for a task. This is not a catalogue of “every forum on the internet” and not a promise of a complete historical archive.
- User-supplied materials
- Documents, files, links, texts and other materials the customer provides for an analytical task. Rights in those materials, and the lawfulness of handing them over, remain with the customer.
What data is extracted
Source fields and Verilyze-derived analytical features need to be kept apart.
Source data, when available: publication text, title, URL, publication date, source identifier, available public metrics and other metadata returned by the resource.
Derived Verilyze features: mentions, entities, events, tone, relations, anomalies, risk scores and model conclusions. URL and publication date are source data. Tone and risk are analytical outputs of the system.
If a source did not return a field, Verilyze does not invent a value. A missing field in a card does not prove the field never existed.
Coverage and collection limits
Verilyze does not publish a coverage percentage for the internet, Telegram, VK or the media, because that number cannot be honestly reduced to a single figure. Coverage depends on the platform, geography, language, APIs, rate limits, changing access rules and the availability of a given source.
- A platform may restrict its API, change the rules or withdraw materials that used to be public.
- A publication may be deleted, edited or moved after collection.
- Regional and language slices are not “the whole information field”.
- Outages, blocks and rate limits affect corpus completeness.
If measurable coverage cuts appear later, they will be published as a table with a date, method and limits. Until then, a polished percentage would be fiction.
What Verilyze does not collect
- Verilyze does not claim access to private personal correspondence.
- Verilyze does not guarantee complete coverage of the internet, all media, all of Telegram or all of VK.
- Verilyze must not bypass authentication or technical protection measures, or obtain private data that was not lawfully provided.
- A source type in the interface does not mean that every object on that platform is already being collected.
These limits are aligned with the notice on processing data from open sources and the terms of use.
Custom sources
A user can add their own documents, files, links and texts to a task. That is the primary way to extend the corpus with materials that are not in the open collection.
Connecting extra source types or individual venues for a specific task is currently a custom setup, not a self-serve catalogue in the MVP interface. If you need a particular source, write to us or request an invite.
Type summary
| Type | Examples | Public data | Limits |
|---|---|---|---|
| Media | news sites, trade outlets | Yes, if the item is publicly available | depends on the site, language and access |
| Telegram | public channels | Yes, available sources only | no private channels or personal chats |
| VK and social | public materials | Yes, within publicity and API limits | API, publicity, platform rules |
| Web | sites and official publications | Yes, if the page is reachable | depends on the resource and URL stability |
| Blogs / forums | public discussions | Yes, if connected for the task | not a complete platform archive |
| Own materials | documents, files, links, texts | supplied by the client | access rights and lawful transfer sit with the client |
Legal and ethics
Collecting from open sources does not cancel the need for a lawful purpose, proportionate processing, or the limits set by platform owners. Users must not set tasks that require bypassing protection, unauthorised access or any other unlawful way of obtaining data.
See the OSINT notice, privacy policy and terms of use. How collected documents become entities, events and risks is described in the methodology.