The client operates one of the world's most widely used data intelligence platforms for the technology domain. Headquartered in the US, the company applies machine learning and large-scale data aggregation to track private companies, venture capital activity, patents, and emerging market trends. Corporates, investors, and analysts worldwide rely on its datasets to understand industries, spot trends early, and map how organizations connect with one another.
In late 2018, the client approached SunTec India with a specific requirement: identify business relationships between two or more organizations from news articles, press releases, company websites, and corporate updates. These relationships (formal or informal agreements ranging from partnerships and licensing deals to vendor-client arrangements) feed the analytics engine at the core of the client's platform.
The source material was as messy as it was abundant. Given the volume and ambiguity of global news data, the client needed a scalable, context-aware partner who could determine the exact nature of each relationship, verify the companies involved, and remove false or duplicate records.
The client engaged our data collection services, data classification services, and data validation services to run a dedicated business relationship (BR) identification workflow. The scope covered:
Turning unstructured media into precise market intelligence requires overcoming significant data noise, ambiguity, and operational volatility. The primary structural obstacles in this project included:
News copy rarely names a relationship in the taxonomy's terms. A "strategic partnership" in a press release may be a routine vendor contract dressed up for publicity. Analysts had to read the full context, assign roles correctly, and reject anything they couldn't substantiate.
Global news is full of lookalike names, subsidiaries reported under parent brands, and mid-coverage rebrands. One misattributed entity could poison the relationship graph we were building, so it was crucial to positively identify every company.
Every time the client added complex new industries or a new dataset of companies, and changed their taxonomy and data validation rules, staying accurate got harder. On top of that, news websites kept redesigning their pages, which broke automated web scrapers and created huge blind spots where wrong or missing data could easily sneak through.
A news-driven data pipeline regularly fluctuated. The client needed a data validation company with a flexible operational model because they didn't want to pay for idle workers when incoming data volume was low. However, they still expected accurate, clean data and fast outcomes when data volume surged back up.
The SunTec India team acted as a human-in-the-loop data validation engine for the client’s market intelligence solution. We engineered our data classification services to meet the client's unique requirements and challenges, ensuring high accuracy and deep entity context throughout the project.
Our data collection team gathered and parsed news articles, press releases, and corporate updates from global online sources, organizing them into structured review queues. After web scraping, each analyst processed roughly 20 articles per hour—170 to 175 articles per day—keeping throughput steady against the client's expected monthly data volumes.
Analysts read each source, identified the organizations involved, attributed roles between them, and classified the relationship into the agreed taxonomy:
We escalated ambiguous records to QA, an in-house subject matter expert, or, when needed, to the client team. This manual data validation layer, backed by our broader data classification services experience, kept contextual accuracy high.
We verified every company’s identity against multiple independent sources, such as corporate websites, registries, and prior coverage, to separate subsidiaries from parent companies, catch rebranded company identities, and differentiate between lookalike names. We added a record to the deliverable dataset only after confirming all entities and their roles.
The client's contract required 98% accuracy. We treated that as a floor, not a target. Sample-based QC reviews, error root-cause feedback to individual analysts, and periodic calibration sessions against the client's own audits enabled our team to deliver data consistently at 99–100% accuracy rates throughout the project’s runtime.
The engagement began in December 2018 with a small team of 4 members. Impressed by the early deliverable quality, the client expanded the project in February 2019, assigning about 200,000 additional records and growing the team to 20 FTEs, a strength maintained from 2019 to 2021. When COVID-related complications led to reduced team volume, the client went back to 4 members (which we accommodated without impacting quality or delivery consistency). As the times stabilized and their data requirements increased again, the client added 5 more resources in February 2025, bringing the team to 9 FTEs.
The strongest proof of a client partnership is that the client keeps trusting you with more of their core business.
Building on nearly eight years of consistent delivery, the client recently transferred a New Business Development (BD) project—previously handled by another vendor—to SunTec India, adding a team of 7 members for the new data processing requirement and increasing total team strength to 16.
Through long-term collaboration, flexible scaling, and rigorous quality controls, the engagement delivered high-value outcomes across the client's platform
1.2 Million+ Records Validated and Classified Scraped, manually reviewed, validated, and categorized, giving the client's corporate customers market intelligence they could act on without second-guessing.
99–100% Data Accuracy Maintained Against a contractual requirement of 98% data accuracy, our team delivered above and beyond, ensuring the client’s solution supported confident decision-making.
Reading a news article and deciding whether two companies genuinely have a business relationship sounds simple. At 37,000 records a month, across every industry, it isn't. Our analysts made that judgment reliably, over a million times. That commitment to quality kept this partnership alive for eight years and counting.
Whether you need data collection services to feed an analytics engine, data research services to enrich your records, or data classification services to bring order to unstructured sources, we build teams that stay accurate at scale and stay with you for years. Request a free sample on your own data, and experience the quality before you commit.