Client Success Story

Powering a Global Market Intelligence Platform with Validated Business Relationship Data

1.2M +

Records Validated

99–100 %

Sustained Accuracy

Platforms

  • Client’s Proprietary Platform
The client

A Leading Global Technology Market Intelligence Platform

The client operates one of the world's most widely used data intelligence platforms for the technology domain. Headquartered in the US, the company applies machine learning and large-scale data aggregation to track private companies, venture capital activity, patents, and emerging market trends. Corporates, investors, and analysts worldwide rely on its datasets to understand industries, spot trends early, and map how organizations connect with one another.

CLIENT OBJECTIVE

Turning Global News into Verified Business Relationship Intelligence

In late 2018, the client approached SunTec India with a specific requirement: identify business relationships between two or more organizations from news articles, press releases, company websites, and corporate updates. These relationships (formal or informal agreements ranging from partnerships and licensing deals to vendor-client arrangements) feed the analytics engine at the core of the client's platform.

The source material was as messy as it was abundant. Given the volume and ambiguity of global news data, the client needed a scalable, context-aware partner who could determine the exact nature of each relationship, verify the companies involved, and remove false or duplicate records.

PROJECT REQUIREMENTS

Data Collection, Validation, and Classification of 1.2 Million+ Records

The client engaged our data collection services, data classification services, and data validation services to run a dedicated business relationship (BR) identification workflow. The scope covered:

  • Record Validation: Verify the integrity and accuracy of business relationship records drawn from global online sources, rejecting inaccurate, ambiguous, or duplicate entries.
  • Relationship Classification: Classify each verified relationship into predefined categories — partnerships (joint ventures, strategic alliances), vendor/client relationships, licensor/licensee agreements, and supplier/distributor agreements.
  • Entity Resolution: Confirm the identity of every company named in a relationship against multiple independent sources.
  • Accuracy Threshold: Maintain 98% accuracy as a minimum first-pass record acceptance threshold across all delivered records.
  • Scalability: Sustain a team capacity that could manage up to 50,000 records per month without compromising delivery speed or quality.
PROJECT CHALLENGES

Extracting Clear Relationships from Ambiguous, High-Volume News Data

Turning unstructured media into precise market intelligence requires overcoming significant data noise, ambiguity, and operational volatility. The primary structural obstacles in this project included:

Ambiguous News Clips

News copy rarely names a relationship in the taxonomy's terms. A "strategic partnership" in a press release may be a routine vendor contract dressed up for publicity. Analysts had to read the full context, assign roles correctly, and reject anything they couldn't substantiate.

Fragmented Company Identities

Global news is full of lookalike names, subsidiaries reported under parent brands, and mid-coverage rebrands. One misattributed entity could poison the relationship graph we were building, so it was crucial to positively identify every company.

High Risk of Data Quality Decay

Every time the client added complex new industries or a new dataset of companies, and changed their taxonomy and data validation rules, staying accurate got harder. On top of that, news websites kept redesigning their pages, which broke automated web scrapers and created huge blind spots where wrong or missing data could easily sneak through.

Absorbing Data Volume Volatility

A news-driven data pipeline regularly fluctuated. The client needed a data validation company with a flexible operational model because they didn't want to pay for idle workers when incoming data volume was low. However, they still expected accurate, clean data and fast outcomes when data volume surged back up.

OUR SOLUTION

Expert-Supervised Data Validation and Entity Classification

The SunTec India team acted as a human-in-the-loop data validation engine for the client’s market intelligence solution. We engineered our data classification services to meet the client's unique requirements and challenges, ensuring high accuracy and deep entity context throughout the project.

1

Structured Data Collection from Global News Sources

Our data collection team gathered and parsed news articles, press releases, and corporate updates from global online sources, organizing them into structured review queues. After web scraping, each analyst processed roughly 20 articles per hour—170 to 175 articles per day—keeping throughput steady against the client's expected monthly data volumes.

2

Manual Data Classification into a Defined Taxonomy

Analysts read each source, identified the organizations involved, attributed roles between them, and classified the relationship into the agreed taxonomy:

  • Partnerships: Identify collaborative agreements, including Joint Ventures (co-created entities or shared-risk ventures) and strategic alliances (co-marketing, co-development, or non-equity technology integration partnerships).
  • Vendor/Client Relationships: Track formal commercial engagements where one organization supplies paid products, services, enterprise software, or operational solutions to another.
  • Licensor/Licensee Agreements: Capture IP-driven arrangements, including patent licensing, software/technology usage rights, brand franchising, and content distribution permissions.
  • Supplier/Distributor Agreements: Map supply chain networks by identifying raw material/component suppliers, original equipment manufacturers (OEMs), authorized wholesalers, and regional distribution partners.

We escalated ambiguous records to QA, an in-house subject matter expert, or, when needed, to the client team. This manual data validation layer, backed by our broader data classification services experience, kept contextual accuracy high.

3

Multi-Source Entity Resolution before Record Approval

We verified every company’s identity against multiple independent sources, such as corporate websites, registries, and prior coverage, to separate subsidiaries from parent companies, catch rebranded company identities, and differentiate between lookalike names. We added a record to the deliverable dataset only after confirming all entities and their roles.

4

Layered Quality Controls to Ensure 98%+ Data Accuracy

The client's contract required 98% accuracy. We treated that as a floor, not a target. Sample-based QC reviews, error root-cause feedback to individual analysts, and periodic calibration sessions against the client's own audits enabled our team to deliver data consistently at 99–100% accuracy rates throughout the project’s runtime.

5

A Team Model that Enabled Quick Scaling (Up and Down)

The engagement began in December 2018 with a small team of 4 members. Impressed by the early deliverable quality, the client expanded the project in February 2019, assigning about 200,000 additional records and growing the team to 20 FTEs, a strength maintained from 2019 to 2021. When COVID-related complications led to reduced team volume, the client went back to 4 members (which we accommodated without impacting quality or delivery consistency). As the times stabilized and their data requirements increased again, the client added 5 more resources in February 2025, bringing the team to 9 FTEs.

EXTENDED PROJECT SCOPE

From Business Relationship Data Analysis to Business Development Data Processing

The strongest proof of a client partnership is that the client keeps trusting you with more of their core business.

Building on nearly eight years of consistent delivery, the client recently transferred a New Business Development (BD) project—previously handled by another vendor—to SunTec India, adding a team of 7 members for the new data processing requirement and increasing total team strength to 16.

PROJECT OUTCOMES

Transforming Data Volatility Into Platform Reliability

Through long-term collaboration, flexible scaling, and rigorous quality controls, the engagement delivered high-value outcomes across the client's platform

1.2 Million+ Records Validated and Classified Scraped, manually reviewed, validated, and categorized, giving the client's corporate customers market intelligence they could act on without second-guessing.

99–100% Data Accuracy Maintained Against a contractual requirement of 98% data accuracy, our team delivered above and beyond, ensuring the client’s solution supported confident decision-making.

Reading a news article and deciding whether two companies genuinely have a business relationship sounds simple. At 37,000 records a month, across every industry, it isn't. Our analysts made that judgment reliably, over a million times. That commitment to quality kept this partnership alive for eight years and counting.

Kuldeep Kumar | DGM-Operations, SunTec India | linkedin-icon

CONTACT US

Build Your Own Reliable Data Pipeline with SunTec India

Whether you need data collection services to feed an analytics engine, data research services to enrich your records, or data classification services to bring order to unstructured sources, we build teams that stay accurate at scale and stay with you for years. Request a free sample on your own data, and experience the quality before you commit.