top of page

Product Matching Using TF-IDF and Cosine Similarity: A Beginner’s Guide for E-commerce Data

  • Writer: Anusha P O
    Anusha P O
  • 7 hours ago
  • 27 min read
Product Matching Using TF-IDF and Cosine Similarity: A Beginner’s Guide for E-commerce Data

In the world of e-commerce product matching, thousands of similar items appear under different names, prices, and brands across online shopping platforms. This creates a challenge when trying to match product listings and identify whether two items refer to the same product. For example, the same air conditioner model may appear on Amazon as “LG AI Convertible 1 Ton 5 Star Split Inverter AC.” It may appear on Flipkart as “LG Super Convertible 1 Ton 5 Star Split Dual Inverter AC.” Although both listings describe the same item, their titles are written differently, which makes product title matching for e-commerce harder with basic text comparison.


To solve this, one of the simplest and most powerful methods is product matching using TF-IDF and cosine similarity. Instead of checking every word by hand, cosine similarity measures how closely two product names relate based on their meaning. It looks at the angle between them in multi-dimensional space, not their exact wording. Here, TF-IDF (Term Frequency–Inverse Document Frequency) converts product titles and brand names into numerical form so a machine can understand and compare them. After converting text into vectors, cosine similarity calculates how similar the texts are. It does this by checking the angle between the vectors. This helps find matching products across platforms like Amazon and Flipkart.


This method is the main part of a real  e-commerce  product matching system and is also widely used in E-commerce Text Classification. It helps remove duplicate products. It also improves product discovery. It supports competitor analysis and price comparison. It is very useful to compare product prices, discounts, and reviews automatically. This helps users make better choices when picking items. Product matching also helps compare prices and discounts. It is good for choosing cheap and high-quality branded products.


The following sections show a simple and beginner-friendly way to build automatic product matching for e-commerce and online shopping. We use scraped product data like air conditioners and coolers from Amazon and Flipkart online shopping site. The goal is to find duplicate product listings. We match similar products across websites. We store the matched results in a database for further analysis. Each step works like a puzzle — starting from cleaning text, converting it into numbers (vectorization), applying similarity scores to the final matched output.


What Is Cosine Similarity? A Beginner's Guide to Product Matching for E-commerce


What Is Cosine Similarity? A Beginner's Guide to Product Matching for E-commerce

At the heart of any product-matching system lies the ability to measure how similar two pieces of text are. Cosine Similarity is a mathematical way to find how close two text strings are. It does not compare words directly. Instead, it looks at their direction in a multi-dimensional space. When product titles are represented as numerical vectors, cosine similarity checks the angle between those vectors. If the angle is small, it means both titles point in nearly the same direction and are therefore very similar. If the angle is large, the texts differ significantly.


The formula behind cosine similarity is straightforward and elegant. It measures the cosine of the angle between two vectors, expressed as:

             

                                                          (A · B) 

      Cosine Similarity =           ___________________

                                   

                                                      (‖A‖ × ‖B‖)


Here, A and B represent two text vectors. The numerator, A · B, calculates how much the two vectors overlap — in other words, how similar the words or features are between the two texts. The denominator, ‖A‖ × ‖B‖, normalizes these values based on the length (or magnitude) of each vector. This ensures that the comparison remains fair, even if one product title is longer than the other.


The final cosine similarity score is always between 0 and 1. Scores near 1 mean the titles are very similar. Scores near 0 mean the texts are quite different.


To understand this intuitively, imagine two arrows drawn from the center of a circle. If both arrows point in the same direction, the angle between them is zero, and the cosine value becomes one—indicating perfect similarity. If they point in completely opposite directions, the cosine value becomes zero, meaning there is no similarity at all.


What makes cosine similarity particularly effective for text is that it focuses on orientation rather than magnitude. In other words, the number of words in a title does not matter as much as the type and order of words used. This is especially useful when comparing product titles that may have different lengths but convey the same meaning. For instance, Samsung 1.5 Ton Split AC” and “Samsung Split Air Conditioner 1.5 Ton” are written differently but represent the same product. Cosine similarity measures this relationship accurately.


Why We Use TF-IDF to Compare Product Titles


Why We Use TF-IDF to Compare Product Titles

Every text-based matching system begins with a fundamental challenge: computers cannot understand words the way humans do. For a machine, the words “Samsung Split AC” and “Split Air Conditioner by Samsung” (product titles) are simply strings of characters. To compare such text meaningfully, these strings must first be transformed into a numerical form that preserves their meaning and context. This is where TF-IDF Vectorization plays an important role.


TF-IDF stands for Term Frequency–Inverse Document Frequency. It is a statistical method that changes text into numbers based on how important each word is in a group of documents. The concept is intuitive. Words that appear often in a product title get higher weight through term frequency. Very common words that appear in almost every title—like “with,” “for,” or “and”—get lower weight through inverse document frequency. The result is a balanced representation where each word’s importance reflects both its presence and its uniqueness.


When applied to product titles, TF-IDF helps highlight the distinctive terms that actually describe a product, such as its brand, model, or key features. For instance, in the product title LG 1 Ton Inverter Split AC”, words like “LG,” “Inverter,” and “Split” carry strong informational value, while filler words are considered less significant. The TF-IDF process turns the text into a vector. This vector is a list of numbers that represents the product title in a form that can be compared mathematically.


Once the product titles from both datasets — Amazon and Flipkart—are transformed into TF-IDF vectors, Cosine Similarity comes into action. Rather than checking whether two titles contain the exact same words, cosine similarity evaluates the angle between their corresponding TF-IDF vectors. A smaller angle, or a higher cosine value, indicates that the titles share similar patterns and vocabulary, even if the wording differs.


TF-IDF and cosine similarity form a strong pair. TF-IDF changes text into numbers. Cosine similarity measures how close those numbers are. Together, they enable efficient and accurate matching of product data across large e-commerce catalogs. The process captures not just exact word overlap. It also captures how close product descriptions are in context. This helps with tasks like E-commerce Text Classification and accurate product matching. This allows the system to find true matches hidden behind different phrasing.


In product-matching workflows, this approach ensures that models focus on the essence of a product rather than superficial text differences. It makes the system easy to grow, understand, and change for different product types. This works whether you compare air conditioners, smartphones, or kitchen appliances.  TF-IDF uses statistical weighting. Cosine similarity measures how similar two things are using geometry. They work together.Together, they form the main analysis method in modern text-based matching systems.


Real-World Product Matching Example: Matching Godrej AC Listings


Consider two product entries taken from Amazon and Flipkart. Both describe the same Godrej air conditioner model but the titles use slightly different wording and extra phrases. The cleaned titles used for comparison are shown here:


  • Amazon cleaned title: godrej 1 ton 5 star 5 in 1 convertible cooling inverter split ac copper i sense technology 2023 model ac 1t ei 12tinv5r32 gwa split ac 1t ei 12tinv5r32 rwb split white



  • Flipkart cleaned title: godrej 5 in 1 convertible cooling 2023 model 1 ton 5 star split inverter i sense technology with blue fin anti corrosive coating ac white gold ac 1t ei 12tinv5r32 gwa split ac 1t ei 12tinv5r32 rwb split copper condenser


The calculation follows the usual steps. First, it breaks text into tokens and creates TF-IDF vectors. Then, it calculates the dot product and Euclidean norms. Finally, it uses the cosine formula. Scikit-learn’s default settings (including smooth_idf=True and norm='l2') are used so vectors are L2-normalized.


Step 1 — TF-IDF vectors and vocabulary


The TF-IDF vectorizer builds a vocabulary of 28 tokens from the two titles. Each title is converted into a TF-IDF vector; because norm='l2', both vectors have Euclidean norm 1.0. A selection of the TF-IDF values and their per-token contributions to the dot product follows (values shown are exact as produced by scikit-learn):

Token

Amazon

Flipkart

Contribution

12tinv5r32

0.29814239699997197

0.25648898295269357

0.0764702401816010

1t

0.29814239699997197

0.25648898295269357

0.0764702401816010

ac

0.4472135954999579

0.3847334744290404

0.1720580404086023

split

0.4472135954999579

 0.3847334744290404

 0.1720580404086023

2023

0.14907119849998599

0.12824449147634678

0.0191175600454003

gwa

0.14907119849998599

0.12824449147634678

0.0191175600454003

inverter

0.14907119849998599

0.12824449147634678

0.0191175600454003

model

0.14907119849998599

0.12824449147634678

0.0191175600454003

ton

0.14907119849998599

0.12824449147634678

 0.0191175600454003

white

0.14907119849998599

0.12824449147634678

0.0191175600454003

  • Several additional shared tokens (e.g., convertible, cooling, sense, split, star, rwb, etc.) each contribute ≈ 0.01911756 or 0.0 if absent in one title (Flipkart contains some extra tokens such as blue, anti, coating, condenser, gold, with, each present only in Flipkart and thus contributing 0 to the dot product


Step 2 — Dot product (numerator of cosine formula)


The dot product sums all per-token contributions. Using the TF-IDF numbers above, the dot product equals:


A⋅B  =  0.8602902020430111


Step 3 — Vector norms (denominator parts) 


Because scikit-learn normalized TF-IDF vectors with L2 norm, each vector has:


      ∥A∥=1.0

      ∥B∥=1.0


Step 4 — Cosine similarity (title) 


Apply the cosine formula:

   

  Cosine_title

                     = A⋅B/∥A∥×∥B∥

                     = 0.8602902020430111/(1.0×1.0)

                     = 0.8602902020430111


Rounded for presentation: 0.86029 (≈ 0.86).


This value shows a strong similarity between the two titles after applying TF-IDF (term frequency-inverse document frequency)  weighting and normalization.


Step 5 — Brand similarity


Both product entries use the same cleaned brand token godrej. Using TF-IDF on the brand strings (a trivial two-document corpus where both items are identical) yields identical normalized vectors, so:

cosine_brand=1.0


Step 6 — Combined similarity with weights


The script combines title and brand similarity using the configured weights title_weight = 0.8 and brand_weight = 0.2


The weighted combination is computed as:


Combined_similarity

             =0.8×cosine_title+0.2×cosine_brand 


"Combined_similarity equals 0.8 times cosine_title plus 0.2 times cosine_brand"


Substituting the numbers:


Combined_similarity

=0.8×0.8602902020430111+0.2×1.0

=0.6882321616344089+0.2

= 0.8882321616344089


Rounded for reporting: 0.888 (≈ 0.89).


Interpretation and practical note


The title cosine (≈ 0.86) shows a very close textual match between the two cleaned product titles. The brand cosine (1.0) confirms brand agreement. After applying the chosen weights, the final combined similarity is about 0.888. This number is higher than a typical matching threshold like 0.7. So, the pair is a strong candidate for being the same product.


This example shows how small wording differences and extra descriptive phrases do not stop a successful match. This happens when TF–IDF vectorization and cosine similarity are used. The same approach scales to large catalogs: text is cleaned, converted to TF–IDF vectors, and compared with cosine similarity to identify likely matches. Later improvements can make the results more accurate. These include adding brand exact-match checks, giving brand tokens more weight, or using transformer embedding for better meaning capture.


Step-by-Step Python Tutorial for Product Matching Between Amazon and Flipkart


Essential Python Libraries Used 


import pandas as pd
import sqlite3
import re
import logging
from datetime import datetime
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
from tqdm import tqdm
import os
import sys

The journey starts with Pandas library, a powerful data-handling library that simplifies reading, cleaning, and transforming datasets. Here, it manages large CSV files containing product information from Amazon and Flipkart. SQLite3 follows, acting as a lightweight database that helps store matched results efficiently for future use or reporting.


Before any meaningful comparison can happen, product text needs to be cleaned and standardized. The regular expression module, re, helps clean text. It removes extra characters, symbols, or spaces. This makes product titles uniform and easier to compare.


Accurate tracking of the script’s progress is equally important, especially when handling large datasets. Logging is included to record the entire process, from start to finish, along with any errors or matches found. To ensure proper time-based tracking, the datetime module helps timestamp events throughout the execution.

Scikit-learn offers two strong tools to find product similarity. They are TfidfVectorizer and Cosine_similarity. The TF-IDF (term frequency-inverse document frequency) vectorizer changes text into number vectors based on word importance. Cosine similarity measures how close two vectors are, so it shows how similar two product titles are. Together, they form the core of the product matching logic.


Since product comparison can be time-consuming, especially with thousands of rows, tqdm introduces progress bars that make it easier to monitor the script’s progress visually. We include os and sys for system tasks. These include file handling, managing paths, and handling unexpected interruptions smoothly.

Together, these imports build a strong base for the product matching process. They combine data handling, text cleaning, similarity calculation, and result management into a clear and efficient system.


Setting Up Configuration Paths

# CONFIG

amazon_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/amazon-ac-v8-cleaned.csv"
flipkart_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/flipkart-AC-full-v1-cleaned.csv"
db_path = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/matched_products_final.db"
log_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/product_match_log_final.txt"

"""
Configuration setup for the product matching workflow.
"""

In this stage of the product-matching pipeline, configuration paths are defined for the main files and databases that form the core of the workflow. Setting up these configurations at the start makes the code run smoothly. It also prevents confusion later. This setup makes the code easier to manage and change when needed.

The first two variables, amazon_file and flipkart_file, point to the cleaned datasets collected from Amazon and Flipkart. These CSV files contain structured product data such as names, specifications, and other details that will later be compared to find similar or matching products. The paths ensure that the code knows exactly where to locate these datasets on the system. Keeping the data in a cleaned and organized format at this stage is essential because it directly affects the quality and accuracy of the matching results.


The db_path variable defines the location of the SQLite database where all final matched product results will be stored. Using a database for storing results rather than saving them as flat files provides more control and flexibility. It allows for efficient querying, filtering, and retrieving specific matches without re-running the entire process. This becomes particularly valuable when working with large volumes of e-commerce data.


Lastly, the log_file variable specifies the location where all logs will be recorded. Logging plays a vital role in maintaining transparency throughout the process. Each event—whether it is a successful match, an error, or a step in progress—is automatically documented in this log file. This record helps in tracking the execution flow and debugging any unexpected issues.

Together, these configuration settings form the backbone of the project. They provide structure and clarity. This makes sure that every later step works well and consistently. These steps include data cleaning, text comparison, and similarity calculation. It also removes the need for manual changes each time the code runs.


Setting the Right Threshold and Weights for Accurate Product Comparison

# cosine similarity threshold and weights
threshold = 0.7
title_weight = 0.8
brand_weight = 0.2

"""
Defines the cosine similarity threshold and attribute weights for product matching.
"""

After defining the core configuration paths, the next step in the product matching workflow involves setting up the similarity parameters. These parameters determine how closely two products need to resemble each other to be considered a match. Cosine similarity measures how similar product titles and product descriptions are. We assign weights to specific attributes. These weights control how much each attribute affects the final matching score.


The threshold value, set at 0.7, acts as a decision boundary. It represents the minimum level of similarity required between two product descriptions for them to qualify as a potential match. For instance, if the cosine similarity between two titles is 0.85, the system recognizes them as highly similar; however, if the score falls below 0.7, the match is discarded. Choosing the right threshold is crucial—it must be balanced enough to avoid both false positives (mismatches) and false negatives (missed matches).


The title_weight and brand_weight variables help fine-tune the importance of different product attributes. In this setup, the product title carries a weight of 0.8, while the brand description contributes 0.2 to the overall similarity score. This weighting reflects a realistic approach to e-commerce product matching, where titles often provide more detailed and descriptive information than brand names. For example, two air conditioners from the same brand can have very different specifications. But their titles mention capacity, model, or features. These product titles give a better way to compare them.

By defining these parameters early in the process, the matching system is guided to make more accurate, data-driven decisions. The process keeps the comparison consistent and objective. It matches real-world product variations. This leads to reliable and meaningful matching results across large datasets.


Creating a Reliable Logging System for Transparent Data Workflows

# LOGGING SETUP

if os.path.exists(log_file):
   os.remove(log_file)

logging.basicConfig(
   level=logging.INFO,
   format="%(asctime)s - %(levelname)s - %(message)s",
   handlers=[logging.FileHandler(log_file), logging.StreamHandler(sys.stdout)]
)

def log(msg, level="info"):
   getattr(logging, level)(msg)

log(" Product Matching Script Started")

"""
Initializes and configures the logging system for the product matching script.
"""

Once the configurations and similarity parameters are set, the next essential component is the logging setup. Logging acts as the project’s internal diary—it records every significant event, tracks progress, and helps in understanding the system’s behavior at each stage. In large data processing tasks, especially when comparing thousands of products, keeping clear and organized logs is essential.


The first step ensures that any existing log file is removed before creating a new one. This keeps the log clean for every fresh run, avoiding clutter from previous executions. The logging.basicConfig() function then initializes the logging configuration with a few key parameters. The logging level is set to INFO, meaning the system will record all informative messages, warnings, and errors. The format specifies how each message should appear in the log, including the timestamp, log level, and the actual message text. The handlers send the output to two places at once. One goes to the log file. The other goes to the console. This lets you track in real time and keep a written record.


To simplify message logging throughout the script, a helper function named log() is defined. This function dynamically calls the appropriate logging method, such as info, warning, or error, depending on the type of message passed to it. The first entry in the log, “ Product Matching Script Started,” marks the beginning of the process. It signals that the system has initialized successfully and is ready to begin the product matching workflow.


This setup ensures that every step, from reading data to calculating similarity scores, is carefully documented. In case of unexpected errors or inconsistencies, these logs serve as a reliable trail for identifying the root cause. This organized approach improves transparency. It also helps with efficient debugging and process improvement in large data projects.


Cleaning and Preparing Text Data for Accurate Comparison


# CLEANING FUNCTION

def clean_text(text):
   if not isinstance(text, str):
       return ""
   text = text.lower()
   text = re.sub(r"[^a-z0-9\s]", " ", text)
   text = re.sub(r"\s+", " ", text).strip()
   return text

"""
Cleans and standardizes raw product text for consistent comparison.

"""

In any text-based data processing task, one of the most crucial steps before analysis begins is data cleaning. Raw product data from different e-commerce platforms often has problems. It includes extra symbols, special characters, or unwanted spaces. These problems can make comparisons inaccurate. To ensure reliable results, all text must be standardized and simplified into a clean, comparable format. This is where the cleaning function plays a central role.


The function named clean_text() is designed to perform this essential preprocessing task. It begins by checking whether the input is a string. This validation step prevents errors when unexpected data types—such as numbers or missing values—appear in the dataset. If the value isn’t a string, the function simply returns an empty string, ensuring the process continues smoothly without interruptions.


Once the input is confirmed as text, the function converts all characters to lowercase. This helps maintain uniformity, so words like “Samsung,” “samsung,” and “SAMSUNG” are treated as identical during similarity calculations. The next step uses regular expressions to remove unwanted characters—anything that isn’t a letter, number, or space. This eliminates punctuation marks and symbols that do not contribute meaningfully to product matching.


After that, multiple spaces are replaced with a single space, and any leading or trailing spaces are removed. The end result is a clean, well-formatted string that focuses only on the meaningful content. For example, a messy product titles like “LG 1.5-Ton Inverter A/C!!!” becomes “lg 1.5 ton inverter a c” after cleaning.


By applying this cleaning process consistently to all product titles and descriptions, the matching algorithm gains a clear and unbiased view of the data. It ensures that every comparison between two product titles is based purely on their relevant content, not on noise or formatting differences. This step lays the foundation for accurate similarity computation and is one of the most important parts of preparing real-world data for analysis.


Measuring Similarity with Jaccard Distance

# JACCARD DISTANCE FUNCTION

def jaccard_distance(str1, str2):
   set1 = set(str1.split())
   set2 = set(str2.split())
   if not set1 and not set2:
       return 0.0
   intersection = set1.intersection(set2)
   union = set1.union(set2)
   return 1 - len(intersection) / len(union)

	"""
	Calculate the Jaccard distance between two strings.

	"""
Jaccard similarity

In data analysis and product matching, comparing how similar two pieces of text are is a common challenge. One practical method for this is the Jaccard distance, which quantifies the difference between two sets of words. At its core, the Jaccard distance focuses on the overlap between the words in two strings, providing a simple yet powerful way to capture similarity.


The function begins by splitting each string into a set of words. This conversion from text to sets is crucial because sets automatically remove duplicate words, which ensures that the comparison considers only unique terms. Once the sets are created, the intersection of the two sets is calculated. This intersection represents the words that appear in both strings, highlighting the shared elements between the two texts. In contrast, the union of the sets includes all unique words present in either string, capturing the total scope of vocabulary used.


The Jaccard distance itself is computed as one minus the ratio of the intersection size to the union size. If two strings are exactly the same, their intersection equals the union, resulting in a distance of zero, which indicates perfect similarity. Conversely, if there are no shared words, the intersection is empty, and the distance reaches one, signaling complete dissimilarity. This simple calculation makes the Jaccard similarity easy to understand. It is useful for tasks like product matching, document comparison, or grouping similar items.


This method works well when the focus is on whether terms are present or not. It gives a different view than methods like TF-IDF or cosine similarity. Using Jaccard distance with other similarity measures helps us better understand text relationships. This ensures accurate and meaningful matches between datasets.

The function itself is concise and efficient, handling edge cases where both strings might be empty by returning a distance of zero. This ensures robustness in real-world scenarios, where missing or incomplete data is often encountered. Jaccard similarity gives a reliable and clear way to compare text. It helps data systems find connections and patterns clearly.


Bringing Amazon and Flipkart Data into the Workflow

# LOAD DATA

try:
   amazon_df = pd.read_csv(amazon_file)
   flipkart_df = pd.read_csv(flipkart_file)
   log(f" Loaded Amazon: {len(amazon_df)} rows, Flipkart: {len(flipkart_df)} rows")
except Exception as e:
   log(f" Error loading CSVs: {e}", "error")
   sys.exit(1)

"""
Loads cleaned product datasets from Amazon and Flipkart using pandas."""

After defining the cleaning function, the next step in the workflow focuses on loading the datasets that will be used for product comparison. This stage acts as the entry point for the actual data that drives the entire matching process. Clean, structured data from multiple sources must be read into the program accurately before any analysis can begin. Even a small error at this point can affect every step that follows, making reliable data loading one of the most important foundations of the pipeline.

In this section, the script attempts to read two CSV files—one from Amazon and the other from Flipkart—using the pandas.read_csv() function. The try-except structure is used to handle this process safely. Inside the try block, the code loads both datasets into memory as DataFrames. DataFrames have a table-like format that makes it easy to work with and analyze the data. Once the data is successfully loaded, a log message records the number of rows in each dataset. This immediate feedback confirms that the files were read correctly and helps verify that the expected amount of data has been imported.


The inclusion of an exception handling block ensures the system remains stable, even if something goes wrong. For instance, if one of the file paths is incorrect or a file is missing, the program logs a clear error message and exits gracefully instead of crashing. This approach adds robustness to the script, making it capable of handling real-world scenarios where data inconsistencies or missing files are common.


By the end of this step, the raw datasets from both e-commerce platforms are securely loaded and ready for the next stages of processing. This careful and organized loading process makes sure the matching system starts with reliable and accessible data. This prepares the system for accurate comparison and analysis in the next steps.


From Raw to Ready: Cleaning and Validating Product Data

# CLEAN & VALIDATE

amazon_df.columns = amazon_df.columns.str.lower()
flipkart_df.columns = flipkart_df.columns.str.lower()

required_cols = ["product_title", "brand"]
for col in required_cols:
   if col not in amazon_df.columns or col not in flipkart_df.columns:
       log(f" Missing required column '{col}'", "error")
       sys.exit(1)

log("🧹 Cleaning text columns...")
amazon_df["clean_title"] = amazon_df["product_title"].apply(clean_text)
amazon_df["clean_brand"] = amazon_df["brand"].apply(clean_text)
flipkart_df["clean_title"] = flipkart_df["product_title"].apply(clean_text)
flipkart_df["clean_brand"] = flipkart_df["brand"].apply(clean_text)


"""
Performs initial data validation and text cleaning on the loaded Amazon and Flipkart datasets.
"""

Once the data has been successfully loaded into memory, the next essential phase is data validation and cleaning. This step makes sure that the datasets from Amazon and Flipkart have the same structure. It also ensures they have the needed information for accurate product matching. Before calculating similarity, we must standardize and check the data. Small differences in column names or missing fields can cause problems.


The first step converts all column names to lowercase. This simple change removes differences caused by naming styles. For example, one dataset might use “Product_Title” while another uses “product_title.” By standardizing column names, the script avoids confusion and ensures smooth access to each field during processing.


Next, the code checks for the presence of key columns—product_title and brand—in both datasets. These two attributes are fundamental to the matching process. The product title provides descriptive details about the item, while the brand description helps confirm its identity. If either of these columns is missing from any dataset, the script immediately logs an error and halts execution. This safeguard prevents incomplete data from entering later stages, where it could lead to inaccurate results or unexpected failures.


Once validation is complete, the focus shifts to text cleaning. Each product title and brand name is passed through the clean_text() function defined earlier. This function removes unwanted characters, converts text to lowercase, and ensures a consistent structure across both datasets. The cleaned versions of these fields are stored in new columns—clean_title and clean_brand—which serve as the standardized input for all future comparisons.


Through this process, both datasets are transformed into a clean, uniform, and validated state. Each entry is now prepared for reliable analysis, free from inconsistencies that could distort similarity scores. Careful attention to data quality is the foundation of any successful product matching process. It makes sure every comparison later is based on correct and meaningful information.


Converting Text Data into Vectors Using TF-IDF


# TF-IDF FITTING

log("📊 Fitting TF-IDF on combined corpus...")
vectorizer = TfidfVectorizer().fit(
   amazon_df["clean_title"].tolist() + flipkart_df["clean_title"].tolist()
)
log(" TF-IDF ready")

After cleaning and checking the datasets, the next important step changes the text data into a form the computer can understand and compare well. Human readers can easily see that two titles like “LG 1.5 Ton 5 Star Inverter AC” and “LG Inverter Split AC 1.5 Ton 5 Star” mean the same product. But a machine only sees both as strings of text. To enable meaningful comparisons, this textual information needs to be converted into numerical representations — and that’s where TF-IDF comes into play.

TF-IDF, short for Term Frequency–Inverse Document Frequency, is a method that converts text into a set of numerical features based on how important each word is within a collection of texts. In simpler terms, it identifies which words carry the most value in describing a product. Common words like “air” or “ac” might appear frequently across many titles, so their importance is lower. However, more specific terms such as “inverter” or “dual cool” appear less often and thus carry greater weight when measuring similarity between two product names.


In this section, the code creates a TF-IDF (term frequency-inverse document frequency) vectorizer and fits it on a combined set of cleaned product titles from both Amazon and Flipkart. By joining text data from both platforms into one group, the model learns one vocabulary. This makes sure the same words are shown the same way in both datasets. This shared representation becomes the foundation for calculating similarity scores later in the process.


Once the fitting is complete, the log confirms that the TF-IDF model is ready. At this stage, every product title is prepared to be converted into its corresponding vector — a numerical form that captures the essence of the text. This change connects human language to machine calculation. Then, cosine similarity measures how closely two products relate based on their vector forms.


Product Matching Using TF-IDF (term frequency-inverse document frequency) and Cosine Similarity


# MATCHING LOOP

matches = []
log("🔍 Starting product matching...")

for i, a_row in tqdm(amazon_df.iterrows(), total=len(amazon_df), desc="Matching"):
   try:
       # TF-IDF vectors for titles
       a_title_vec = vectorizer.transform([a_row["clean_title"]])
       f_title_vecs = vectorizer.transform(flipkart_df["clean_title"])
       title_sim = cosine_similarity(a_title_vec, f_title_vecs).flatten()

       # TF-IDF vectors for brands
       brand_vectorizer = TfidfVectorizer().fit(
           [a_row["clean_brand"]] + flipkart_df["clean_brand"].tolist()
       )
       a_brand_vec = brand_vectorizer.transform([a_row["clean_brand"]])
       f_brand_vecs = brand_vectorizer.transform(flipkart_df["clean_brand"])
       brand_sim = cosine_similarity(a_brand_vec, f_brand_vecs).flatten()

       # Combined cosine similarity
       combined_sim = (title_sim * title_weight) + (brand_sim * brand_weight)
       best_idx = combined_sim.argmax()
       best_score = combined_sim[best_idx]

       if best_score >= threshold:
           # Jaccard distance for the matched pair
           combined_text_a = f"{a_row['clean_title']} {a_row['clean_brand']}"
           combined_text_f = f"{flipkart_df.iloc[best_idx]['clean_title']} {flipkart_df.iloc[best_idx]['clean_brand']}"
           jdl_score = round(jaccard_distance(combined_text_a, combined_text_f), 3)

           matches.append({
               "amazon_index": i,
               "flipkart_index": best_idx,
               "cosine_similarity": round(best_score, 3),
               "jaccard_distance": jdl_score
           })

           log(f" Match (Cosine {best_score:.2f} | JDL {jdl_score:.2f}): "
               f"{a_row['product_title'][:60]} ↔ {flipkart_df.iloc[best_idx]['product_title'][:60]}")

   except Exception as e:
       log(f" Error on Amazon row {i}: {e}", "warning")


"""
This loop performs product matching between the Amazon and Flipkart datasets by comparing
both the product titles and brands using TF-IDF vectorization and similarity metrics. 
"""

When comparing products from different platforms, the main challenge is to find which items match between datasets. This is achieved by evaluating the similarity between product titles and brands, which are often the most descriptive identifiers. The matching loop does this task efficiently. It uses math methods to measure similarity.


Each product from the Amazon dataset is processed one by one. The product title is first changed into numbers using a method called TF-IDF. TF-IDF stands for Term Frequency-Inverse Document Frequency. This technique converts text into vectors that capture the importance of each word relative to the entire dataset. A similar transformation is applied to all product titles from Flipkart, allowing a direct comparison between the two platforms. We use cosine similarity to calculate how similar two vectors are. This measures how closely two vectors point in the same direction. A high cosine similarity indicates that the product titles are closely related, which is crucial for accurate matching.


Brands are treated in a similar manner. A separate TF-IDF vectorizer is fitted specifically for the brand names and brand descriptions of the current Amazon product and all Flipkart products. Once again, cosine similarity is used to evaluate how closely the brand names match. By combining the similarities from both the title and brand, a weighted score is computed. The weighting allows the title to have a stronger influence while still considering the contribution of the brand. The highest combined similarity score identifies the most likely matching product on Flipkart for the current Amazon product.


In addition to cosine similarity, Jaccard distance is calculated for the matched pair. This metric provides a different perspective by evaluating the overlap of unique words between the combined title and brand texts of the two products. Cosine similarity uses vector space. Jaccard distance measures how much vocabulary is shared. Jaccard distance helps find partial matches or small text differences. The computed Jaccard distance is rounded and stored alongside the cosine similarity, providing a dual-metric assessment of how closely the products align.


Each successful match is logged, showing both the cosine similarity and Jaccard distance alongside a snippet of the product titles. This not only ensures transparency in the matching process but also helps in debugging and verifying the results. The script catches and records any errors during processing, such as missing data or unexpected text formats. It does this without stopping the overall execution. This makes the script strong for large datasets.


This loop forms the core of the product matching system, combining advanced text representation techniques with practical similarity measures. The system compares every product and records both cosine similarity and Jaccard distance. This method helps match products across platforms reliably. It also prepares the data for further analysis or use in e-commerce workflows.


Storing Amazon–Flipkart Matched Results Safely Using SQLite


# DATABASE SAVE

log(" Saving results to database...")
conn = sqlite3.connect(db_path)

if matches:
   matches_df = pd.DataFrame(matches)
   amazon_matched = []
   flipkart_matched = []
   combined_rows = []

   all_columns = list(set(amazon_df.columns.tolist() + flipkart_df.columns.tolist()))
   key_columns = [
       "product_url", "product_title", "company", "brand", "mrp", "sales_price", "similarity", "jaccard_distance"
   ]
   ordered_cols = key_columns + [c for c in all_columns if c not in key_columns]

   for _, match in matches_df.iterrows():
       a_row = amazon_df.loc[match["amazon_index"]].to_dict()
       f_row = flipkart_df.loc[match["flipkart_index"]].to_dict()

       a_row["similarity"] = match["cosine_similarity"]
       a_row["jaccard_distance"] = match["jaccard_distance"]
       a_row["company"] = "Amazon"

       f_row["similarity"] = match["cosine_similarity"]
       f_row["jaccard_distance"] = match["jaccard_distance"]
       f_row["company"] = "Flipkart"

       # Fill missing columns with None
       for col in ordered_cols:
           a_row.setdefault(col, None)
           f_row.setdefault(col, None)

       amazon_matched.append(a_row)
       flipkart_matched.append(f_row)
       combined_rows.extend([a_row, f_row])

"""
DATABASE SAVE SECTION

This section saves the matched product data between Amazon and Flipkart into an SQLite database.
"""

When handling product matching between two major e-commerce marketplace, the final step is often saving the results in a structured and accessible way. After we calculate similarity measures like cosine similarity and Jaccard distance, we organize the data into clear tables. This is important for future analysis or reports. The process begins by establishing a connection to a SQLite database, which serves as a lightweight and reliable storage solution for structured data.


Once the database connection is ready, the script checks if there are any matched products. Each match contains references to the corresponding entries in both the Amazon and Flipkart datasets, along with the computed similarity scores. The system goes through these matches one by one. It changes the data into dictionaries. This keeps the similarity and Jaccard distance scores with the original product details. Assigning a company label to each entry clearly distinguishes between the two platforms.


To maintain consistency and prevent missing values, the script ensures that every column in the final tables is filled. Any missing field is explicitly set to None, guaranteeing that the database structure remains uniform across all entries. We collect matched products from Amazon and Flipkart separately. We also create a combined dataset for full analysis. This method allows flexibility. It lets you see platform-specific matches separately. It also gives a combined view of all product comparisons.


The final stage involves writing these datasets to the database. Separate tables are created for Amazon matches, Flipkart matches, and a combined table containing all matched products. We keep a special table for Jaccard distance. This table does not include the cosine similarity column. It gives another way to look at product similarity. This organized method stores the results efficiently. It also makes it easy to search and analyze the data for insights, reports, or further work.


Using this method, we store matched product data in a professional and easy-to-access format. This keeps all important details and lets users explore and understand product relationships across platforms in many ways. This setup ensures that any analysis performed on the matched data is accurate, reproducible, and easy to manage.


Creating Separate Database Tables for Amazon, Flipkart, and Combined Matches


# Convert to DataFrames
   amazon_df_out = pd.DataFrame(amazon_matched)[ordered_cols]
   flipkart_df_out = pd.DataFrame(flipkart_matched)[ordered_cols]
   combined_df_out = pd.DataFrame(combined_rows)[ordered_cols]

   # Save tables
   amazon_df_out.to_sql("amazon_matched_only", conn, if_exists="replace", index=False)
   flipkart_df_out.to_sql("flipkart_matched_only", conn, if_exists="replace", index=False)
   combined_df_out.to_sql("matched_combined", conn, if_exists="replace", index=False)

   # Save Jaccard distance table (exclude cosine similarity)
   jdl_df_out = combined_df_out.drop(columns=["similarity"])
   jdl_df_out.to_sql("matched_jaccard", conn, if_exists="replace", index=False)

   log(f" amazon_matched_only: {len(amazon_df_out)} rows")
   log(f" flipkart_matched_only: {len(flipkart_df_out)} rows")
   log(f" matched_combined: {len(combined_df_out)} rows")
   log(f" matched_jaccard: {len(jdl_df_out)} rows")

else:
   log(" No matches found above threshold")

"""
DATAFRAME CONVERSION AND DATABASE SAVE

This section converts the matched product data stored in Python lists into Pandas DataFrames
and saves them into an SQLite database. 
"""

Once the product matching process is complete, the results are organized and stored in a structured format for future analysis. This step ensures that all matched product information between Amazon and Flipkart is captured in a way that is easy to query, compare, and visualize. The first action involves converting the collected match lists into Pandas DataFrames. Separate DataFrames are created for Amazon products, Flipkart products, and a combined view that brings together all matched rows. This separation lets us analyze each platform alone. It also lets us see the overall matching in one dataset.


Saving these DataFrames to an SQLite database provides a reliable, persistent storage mechanism.


  • amazon_matched_only – Stores all matched products from Amazon.

  • flipkart_matched_only – Stores all matched products from Flipkart.

  • matched_combined – Contains all matched rows from both platforms.


This structured approach ensures that queries can be tailored based on specific requirements, such as platform-specific insights or cross-platform comparisons.

In addition to cosine similarity, the Jaccard distance between products is also stored. To focus on this metric, we create a separate table named matched_jaccard. We make it by removing the cosine similarity column from the combined DataFrame.


This separation shows the Jaccard distance clearly. It does not mix with other similarity measures. This makes it easier to understand and use for further analysis.

Finally, the script logs the number of rows in each table, providing a quick overview of how many matches were successfully identified and stored. These logs serve both as a checkpoint and a record for auditing purposes. By organizing and storing the matched product data in this systematic way, the dataset becomes a powerful resource for insights, reporting, or downstream processing tasks.


Final Step: Safely Closing SQLite Connection and Ending the Product Matching Process


conn.close()
log(" Script finished successfully")
log(f" Completed at {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")

"""
Closes the active database connection and logs the completion of the script.
"""

After the matched data has been successfully stored in the database, the final step gracefully concludes the entire process. The database connection, which has been open throughout the matching and saving operations, is now closed to ensure that all resources are properly released. This is an important best practice in any data-handling workflow, as keeping connections open unnecessarily can lead to issues such as memory leaks or database locks. The script closes the connection explicitly. This keeps the process efficient and stable. It shows that all operations finished safely.


Once the connection is closed, a series of final log messages marks the completion of the entire workflow. These messages not only confirm that the script has run successfully but also record the exact time of completion. This acts as a clear sign for tracking how well the process runs or for fixing problems. This is especially useful when working with large datasets.


This closing stage gives the entire process a sense of completion and reliability. Loading and preparing data, calculating similarities, finding matches, and saving results all help build a smooth, automated product matching system. Ending the script with clean closure and clear logs helps the pipeline run well. This makes the process professional, easy to follow, and reliable. These are important parts of good data engineering.


Input Data


Before starting the product matching process, we used two cleaned datasets as inputs:



Output Results

These files contain the matched product pairs along with their computed cosine similarity scores and related details:


Table Name

What it Stores

Includes the Amazon product details that were successfully matched with Flipkart listings based on the cosine similarity threshold

Contains only the Flipkart details of the matched items, aligned with their corresponding Amazon listings

This file consolidates all matched records from both Amazon and Flipkart, making it easier to analyze the overall matching accuracy

Contains all matched product records from both Amazon and Flipkart, it focuses on the Jaccard distance metric. This table is useful for analyzing matches using set-based similarity rather than vector-based similarity



AUTHOR


I’m Anusha P O, Data Science Intern at Datahut. I build smart data-processing workflows and product-matching systems for e-commerce sites like Amazon and Flipkart. In this blog, I explain how we used TF-IDF vectorization, cosine similarity, and Jaccard distance to find matching products in large datasets. We turned messy, inconsistent product listings into clean, organized records ready for analysis.


At Datahut, we help businesses use e-commerce data. We design strong and scalable solutions for product matching. We also work on removing duplicate items from catalogs. We gather information about competitors too. Learn more about our expertise through our Datahut Services and discover how we extract and deliver data. You can also get to know our team and mission on the About Datahut page. For more insights and tutorials, visit the Blog to explore all posts. If you want to use data-driven methods to improve product discovery, pricing analysis, or catalog management, contact us through the chat widget on the right side of our website. Let’s turn your e-commerce data challenges into actionable insights.



FAQs


1. What are the best practices for tuning cosine similarity thresholds in product matching tasks?

The best practices include testing different threshold values on a labeled sample, analyzing false positives and false negatives, adjusting thresholds based on product category sensitivity, and using validation data to find a balanced score that maximizes accuracy while minimizing mismatches.


2. What are common issues in product matching implementations using cosine similarity and TF-IDF, and how can they be resolved?

Common issues include mismatched titles due to noisy text, missing brand/model information, and low similarity scores for products with different wording. TF-IDF may also over-weight rare words or under perform on short titles. These issues can be resolved by cleaning text (normalization, stop-word removal), adding more attributes (brand, model, specs), using better weighting, applying thresholds carefully, and combining cosine similarity with advanced methods like embedding for improved accuracy.


3. What challenges does product matching address in e-commerce platforms?

Cosine similarity combined with TF-IDF improves product matching. It turns product titles into numerical vectors. Then it measures how closely these vectors align. This improves search accuracy, price comparison, and overall shopping experience.


4. How does cosine similarity combined with TF-IDF help improve product matching in online stores?

Cosine similarity combined with TF-IDF improves product matching by converting product titles into numerical vectors and measuring how closely they align. TF-IDF highlights important keywords, while cosine similarity checks their directional similarity, allowing the system to identify matching products even when titles are written differently.


5. Product matching helps fix problems like duplicate listings and inconsistent product titles. It also handles different naming styles across sellers. It makes it easier to tell if two products are the same.

Yes, cosine similarity works well. It measures how closely two text vectors align. This makes it great for matching product titles, descriptions, and other e-commerce data. It works even when the wording is different.

Do you want to offload the dull, complex, and labour-intensive web scraping task to an expert?

bottom of page