# Web Scraping Guides & E-Commerce Data Insights | Datahut Blog > Tutorials, case studies, and data analysis on web scraping, e-commerce intelligence, and pricing trends — from the team that runs managed extraction at scale. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages _No public content available._ ## Posts ### When Your Competitor Runs Out of Stock, Their Customer Is Up for Grabs URL: https://www.blog.datahut.co/post/competitor-stockout-bid-trigger/ Last updated: 2026-09-07T07:58:02.000Z Most paid media teams treat ad spend as a static allocation problem: pick the top keywords or audiences, set bids, review performance weekly. But there's a pocket of high-intent demand that only exists for a few hours or days at a stretch, and it's sitting in plain view on your competitors' product pages every time they run out of stock. Call it availability-triggered conquesting. The signal isn't a keyword or an audience segment. It's a competitor SKU flipping to out-of-stock. It belongs to the same family as everything else in [web scraping for marketing](https://www.blog.datahut.co/post/webscraping-for-marketing-2025/): external data that your internal analytics will never show you, because it isn't about you. ## The idea in one sentence When a competing SKU goes dark, the shopper who wanted it doesn't disappear. They go looking for the next-best option, and if your ad isn't in front of them in that window, someone else's is. The behaviour is documented. DOSS surveyed 1,000 US adults in April 2026 and found that when shoppers hit a stock-out, 45% buy from a different retailer and 32% switch to a competing brand, at least temporarily ([Stockout Stigma Index](https://www.doss.com/research/stockout-stigma-index?ref=blog.datahut.co); [covered by Chain Store Age](https://chainstoreage.com/survey-out-stocks-often-lead-consumers-switch-brands?ref=blog.datahut.co)). "Temporarily" is doing real work in that sentence, and it's the reason speed matters: you're not buying a customer for life, you're buying the one order that was already decided. Take a category with a clean one-to-one competitive map. Running shoes, say. If you sell 100 products that go head-to-head with roughly 300 equivalent SKUs across three competitors, at any given moment some share of those 300 are unavailable: a size run sold out, a colour discontinued mid-restock, a promotion that emptied inventory faster than the supply chain could refill it. Each of those is a shopper who was ready to buy and just got blocked. ## Why this beats always-on conquesting Standard conquesting means bidding on a competitor's brand terms and running comparison ads against them. It works, but it's blunt. You pay to compete for that demand whether or not the competitor can actually fulfil it, and most of the time they can. Availability-triggered conquesting only spends aggressively in the window where the competitor is structurally unable to serve the customer. That changes two things. **Conversion intent is higher.** The shopper isn't choosing between brands. They were already sold and hit a wall. Your ad removes a blocker rather than manufacturing interest. **You get a defensible reason to flex bids.** "Competitor SKU X went out of stock at 09:40" is a clean, auditable trigger. That's far easier to greenlight in a budget review than an unexplained bid spike. One claim we'd be careful with: it's tempting to assume the auction thins out because a competitor with no stock stops defending the page. In practice plenty of retailers leave ads running on out-of-stock SKUs, either because nobody wired inventory into the ad account or because the campaign is at the category level. Don't build your business case on a cheaper auction. Build it on the conversion rate. ## Doesn't Smart Bidding already do this? Fair question, and the honest answer is partly. Target CPA and Target ROAS bidding watch conversion rate and bid up when it rises. If a competitor stock-out lifts your conversion rate, Google will eventually notice and pay more for that traffic on its own. So what does an explicit trigger add? Three things. **Latency.** Automated bidding is reactive. It needs conversion volume to accumulate before the signal clears noise, and on a single SKU's keyword set in a mid-volume category that can take most of a six-hour window. By the time the algorithm has enough data to move, the window is closing. A stock-out trigger fires on the first pageview of the competitor's product page, not the fifteenth conversion of yours. **It has no idea why.** Smart Bidding sees a conversion rate change. It doesn't know the cause, so it can't distinguish a competitor stock-out from a seasonal bump, and it can't act on the restock. When the competitor's inventory returns, the algorithm keeps bidding at the elevated level until performance decays enough to correct. An explicit trigger pulls spend back the moment the SKU flips green. **It can't touch creative or campaign state.** Bidding automation adjusts bids. It won't unpause a campaign, swap in a Responsive Search Ad variant that leans on availability, or shift budget between SKU pairings. If you run manual or Enhanced CPC, none of this is a debate. If you run Smart Bidding, the practical implementation isn't fighting the algorithm with manual bids. It's using seasonality adjustments, campaign state, and creative swaps as your levers, and letting the bidding do what it's good at. ## What the CPA math actually looks like Worth being precise here, because the arithmetic is easy to present dishonestly. Take a single mapped SKU pair in a mid-volume category: a $1.20 CPC and a 2% conversion rate on shoppers who land on your product page while the competitor's equivalent is in stock. That's a $60 cost per acquisition. Now the competitor goes dark on that SKU for six hours, and conversion rate on that traffic goes to 11%. (That number is illustrative, not measured. Your own before-and-after data is the only version of it worth planning against, and measuring it is the first thing to do before you automate anything.) The important point is what causes what. **The higher bid doesn't produce the higher conversion rate.** The stock-out does, and it would do so whether you bid a cent more or not. What the elevated conversion rate buys you is headroom: - At 2% conversion, a $60 target CPA supports a maximum CPC of $1.20. - At 11% conversion, the same $60 target CPA supports a maximum CPC of $6.60. So a 3x bid to $3.60 during the window isn't paying a premium. It's spending about half the headroom the window opened up, buying impression share you'd otherwise lose, and landing at roughly $33 CPA — nearly half your baseline. ![Line chart and horizontal bar graph illustrating the impact of a 6-hour competitor stock-out on Cost Per Acquisition (CAC) and demand captured across three strategies: Stock-out trigger, Smart Bidding, and Static bid.](https://www.blog.datahut.co/content/images/2026/08/stockout-cac-window.webp) Line chart and horizontal bar graph illustrating the impact of a 6-hour competitor stock-out on Cost Per Acquisition (CAC) and demand captured across three strategies: Stock-out trigger, Smart Bidding, and Static bid. That's the whole argument. Static bidding can't tell the difference between a shopper who's still comparing and one who's already been blocked once. A stock-out trigger can, and it can tell the difference in minute one rather than hour five. Miss the first hour because a human is reading a weekly report and you've spent the cheapest part of the window on nothing. ## How to wire the trigger into Google, Meta, and retail media The trigger itself is platform-agnostic: a webhook or API call fired the moment a mapped SKU flips state. What it calls differs. ![](https://www.blog.datahut.co/content/images/2026/08/slack-stockout-alert.webp) Stock out Alerts in Slack / Mattermost / Whatsapp / Telegram or whatever messaging tools you use **Google Ads.** [Google Ads Scripts](https://developers.google.com/google-ads/scripts/docs/start?ref=blog.datahut.co) run on a time-driven schedule down to hourly and can adjust keyword bids, toggle a paused campaign, or swap an ad variant. The Google Ads API gives you the same control on demand rather than on a schedule, which matters when your monitoring is more frequent than hourly. Scripts are the lower-lift path if you're not ready to build against the full API. **Meta.** The [Marketing API](https://developers.facebook.com/docs/marketing-api/?ref=blog.datahut.co) supports [rule-based budget and bid changes](https://developers.facebook.com/docs/marketing-api/ad-rules?ref=blog.datahut.co). Pair it with dynamic creative slots so "in stock now" messaging swaps in automatically instead of needing a hand-built asset per SKU. **Retail media (Amazon Ads, Criteo, and similar).** These platforms increasingly expose bid and budget APIs for exactly this kind of rule, but maturity varies. Check current API coverage before assuming parity with Google or Meta. You don't need any of this on day one. A daily Slack alert carrying the mapped SKU, the stock-out duration, and a suggested bid multiplier, actioned by whoever owns the account, is a reasonable way to prove the case before you build the automated version. Do that first. It also gives you the conversion-rate delta you need to replace the illustrative 11% above with your own number. This is a different mechanism from feeding your *own* stock status into Google Shopping or Meta dynamic product ads (see [our post on product feeds](https://www.blog.datahut.co/post/how-to-leverage-product-feeds-to-boost-your-e-commerce-business/)). That's your inventory driving your ads. This is the competitor's inventory driving your bids. ## Where this fits if you already run retail media tools If you sell on Amazon or Walmart and pay for an enterprise retail media platform, some version of this is probably already in your stack. Those tools adjust marketplace bids when a competitor ASIN goes dark, and Walmart's retail media ecosystem has published case studies on the uplift. This post isn't for that audience. It's for everyone selling on their own storefront, or running Google and Meta campaigns outside a marketplace, who doesn't have a six-figure retail media contract. The DIY version, built on scraped competitor data and a script instead of a bought platform. ## How to monitor competitor stock at the variant level Doing this for one product against one competitor is a five-minute manual check. Doing it continuously across a real catalogue needs three pieces. **1\. A maintained product map.** Your SKUs mapped to the closest equivalent at each competitor: same category, comparable price point, same use case. Get this wrong and you trigger conquesting on irrelevant signals. A $180 stability shoe going out of stock has nothing to do with your $90 neutral trainer. Two things make this harder than it sounds. First, you have to be right about who your competitors actually are, which is a data question rather than a gut one ([how to identify your true ecommerce competitors](https://www.blog.datahut.co/post/ecommerce-competitor-analysis/)). Second, catalogues shift constantly, and a SKU map that breaks every time a competitor renames a colourway isn't a map. The same stable-mapping problem shows up when teams try to [build a category price index](https://www.blog.datahut.co/post/category-price-index/), and the answer is the same: the mapping needs maintenance as a standing job, not a spreadsheet built at launch and never touched. **2\. Variant-level availability, not page-level.** This is where most first attempts break. A shoe sold out in sizes 8 to 11 but still showing "in stock" because a 6 and a 13 remain is functionally unavailable to most of the demand you care about. The aggregate badge on a product page lags reality more often than it reflects it. In practice the reliable signal usually isn't the badge at all. Most modern storefronts ship variant availability in structured form: `schema.org` product markup with an `offers.availability` field per variant, or a JSON blob in the page payload that the front end reads to grey out size buttons. Shopify stores expose a variants array with `available: true/false`. Read those, not the rendered HTML, because the rendered state depends on JavaScript that may not run the same way in your crawler as in a browser. Where a site genuinely renders availability client-side only, you need a headless browser for those pages, which changes the cost profile enough that it's worth knowing before you scope the project. Set monitoring frequency to match how fast the category turns over: hourly for fast-moving items, daily for slower ones. Hourly polling across a few hundred SKUs on a well-defended retail site is where the boring infrastructure problems start, which brings us to the rest of it. **3\. An automated bridge into the ad platform.** The moment a mapped SKU flips, that should become an action with no human in the loop: raise bids on the corresponding keyword set, unpause a campaign, swap in availability-led creative, or reallocate budget from a lower-priority pairing. By the time someone spots the gap in a weekly report, it's closed. ## Is it legal to monitor a competitor's stock levels? Publicly listed product pages are among the least contentious things to collect. There's no personal data involved, no login wall, and the information is displayed to every shopper who visits. Prices, availability, and product attributes on public catalogue pages have been scraped commercially for two decades. The constraints that actually bite are practical rather than legal: respect robots.txt, keep request rates low enough that you're not degrading someone's site, don't route around authentication or anti-bot measures in ways that breach a site's terms, and don't republish competitor content. If your legal team wants a sharper answer than this, get one before you start rather than after your crawler has been hammering a site for three months. ## The details that decide whether this works **Restocks deserve the same rigour as stock-outs.** Pull conquest spend back the moment inventory returns, otherwise you're defending a window that already shut. This is the half teams skip when they build the first version, and it's the half that quietly eats the gains. **Attribution needs its own tag.** If you can't separate revenue from stock-out-triggered conquesting from your always-on campaigns, you can't show the approach is incremental, and it gets cut in the next budget cycle for the wrong reason. **It generalises past footwear.** Electronics, beauty, home goods, and anything with seasonal or promo-driven demand are strong fits: categories where stock-outs are frequent, publicly visible, and reflect real unmet demand rather than deliberate scarcity marketing. DOSS's Reddit analysis found fashion and apparel generating the highest rate of stock-out complaints of any category at 7.9% of posts, with beauty and personal care (7.5%) and electronics (7.4%) close behind. Treat that as a rough proxy for where the frustration is loudest, not for where stock-outs actually happen. DOSS is explicit that the Reddit data doesn't confirm real stock-out events. ## Conquesting cuts both ways Everything above is offence. A mature programme also runs defence, because once you start bidding aggressively whenever a competitor stocks out, expect the reverse. The same monitoring that flags a competitor's stock-out should flag when your own SKUs are exposed: low stock, a price gap, a lapsed brand-term defence. Better that your team sees it before a rival's system does. If you're already running [competitor price monitoring](https://www.blog.datahut.co/post/how-to-leverage-web-scraping-to-create-a-competitor-price-monitoring-strategy/), you have most of this pipeline already. Availability is one more field on pages you're crawling anyway, and the marginal cost of collecting it is close to zero compared to standing the crawl up in the first place. ## The part that's actually hard The strategy is simple once you frame it as an inventory problem instead of a bidding problem. What's hard is keeping structured, near-real-time competitor availability data flowing accurately, at variant level, mapped correctly as catalogues change underneath you, across hundreds of SKUs, every day, indefinitely. That's a data infrastructure problem before it's a marketing one. Continuous crawling of competitor product and category pages, variant-level parsing that survives a front-end redesign, normalisation into a feed your campaign automation can consume, and monitoring that tells you when a parser silently starts returning "in stock" for everything. Get that wrong — stale data, broken variant detection, an unmaintained SKU map — and the signal you're trading on is noise dressed up as a trigger. Most paid media teams don't want to own that pipeline, and shouldn't have to. This is what our [ecommerce scraping service](https://www.datahut.co/solutions/ecommerce-web-scraping?ref=blog.datahut.co) exists for: you get the clean, mapped, variant-level availability feed, not another crawler to babysit. [Talk to a Datahut data expert](https://www.datahut.co/contact?ref=blog.datahut.co) about what your competitor set would take to monitor, and we'll come back with a feasibility read on the specific sites. ## Frequently asked questions **What is availability-triggered conquesting?** Bidding aggressively on a competitor's demand only during the windows when they're out of stock and can't serve it. Standard conquesting runs continuously, whether or not the competitor can fulfil the order. Availability-triggered conquesting uses their inventory state as the on/off switch, so spend concentrates in the hours when their shopper has nowhere else to go. **How do I know when a competitor is out of stock?** Read the structured data on their product pages, not the visible badge. Most storefronts publish variant-level availability in schema.org product markup under `offers.availability`, or in a JSON payload the front end uses to grey out size buttons. Shopify sites expose a `variants` array with an `available` flag per variant. The rendered "In Stock" badge is the least reliable signal on the page — it usually reflects the parent product rather than the specific size or colour a shopper wants. **How often should I check competitor stock levels?** Match the polling interval to how fast the category turns over: hourly for fast-moving items, daily for slower ones, every 15 minutes for high-value SKUs in volatile categories. The interval sets a ceiling on how much of any window you can act on. If a typical stock-out lasts six hours and you poll daily, you'll usually find out after it's closed. **Does this work with Google Smart Bidding?** Yes, but the lever changes. You don't override Smart Bidding with manual bids — you feed it a seasonality adjustment for the window, or you change campaign state and creative, which bidding automation can't touch on its own. Smart Bidding will eventually detect the conversion rate lift, but it needs conversion volume to accumulate first, and it can't act on the restock because it never knew the cause. **Is it legal to monitor a competitor's stock levels?** Collecting publicly listed product pages is about as uncontentious as web data collection gets: no personal data, no login wall, and the information is shown to every shopper who visits. The constraints that matter are operational — respect `robots.txt`, keep request rates modest, don't route around authentication, and don't republish competitor content. Get a view from your own legal team before you start, not after. **How large a bid increase is justified during a stock-out?** Work backwards from your target CPA, not from a multiplier. Your maximum defensible CPC is target CPA × the conversion rate you see during the window. At a $60 target and an 11% in-window conversion rate, that's $6.60 — five and a half times a $1.20 baseline. Most teams should spend well under the ceiling. The goal is capturing volume at an improved CPA, not spending the entire margin the window opened. **What if I don't have a competitor SKU map?** Build one for a narrow slice first — ten to twenty of your highest-margin products against a single competitor. A full-catalogue map is a large project and unnecessary before you've measured whether the conversion lift is real for your category. The map is also the part that decays fastest as catalogues change, so it needs an owner and a review cadence, not a one-time build. **Which categories does this work best in?** Categories where stock-outs are frequent, publicly visible, and reflect genuine unmet demand: footwear and apparel, beauty, consumer electronics, and anything with seasonal or promotional demand spikes. It works poorly where scarcity is deliberate — limited drops, manufactured exclusivity — because the shopper who wanted that item usually isn't looking for a substitute. **How do I prove the approach worked?** Tag stock-out-triggered spend separately from your always-on campaigns before you start, not after. Without that separation you can't show the revenue was incremental, and the approach tends to get cut in a budget review for reasons unrelated to performance. The comparison worth tracking is CPA and conversion volume inside triggered windows against the same SKU set outside them. ## References - [Stockout Stigma Index](https://www.doss.com/research/stockout-stigma-index?ref=blog.datahut.co) — DOSS, survey of 1,000 US adults (April 2026) plus analysis of 8,679 Reddit posts, on switching behaviour and category-level complaint rates. [Coverage in Chain Store Age](https://chainstoreage.com/survey-out-stocks-often-lead-consumers-switch-brands?ref=blog.datahut.co). - [Google Ads Scripts documentation](https://developers.google.com/google-ads/scripts/docs/start?ref=blog.datahut.co) — scheduling bid, campaign, and creative changes. - [Meta Marketing API](https://developers.facebook.com/docs/marketing-api/?ref=blog.datahut.co) and [Ad Rules Engine](https://developers.facebook.com/docs/marketing-api/ad-rules?ref=blog.datahut.co) — rule-based bid and budget automation. - [What brands need to know about conquesting at Walmart](https://www.profitero.com/blog/walmart-competitor-conquesting-brand-strategies?ref=blog.datahut.co) — Profitero, on inventory-aware conquest timing in retail media. ### How India's DPDP Act Is Changing Web Scraping: What Every Business Should Know URL: https://www.blog.datahut.co/post/how-indias-dpdp-act-is-changing-web-scraping-what-every-business-should-know/ Last updated: 2026-09-07T08:10:16.000Z ## India's DPDP Act: Introduction Imagine your marketing team has been scraping contact data from public directories for years. Then your legal team forwards you a notification about the DPDP Act. Suddenly, that routine pipeline is a compliance question nobody has a clean answer to. That's the situation many Indian businesses are walking into right now. [India's DPDP](https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf?ref=blog.datahut.co) Act has technically been in effect since August 2023\. For two years, it sat on paper with no regulator, no rules, and no penalties. That changed in November 2025\. The rules are now live, the Data Protection Board is operational, and full enforcement begins in May 2027\. For businesses relying on web scraping, that's not as much time as it sounds. Here's what[ the DPDP Act ](https://www.dpdpa.com/?ref=blog.datahut.co)actually means for your web scraping and what your business needs to do before enforcement begins. ## **What Impact Does the DPDP Act Have on Web Scraping?** The DPDP Act changes how businesses collect, store, and use personal data, and web scraping operations are directly in scope. Following[ ](https://static.pib.gov.in/WriteReadData/specificdocs/documents/2025/nov/doc20251117695301.pdf?ref=blog.datahut.co)[MeitY's notification of the DPDP Rules](https://static.pib.gov.in/WriteReadData/specificdocs/documents/2025/nov/doc20251117695301.pdf?ref=blog.datahut.co) in November 2025, businesses can no longer treat data collection as a purely technical decision. There are legal obligations attached to it now. ![What Impact Does the DPDP Act Have on Web Scraping?](https://www.blog.datahut.co/content/images/2026/08/b3461d_61427ecb71864bc385b1ca033c8d61f6-mv2.webp) ## Is Your Web Scraping Safe or Risky Under the DPDP Act? | Business Use Case | Usually Safe ✅ | Higher DPDP Risk ❌ | | ---------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------- | | Competitor pricing | Product prices, discounts, stock availability | Customer accounts, buyer profiles | | Product catalogue monitoring | SKUs, descriptions, specifications | Customer wishlists, order history | | Airline fare tracking | Ticket prices, routes, schedules | Passenger names, booking IDs, traveller details | | Hotel price monitoring | Room rates, availability, ratings | Guest names, booking details, reviews with identifiable information | | FMCG market intelligence | Product listings, pricing across platforms | Loyalty programme data, customer purchase history | | Real estate analytics | Property prices, amenities, location details | Owner names, broker phone numbers, personal email addresses | | Recruitment analytics | Job titles, salaries, required skills, company names | Candidate resumes, applicant contact details | | Financial services | Interest rates, loan eligibility criteria, product features | Customer financial records, personal account details | | Healthcare market research | Medicine prices, stock availability across pharmacies | Patient names, prescriptions, health records | | News aggregation | Article headlines, metadata, publication dates | Subscriber information, reader profiles | | Marketplace seller analytics | Product ratings, pricing, inventory levels | Seller phone numbers, personal email addresses | **Key insight:** The difference between compliant and non-compliant scraping often comes down to one question: *are you collecting commercial information or information about an identifiable individual?* Stay on the left side of this table, and your DPDP exposure will be significantly reduced. ## What Your Business Must Get Right: Consent Under DPDP Under the DPDP Act, consent must be clear and in plain language: explaining exactly what data is collected and why. Buried checkboxes and vague terms no longer hold up legally. The Act also gives individuals the right to access, correct, or delete their data whenever they choose. Users must also be able to withdraw consent just as easily as they gave it. *(Section 6, Digital Personal Data Protection Act, 2023)* ## What Happens If Your Business Suffers a Data Breach Under DPDP? The DPDP Act requires businesses to notify the Data Protection Board and affected individuals immediately upon discovering a breach. *(Section 8(6), DPDP Act 2023)*. The [DPDP Rules 2025](https://www.meity.gov.in/static/uploads/2025/11/53450e6e5dc0bfa85ebd78686cadad39.pdf?ref=blog.datahut.co) specify this window as 72 hours; missing that deadline carries penalties up to ₹200 crore. ## What Are the Rules for Collecting Children's Data Under DPDP? ## If your platform could in any way be accessed by anyone under 18, whether it's gaming, education, or e-commerce, verifiable parental consent is required before processing their data. The Act also strictly prohibits tracking, behavioural monitoring, and targeted advertising directed at children. *(Section 9, DPDP Act 2023)* ## How Does the DPDP Act Define Significant Data Fiduciaries? Large platforms processing personal data at scale - e-commerce companies, social media platforms, and similar businesses - may be classified as Significant Data Fiduciaries. This comes with stricter obligations: appointing a Data Protection Officer based in India and undergoing independent data audits. (Section 10, DPDP Act 2023) ## What Are the Penalties for Non-Compliance Under the DPDP Act? Penalties under the DPDP Act can reach up to ₹250 crore per instance. The Data Protection Board is operational as of November 2025 - it has the power to investigate, summon, and impose fines. This is not a future risk. It is a present one.(Section 33(1), DPDP Act 2023) The Board weighs the severity and your corrective actions before imposing a fine. (Section 33(2), DPDP Act 2023) ## **What are the Major Risks of Non-Compliant Web Scraping?** ![What are the Major Risks of Non-Compliant Web Scraping?](https://www.blog.datahut.co/content/images/2026/08/b3461d_7f455217b1d748378f6dcb9cc037fe38-mv2.webp) ### **Risk 1: Heavy Financial Penalties** The fact that the DPDP Act has a tiered system of penalties is important. | Violation Type | Maximum Penalty | | -------------------------------------------------- | ---------------------------- | | Failure to maintain reasonable security safeguards | **₹250 crore per violation** | | Consent obligation violations | **₹50 crore per instance** | | Failure to notify the Board of a breach | **₹200 crore** | ### **Risk 2: Purpose Limitation Violations** Data collected for one purpose cannot be used for another without fresh consent. It’s easy for teams to quietly repurpose existing scraping pipelines for new internal uses, completely unaware that doing so can trigger legal and compliance risks. A marketplace intelligence company tracks seller performance across platforms such as product ratings, pricing, and inventory levels. That's compliant commercial data. But when the same dataset is handed to the sales team to build outreach lists using seller phone numbers and personal email addresses, the original purpose has been crossed. The data was collected for market analytics; using it for direct outreach is a separate act that requires fresh consent under the DPDP Act. ### **Risk 3: Third-Party and Vendor Liability** Outsourcing your scraping doesn't outsource your legal risk. Under the DPDP Act, your business remains legally accountable for how your data is collected. This makes vendor selection a legal decision, not just a commercial one. ### **Risk 4: Reputational Damage** A Data Protection Board investigation is public. Even if it ends without a fine, it will still be part of your company's public record. A compliance investigation surfacing during a procurement process can cost you a deal worth far more than any fine. **Key takeaway:** The ₹250 crore figure gets attention, but for most businesses, the bigger day-to-day risks are vendor liability, pipeline disruptions, and what happens when a Board investigation goes public. ## **How to Stay Compliant: Web Scraping Under the DPDP Act** ![How to Stay Compliant: Web Scraping Under the DPDP Act](https://www.blog.datahut.co/content/images/2026/08/b3461d_a9ce6ab810914db9aaa5156aea8139ba-mv2.webp) ### **S**tep 1: Audit your existing data pipeline. Start by mapping every active data pipeline in your business. What data are you collecting, from which websites, where is it stored, and who can access it? Most businesses can't answer all four questions cleanly, and that gap is exactly where compliance risk hides. An Amazon seller scrapes competitor product prices every morning to adjust their own pricing automatically. This is generally low risk: product prices, discounts, and stock availability are commercial information, not personal data. But if that same scraper picks up customer review names or seller contact details, those elements fall under the DPDP Act. ### Step 2: Identify whether your data is personal or non-personal Personal data is any digital information that can identify an individual directly or indirectly. In web scraping, this includes names, emails, phone numbers, and location data. A travel aggregator monitors flight prices every hour to recommend the cheapest fares. Collecting route information, airline names, schedules, and ticket prices is different from collecting passenger names, booking IDs, or traveller contact details. The latter introduces personal data obligations. Takeaway: Monitor fares - not travellers ### Step 3: Make Consent Clear, Specific and Withdrawable Valid consent under the DPDP Act must be clear, specific, and deliberate, not buried in a terms document or pre-ticked by default. Users must also be able to withdraw consent just as easily as they gave it. ### Step 4: Establish a Clear Purpose and Stick to It ### The DPDP Act requires a clear, lawful purpose for every data collection activity. Document exactly why you need the data before you start scraping, not after. And if that purpose changes, you need fresh consent before repurposing it. A recruitment platform analyzing job postings for hiring trends is generally on safe ground. The risk starts when the same pipeline begins pulling candidate resumes or applicant profiles. That data requires explicit consent. And if it was collected for talent matching but later used to build a marketing list, that second use is a separate DPDP violation, even if the original collection was legal. ### Step 5: Target Site Pre-Screening Before running any scraper, check the target site's robots.txt file and Terms of Service. Bypassing these restrictions can constitute unauthorised access under Section 43 of the IT Act, which immediately undermines any DPDP compliance claims your business might otherwise have. ### Step 6: Automate Data Expiry and Deletion Under the DPDP Act, holding onto personal data indefinitely isn't an option. Once data has served its purpose, it needs to go. Set clear shelf lives for every dataset, build automatic deletion routines, and keep clean logs, as you may need to show them during an audit. ### Step 7: Set Up a Breach Response Plan Write your breach response plan before you need it. Be specific: who detects the breach, who escalates it internally, who notifies the Data Protection Board, and who communicates with affected individuals. The DPDP Act gives you just 72 hours. Missing that deadline carries a penalty of up to ₹200 crore. ## **What Should Web Scraping Businesses Prioritise Under the DPDP Act?** If your business depends on web scraping for day-to-day decisions, these four areas need leadership attention. ![What Should Web Scraping Businesses Prioritise Under the DPDP Act? ](https://www.blog.datahut.co/content/images/2026/08/b3461d_b18b3e6ad4ca42ad9f607778b1ca3793-mv2.webp) ### 1\. Shift Focus to Non-Personal Data Commercial scraping remains fully viable if you avoid personal data. Pricing intelligence, market trends, product catalogues, news monitoring- none of these requires touching individually identifiable information. Build your pipelines around non-personal data wherever possible, and you significantly reduce your compliance exposure. ### 2\. Maintain an Internal Data Mapping Registry Know exactly where scraped data enters your business, who accesses it, where it lives, and when it gets deleted. Without this registry, you cannot respond to a Data Protection Board inquiry. ### 3\. Build a "Data Principal Rights" Portal If personal data is stored in your databases, individuals have a legal right to demand its deletion under the DPDP Act. Build a clear, accessible mechanism for users to make that request, and make sure someone inside your organisation is responsible for acting on it promptly. ### 4\. Audit Third-Party Data Vendor Buying ready-made scraped datasets or lead lists from vendors carries real legal risk. Under the DPDP Act, if their data was collected illegally, your business shares the liability, not just the vendor. Make them show documented proof of how consent was obtained or how PII was removed before you sign anything. ## Conclusion The DPDP Act is already active, and it does not prohibit web scraping. But it does require businesses to collect data with a clear purpose, handle it transparently, and take full accountability for how it is used. Businesses that build compliance into their data processes today will face far less disruption when full enforcement begins on 13 May 2027\. Those that don't will be scrambling to catch up. ## **Keep Scraping. Stay Compliant. Here's How.** ![](https://www.blog.datahut.co/content/images/2026/08/b3461d_c6404399d6324292aaefb3654a012174-mv2--1-.webp) A consumer electronics brand wanted to monitor competitors across Amazon, Flipkart, and Croma. Their scraping pipeline was configured to collect product names, specifications, prices, discounts, and ratings, but deliberately excluded customer names, reviewer profiles, and seller contact details. By designing around commercial data from the start, the company got the market intelligence it needed without touching a single data point that triggers DPDP obligations. [Datahut](https://www.datahut.co/services/data-as-a-service?ref=blog.datahut.co) has been helping businesses extract web data for 14 years. As India's DPDP Act moves toward full enforcement, we handle the compliance, so your team doesn't have to. From PII filtering to purpose-limited pipelines, we deliver clean, compliant data while you focus on decisions that grow your business. [Talk to a Datahut expert today](https://www.datahut.co/contact?ref=blog.datahut.co) to get a free consultation and build a DPDP-compliant data strategy for your organisation. ## Frequently Asked Questions( FAQs) **1\. What is the DPDP Act?** The Digital Personal Data Protection (DPDP) Act, 2023 is India's first comprehensive law governing how digital personal data is collected, stored, and used. It gives individuals greater control over their data and places clear legal obligations on businesses that process it. It moves to full enforcement in May 2027. **2\. Is Web scraping illegal under the DPDP Act?** Web scraping is not illegal under the DPDP Act, provided that you are scraping non-personal data. For example, e-commerce product prices, stock availability, or information from public business directories. If you want to carry out a thorough examination of global precedents, see our guide on whether[ web scraping is legal](https://www.blog.datahut.co/post/is-web-scraping-legal/). **3\. What counts as personal data under the DPDP Act?** Under the DPDP Act, personal data is any digital data that can identify an individual directly or indirectly. This includes obvious identifiers like names, email addresses, and phone numbers, as well as less obvious ones like location data, financial details, and behavioural profiles built by combining multiple data points. **4\. Does the DPDP Act apply if we scrape data that is already public?** **Yes.** Don't fall into the trap of assuming "publicly visible" means "free to collect." Under the DPDP Act, publicly accessible personal data like user profiles, contact details, or reviews is still protected. If your scrapers touch personal details, you are still legally required to follow core privacy rules. **5\. What happens if my business is not DPDP compliant by May 2027?** Non-compliance after May 2027 exposes your business to financial penalties, regulatory investigation, and reputational damage. The smarter move is building compliance into your data processes now, while there's still time to do it without pressure. ### Olay vs. L'Oréal Paris on Amazon: Which Beauty Brand Really Wins? URL: https://www.blog.datahut.co/post/olay-vs-loreal-paris-on-amazon/ Last updated: 2026-09-07T08:18:35.000Z In short, neither brand clearly comes out on top since they succeed in different ways. Olay does better in terms of price and in building customer loyalty based on the sensory experience of skincare, having an average price of $25.82 and receiving reviews that focus on the scent and the softness of the skin. L'Oréal Paris, on the other hand, excels in terms of scale and consistency, with 4.3 times as many reviews (2.8 million compared to 651 thousand), a much broader range of 27 products, and a more tightly grouped set of ratings. Although both have the same average rating of 4.44, they achieve this result by following completely different strategies: Olay by concentrating on premium-mid range skincare, and L'Oréal by offering a wide range of products across skincare, haircare, and makeup. The analysis is based completely on structured [Amazon](https://www.blog.datahut.co/post/what-can-you-get-from-scraping-data-from-amazon/) catalog and review data obtained by Datahut - the same underlying datasets that are used in our individual breakdowns of the [Olay ](https://www.blog.datahut.co/post/olay-on-amazon/)and [L'Oréal](https://www.blog.datahut.co/post/l-or%C3%A9al-paris-on-amazon/) Paris brands. The questions that a marketer, a category manager, or a member of a [competitor-research](https://www.blog.datahut.co/post/ecommerce-competitor-analysis/) team would actually ask when comparing the two are set out below. ## **What Is This Comparison Based On?** [The comparison is based on data from the Amazon](https://www.blog.datahut.co/post/amazon-product-data-scraping/) US catalog for both brands, including information on active product listings, prices, discount percentages, star ratings, the number of reviews, and analysis of review text to identify factors that led to purchases. The dataset for Olay includes 266 active products in 14 different product forms, while that for L'Oréal Paris includes 300 active products in 27 product forms. The same methodology was applied in carrying out both analyses, and this common methodology enables the comparison to be direct rather than comparing apples to oranges. ## **How Do Olay and L'Oréal Paris Compare in Catalog Size?** [L'Oréal Paris](https://www.loreal.com/?ref=blog.datahut.co) runs a broader, more fragmented catalog. [Olay](https://www.olay.com/?ref=blog.datahut.co) runs a narrower, more concentrated one. ![Olay vs Loreal Paris Products](https://www.blog.datahut.co/content/images/2026/08/Olay-Loreal-2.webp) L'Oréal's 27 product lines include those for the skin, hair, and color cosmetics - the reason for the brand's existence is so that it will have a product available for almost every beauty-related search on Amazon. Olay remains focused within the field of skincare and personal care, with nearly half of its range devoted just to Cream and Body Wash. [**When the question is which brand covers more search intent, L'Oréal comes out on top in terms of range. But when the question is which brand goes the deepest into one category, Olay is the winner in terms of depth.**](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/) ## **Which Brand Has More Reviews, and Why Does That Matter?** L'Oréal has around 4.3 times as many reviews as Olay -2,824,778 as compared to 651,038, even though its catalog is only about 13% bigger when measured by the number of SKUs. Review volume serves as a rough estimate of purchase volume and market reach. The review base for L'Oréal is due to its extensive distribution of colour cosmetics and haircare products, categories in which Olay does not compete - along with high-frequency categories such as Liquid (with 851,000 or more reviews across 97 SKUs) and Cream (with 728,000 or more reviews across 85 SKUs). Nearly all of Olay's review base is concentrated in a single format: Cream alone makes up 363,514 reviews, which is 55% of all the reviews the brand has gathered. **A category analyst can conclude from this that Olay's growth is limited by the breadth of the category since it has already reached the limits of its influence in the "Cream" segment, and that L'Oréal's growth is hindered by its performance across a large number** of smaller formats, each of which has its own set of competitive dynamics. ## **How Do Their Pricing Strategies Differ?** Olay sits in premium-mid skincare pricing at $25.82 average. L'Oréal Paris sits in mass-market pricing at $15.41 average - roughly 40% cheaper. ![Olay vs Loreal Price comparison](https://www.blog.datahut.co/content/images/2026/08/467a063c-cb78-48df-a9c9-2d254154f2d6.webp) It is interesting that the same $15 to $30 mid-range segment is the one in which shoppers in the beauty sector report the highest level of satisfaction in relation to the price they pay for both brands. However, the two brands differ at the top end: in Olay's premium range (products priced at $30 and above) the ratings fall to 4.29, indicating that consumers compare more expensive Olay products with those of luxury brands. In contrast, L'Oréal's premium range maintains a score of 4.42, which is close to the average across its range - this is what Datahut's analysis describes as a "masstige" positioning that does not break down at higher price levels. ## **Which Brand Discounts More Aggressively and Why?** L'Oréal discounts significantly harder - 15.23% average versus Olay's 9.44% - but both brands use discounting surgically, not as a blanket strategy. ![Olay vs Loreal paris discounts](https://www.blog.datahut.co/content/images/2026/08/OLay-loreal-3.webp) The two brands adopt the same basic approach: they maintain their profits on the SKUs that are already successful and use discounts to gain trial in those categories where they are still building up their market share. Olay applies the largest discounts to its Serums since it is putting money into this higher-margin but more competitive area. With respect to L'Oréal, it is giving the biggest discounts on Masks (55%) and Kits (36%) because these products are low-commitment entry points intended for attracting new customers into the wider range of products - a strategy that Datahut's analysis refers to as the "gateway product" formula. **The extent of L'Oréal's discounting (with some items being as high as 55%) is in line with the overall range of its product offerings, since it has more "gateway" categories to introduce customers to than Olay does.** ## **Which Brand Is More Resilient to Losing a Bestseller?** Both brands are unusually well-distributed - no single product dominates either portfolio but L'Oréal is marginally more resilient. ![Olay vs loreal reviews](https://www.blog.datahut.co/content/images/2026/08/olay-loreal-4.webp) Compared with the review concentration observed in many consumer-product portfolios, both Olay and L'Oréal show unusually distributed review volume, indicating that the disappearance of one or two of its key products would indeed represent a real business risk. This is not the case with either Olay or L'Oréal; both of them spread about three-quarters of their review volume over 250 or more other products and are therefore not vulnerable to any single product failing. It is a genuinely rare feature at this level of product range, and it is probably the most significant similarity between the two brands. ## **What Actually Drives Purchase Decisions for Each Brand?** **I**n short, customers who shop at Olay base their purchases on sensory and functional trust; this includes scent, quality, and skin softness, while those who shop at L'Oréal buy primarily for visible results at first, supported by the product's effectiveness. ![](https://www.blog.datahut.co/content/images/2026/08/Olay-Loreal-5.webp) The most distinct difference in strategy between the two brands is this: Olay is offering a promise that is both sensory and functional, focusing on how the product feels and performs on the skin, and doing so at a price point that matches that of top-end alternatives. In contrast, L'Oréal starts by promoting a visible change (such as hair color or cosmetics), with product effectiveness and quality serving as the supporting evidence. **The two types of strategic approach are not transferable between the brands' main product areas, which is why they can both exist in the same market without one eating into the other.** ## **Which Brand Should You Benchmark Against?** This depends on what you're trying to learn: - **Benchmarking a premium-adjacent skincare brand?** Use Olay. Its pricing discipline, mid-tier sweet spot, and sensory-driven review language are the more relevant comparison set. - **Benchmarking a mass-market, multi-category beauty brand?** Use L'Oréal Paris. Its catalog breadth, tiered discount strategy, and consistency-at-scale (a 0.16-point ratings spread across 27 product forms) are the harder benchmark to hit and the more instructive one if your brand also spans skincare, haircare, and cosmetics. - **Studying portfolio resilience?** Either brand is a strong reference both keep top-10 concentration under 26% of total reviews, a genuinely uncommon result worth modeling regardless of category. ## **Turn Comparisons Like This Into a Repeatable Process** This analysis took two structured Amazon datasets and a shared methodology to produce. If you're tracking competitors across a category - not just one brand at a time ; the real value comes from doing this continuously: fresh pricing, fresh discount data, fresh review-driver mining, refreshed on a schedule instead of a one-time pull. Datahut builds exactly this kind of structured, automatically refreshed marketplace data for teams doing competitive and category benchmarking. If you want a similar side-by-side built for your own category,[ request a free data audit](https://www.datahut.co/contact?ref=blog.datahut.co) or talk to a Datahut strategist about ongoing category monitoring. ## **Frequently Asked Questions** **Is Olay or L'Oréal Paris rated higher on Amazon?** They're effectively tied - both average 4.44 out of 5.0\. Olay's rating varies more by product form (a 0.46-point spread), while L'Oréal's is more consistent (a 0.16-point spread across 27 forms). **Which brand is cheaper on average?** L'Oréal Paris, at $15.41 average versus Olay's $25.82 , roughly 40% less expensive per product. **Which brand has more Amazon reviews?** L'Oréal Paris, with 2,824,778 reviews compared to Olay's 651,038 - about 4.3x more, driven by its wider catalog spanning color cosmetics and haircare in addition to skincare. **Does either brand rely too heavily on one hero product?** No. Olay's top 10 products account for 25.32% of all reviews; L'Oréal's top 10 account for 24.09%. Both are unusually well-distributed compared to typical consumer brands, where a handful of SKUs often drive the majority of engagement. **What price tier performs best for both brands?** The $15–$30 mid-range tier is the strongest for both - Olay rates 4.51 there and L'Oréal rates 4.48\. It's the one place both brands' data agrees almost exactly. **Which brand discounts more?** L'Oréal Paris, at a 15.23% average discount versus Olay's 9.44%. L'Oréal also uses much deeper discounts on specific categories (up to 55% on Masks) to drive trial in niche formats. ### Amazon's Beauty Market: What 60 Bestsellers Reveal! URL: https://www.blog.datahut.co/post/amazons-beauty-market-bestsellers/ Last updated: 2026-09-07T08:18:37.000Z The [Beauty & Personal Care](https://sell.amazon.com/blog/products-to-sell?ref=blog.datahut.co) category on Amazon represents one of the platform's most competitive and consumer-driven digital landscapes. With 40 active brands across 60 top-performing listings, the digital shelf is anchored by an accessible category average sale price of $13.47 and a remarkably high average rating of 4.59 out of 5.0 - setting one of the highest satisfaction benchmarks across all major Amazon categories. In a market this dense, capturing and maintaining consumer attention requires deliberate strategic positioning. Category leaders like essence and CeraVe currently dominate the space, but they do so through opposing playbooks — one winning on hyper-focused Hero SKU efficiency and the other on a multi-listing portfolio approach. For mid-tier and emerging brands, understanding how pricing discipline, [discount strategy](https://www.blog.datahut.co/post/6-signs-your-business-has-pricing-fatigue-and-how-to-fix-it/), and consumer sentiment shape conversions is the only proven path to gaining visibility and stealing market share on the [Amazon](https://www.blog.datahut.co/post/amazon-product-data-scraping/) digital shelf. ## **Key Findings from Amazon's Beauty Market** Based on our analysis of 60 top-ranking product listings from Amazon's Beauty Market, we find the following key structural themes: - **The Massive Review Moat:** A tiny fraction of products controls the market share of attention. The Top 10 products capture 44.49% of all [customer reviews](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/) in the category - a concentration level that makes organic displacement nearly impossible without a purpose-built Hero SKU strategy. - **Budget Dominance:** Unlike other premium-driven categories, the budget tier ($15 or less) dominates consumer choices, accounting for 70% of all active [best-selling](https://sell.amazon.com/blog/amazon-best-sellers-rank?ref=blog.datahut.co) listings (42 SKUs). Critically, these budget products maintain the highest average rating in the category at 4.62 stars, proving that price reduction is not a trade-off for quality perception. - **Engagement Efficiency:** Volume leaders list multiple products, but efficiency leaders like essence, Mighty Patch, and MAYBELLINE generate massive review numbers on single-SKU listings. This proves that a hyper-focused "Hero Product" strategy is a fully viable and often superior path to search dominance. ## **Amazon Beauty & Personal Care Category Overview** ![Amaon bestsellers category overview](https://www.blog.datahut.co/content/images/2026/08/b3461d_5ff85664c4b74ef1bf3a2174e84e59db-mv2-1.webp) **Category Leaders** **essence** is the undisputed leader in engagement efficiency. By focusing on a lean, highly optimized presence, they have built a review barrier of 413,046 reviews on a single SKU - making them nearly impossible to outrank organically. Their strategy prioritizes hyper-focused "Hero SKU" conversion over mass catalog variety. **CeraVe** holds close second position, balancing shelf space and high velocity with 3 active listings yielding 337,069 total reviews. Their multi-SKU approach ensures they capture maximum share of consumer attention across skin type and product-format searches. ## **How Brands Compete: Product Distribution** This analysis details which brands occupy the most room on the digital shelf. By keeping multiple products in the bestseller ranks, these brands capture the highest share of consumer attention before a click even happens. ![Amazon SKU count](https://www.blog.datahut.co/content/images/2026/08/b3461d_e58431bc237b4120a3d97a19e1258808-mv2.webp) **Key Insights** - **Shelf Space Leaders:** Amazon Basics and The Ordinary each hold 4 active bestseller listings - the widest portfolio footprints in the category. This breadth ensures maximum search visibility across multiple sub-categories. - **Multi-SKU Anchors:** CeraVe, La Roche-Posay, Dove, and medicube each maintain 3 active bestseller listings, forming a dense mid-tier of brands that dominate shelf presence through consistent multi-product representation. - **Attention Economics:** If a brand owns 4 out of 60 bestseller listings, they effectively control approximately 7% of consumer attention before a single click occurs. In a category with 40 competing brands, this scale advantage is a powerful compounding force for organic discovery. ### **Which Brands Have the Highest Average Price?** ![Amazon bestsellers average price](https://www.blog.datahut.co/content/images/2026/08/b3461d_f4984821671046cbb1f0cc39dd63b3e1-mv2.webp) **Key Insights:** - **The Premium Peak:** EltaMD ($45.00) sits at the top of the pricing spectrum, representing the luxury baseline for this category - nearly 3.3x higher than the market average of $13.47\. Their premium positioning is sustained by strong dermatologist recommendations and clinical SPF formulations. - **The Competitive Sweet Spot:** The majority of consumer brands cluster below the $20 mark. Premium challengers like Paula's Choice ($25.90) and La Roche-Posay ($21.32) must lean heavily on specialized, clinically-validated formulations to justify prices above this cluster and maintain search relevance. - **Value Cluster:** grace & stella ($20.36) and MEDITHERAPY ($19.99) represent the top of the mid-range "aspirational value" band - offering perceived premium quality at accessible price points, a positioning that drives strong conversion in a budget-dominant category. ## **Brand Discount Strategy Snapshot** [Promotional behavior](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-dos-and-donts-of-dynamic-pricing-in-retail?ref=blog.datahut.co) across the Beauty & Personal Care category reveals a deliberate and calculated split between brands driving aggressive high-volume trial and those [protecting hard-built premium margins](https://www.mckinsey.com/industries/retail/our-insights/solving-the-paradox-of-growth-and-profitability-in-e-commerce?ref=blog.datahut.co). While the category sees a moderate promotional baseline, individual brand strategies vary dramatically based on their market positioning and acquisition goals. ![Amazon bestsellers discount strategy](https://www.blog.datahut.co/content/images/2026/08/b3461d_321e26a7313b47438cc48ca4068af10e-mv2.webp) **Key Insights** - **The Aggressive Recruiters:** KAHI (38.00%) and Paula's Choice (30.00%) deploy the deepest promotional cuts in the category. This signals aggressive customer acquisition tactics - using steep price incentives to drive quick trial volume among price-sensitive shoppers and rapidly build review momentum on newer or lower-visibility listings. - **Mainstream Challengers:** Neutrogena (26.50%) and Garnier (26.00%) maintain elevated discount rates consistent with their high-volume, mass-market positioning. These brands use broad promotional coverage to compete in a crowded mid-shelf environment and defend market share against private-label alternatives. - **Controlled Discounting:** L3 (25.00%) rounds out the high-discount tier, using strategic promotions to maintain shelf visibility without fully eroding brand equity - a tactic common among brands positioned between budget and mid-range. **The Strategic Pattern** - **Promote for Trial:** High-discount brands like KAHI use aggressive cuts (near 40%) to funnel price-sensitive shoppers into their listings and accelerate early review accumulation. - **Protect Established Equity:** Established premium brands with strong review moats rely on social proof and clinical positioning rather than promotional depth to sustain conversions - keeping discounts minimal to protect brand desirability. ## **Price Band Analysis - Market Distribution & Satisfaction** ![Amazon SKU price brand](https://www.blog.datahut.co/content/images/2026/08/b3461d_0594b16d70c6434e8fb71a36a6e3e825-mv2.webp) **Key Insights** - **Value Dominant:** The Budget tier ($15 or less) represents the clear baseline of the category, holding an overwhelming 70% of bestseller listings (42 SKUs) at an average price of $10.10\. This is not a temporary trend - it reflects a fundamental consumer preference structure where accessible, high-performing products define the default purchase decision. - **High-Satisfaction Baseline:** Consumer expectations in the budget space are extraordinarily high, with a 4.62-star average - the highest of any price tier. Any brand entering the Budget segment must deliver pristine quality at a low price simply to survive on the digital shelf. There is no performance discount for low pricing in this category. - **Mid-Range Competition:** The Mid-Range tier ($15–$30) holds 28.3% of listings (17 SKUs) at an average price of $19.92 and a 4.52-star rating. This segment is where specialty brands and[ skincare-focused](https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/) players compete, requiring a clear functional differentiation story to justify the price premium over budget alternatives. - **Premium Exclusivity:** The Premium segment ($30+) is the smallest and most exclusive tier, representing just 1.7% of listings (1 SKU) at $45.00 average. Despite the significant price jump, the 4.60-star rating demonstrates that premium products can maintain strong satisfaction scores - but must rely entirely on clinical positioning and brand authority rather than pricing accessibility to drive conversion. ## **Review Concentration Analysis - How Consumer Attention is Distributed** ![Amazon bestsellers review dominance](https://www.blog.datahut.co/content/images/2026/08/b3461d_cb07ca3ceff946778400d2f4d7e4fcf2-mv2.webp) **Key Insights** **Risk Profile: High-Dependency.** The top 10 products command 44.49% of all reviews across the category's 60 listings. This reveals a market where attention is dangerously concentrated - and where a small number of Hero SKUs function as the de facto gatekeepers of the entire category's organic visibility. **Catalog Tail: 55.51%.** While the majority of total engagement volume technically sits outside the top 10, the remaining 50 SKUs must collectively compete for just over half of available attention. This creates a highly fragmented "long tail" where individual listing performance is severely diluted by competition. **Hero Impact: Nearly 1 in 3\.** The single top-performing product - essence's mascara - alone accounts for 11.28% of all category reviews, and the Top 5 products hold 29.04% of total engagement. These top performers don't just provide social proof; they structurally define the category's search landscape. **The "Skyscraper" Effect:** A challenger brand cannot unseat incumbents by launching a wide product catalog. The category functions less like a dense city block and more like a few massive skyscrapers surrounded by low-rise buildings. To compete meaningfully, a challenger must build its own "skyscraper" - a single, high-intensity Hero SKU capable of generating the concentrated review volume required to penetrate the top tier of category attention. ## **Voice of the Customer - Average Review Count Across Brands** ![Amazon bestsellers average reviews](https://www.blog.datahut.co/content/images/2026/08/b3461d_e8070fc7e0fa45609319b331a58ceddd-mv2.webp) **Key Insights** - **Most Validated:** essence (413,046 avg reviews per SKU). The strongest consumer trust signal in the entire category. By focusing on a single, hyper-optimized bestseller listing, essence has built a Review Moat that is structurally impenetrable for organic challengers. Each review compounds the algorithm's confidence, reinforcing placement in top-of-search results. - **High-Engagement Champions:** Mighty Patch (183,491 avg reviews) and MAYBELLINE (183,110 avg reviews) follow as the category's second tier of engagement dominance. Both brands demonstrate that sustained product quality combined with strong brand awareness can sustain massive per-SKU review density - keeping them permanently embedded in the top search positions. - **Trust Anchors:** Aquaphor (139,357 avg reviews) and Paula's Choice (114,378 avg reviews) round out the top tier. These brands leverage clinical credibility and word-of-mouth recommendation patterns to sustain high review velocity, despite occupying a narrower audience than mass-market beauty brands. ## **Average Ratings - Where Brands Stack Up** The ratings chart reveals an impressively tight band of quality across the entire Beauty & Personal Care category. With minimal variation among the top-performing brands, maintaining a high satisfaction baseline is the mandatory requirement for digital shelf survival - and the entry ticket for any brand aiming to compete in top search positions. ![Average rating across Amazon brands](https://www.blog.datahut.co/content/images/2026/08/b3461d_0d4091ab7a7443bcba6f4ce576884119-mv2.webp) **Key Insights** - **The Highest-Rated Leaders:** Q-tips (4.85 stars), CLEAN SKIN CLUB (4.80 stars), Aquaphor (4.80 stars), and Dove (4.80 stars) represent the category's elite satisfaction performers. Their consistently high ratings across multiple use cases and product forms prove that delivering a predictable, premium user experience at scale is achievable across both mass-market and specialty positioning. - **Consistent Quality Across the Board:** eos (4.75 stars) rounds out the top five, demonstrating that even high-volume, wide-distribution brands can maintain elite satisfaction levels when product fundamentals - texture, efficacy, and sensory experience - are consistently delivered. - **The Satisfaction Baseline:** Across all 40 tracked brands, ratings cluster tightly between 4.2 and 4.85 stars. In this environment, a rating below 4.5 stars represents a meaningful competitive disadvantage - not just a statistical difference. Any brand dropping below this threshold faces increased risk of losing organic search visibility and Buy Box position at scale. ## **Revenue Drivers: What Influences Customer Decisions?** Review-text mining of consumer feedback across all 60 top-performing listings reveals the explicit product attributes and functional outcomes that drive purchase decisions, repeat conversions, and long-term brand loyalty in the Beauty & Personal Care category. ![Amazon bestsellers review tags and purchase drivers](https://www.blog.datahut.co/content/images/2026/08/b3461d_aa2e5676a74b45bf843ae6ed1e676f1b-mv2.webp) ![Amazon bestsellers voice of the customer](https://www.blog.datahut.co/content/images/2026/08/b3461d_3f8730d3f3d6435a8fc0e02ddfde382a-mv2.webp) **Key Insights** - **The Rational Anchors - Quality (64,051 mentions) and Effectiveness (54,998 mentions):** These are the two non-negotiable baseline execution elements. Quality is the primary conversion anchor - if a product doesn't perform at a perceptibly high standard, no promotional depth or packaging investment can rescue its retention rate. Effectiveness functions as the performance baseline: shoppers in this category demand clear, visible outcomes. A product that fails to deliver on its primary functional promise cannot sustain a 4.5+ star rating long-term, regardless of brand equity. - **The Core Skincare Driver - Moisturizing (45,544 mentions):** Moisturizing capability stands out as the single most critical functional benefit in the category, far outpacing any other specific ingredient or treatment outcome. This signals that shoppers across all sub-categories — from facial skincare to body care — prioritize immediate, perceptible hydration outcomes as their primary measure of product value. - **Value for Money (38,485 mentions):** Even in a budget-dominant category, perceived value - the ratio of functional outcome to price paid - remains a foundational purchase driver. This finding reinforces the pricing data: shoppers are not simply choosing the lowest-priced option; they are actively evaluating whether the product's efficacy justifies its cost relative to available alternatives. - **The Sensory Component - Fragrance (35,115 mentions):** Fragrance remains a key emotional conversion trigger and brand differentiation lever. A pleasant or distinctive scent profile allows brands to build strong emotional attachment beyond the functional transaction - transforming a routine purchase into a multi-sensory brand experience that drives repeat purchase and loyalty. - **Active Risk Factors - Clumping (9,800 mentions):** For mascara and cosmetic products, clumping is the primary negative sentiment driver. Brands in the eye makeup sub-category must actively address and counter clumping concerns - both in product formulation and in their review response strategy - to protect ratings and conversion rates. **Summary:** The data reveals that the modern Beauty & Personal Care shopper is making decisions at the intersection of functional performance and sensory experience. Quality and Effectiveness set the baseline threshold for consideration; Moisturizing and Fragrance close the conversion. Brands that can simultaneously deliver clinical-grade functional outcomes and a premium sensory experience - regardless of price tier - will consistently outperform category benchmarks on both ratings and review velocity. ### **Final Thoughts - What the Data Really Tells Us** The Amazon Beauty & Personal Care marketplace is defined by two parallel competitive realities. Budget-tier brands like Amazon Basics and The Ordinary dominate shelf space through portfolio breadth and accessible pricing. Meanwhile, efficiency-first brands like essence and Mighty Patch dominate consumer trust through concentrated, hyper-validated Hero SKU performance. Both strategies are viable paths to category leadership - but they require fundamentally different resource allocation, launch sequencing, and organic optimization approaches. For any brand seeking to enter or scale in this category, the data delivers a clear mandate: you cannot win through catalog breadth alone without the review density to back it. A single product built to dominate its sub-category niche - with exceptional quality, strong sensory profile, and a pricing position that delivers tangible value - is worth more than a dozen moderately-reviewed listings. The Review Moat is the most durable competitive advantage on Amazon's digital shelf, and in this category, it is already deeply entrenched. # **Call to Action - Turn These Insights Into Growth** If you want to: - Track competitors' pricing, discounts, and review velocity in near real-time. - Benchmark your brand's share of reviews and organic visibility against category leaders. - Identify pricing gaps, discount opportunities, or emerging Beauty & Personal Care trends before your competitors do. [**Datahut**](https://www.datahut.co/solutions/ecommerce-web-scraping?ref=blog.datahut.co) **can deliver that data.** Get a free 10-minute data audit to see how your catalog stacks up. Talk to a Datahut Data Strategist to explore pricing intelligence, review analytics, and custom data pipelines tailored to the Beauty & Personal Care category. ![Contact Datahut](https://www.blog.datahut.co/content/images/2026/08/b3461d_536d8d5993604205bbdb79747b6ae3d6-mv2-1.webp) # **Frequently Asked Questions (FAQs)** **1\. Q: Is the market too crowded for new brands?** A: It is highly concentrated at the top. With 44.49% of reviews held by just 10 products, new entrants must invest in sustained paid visibility and aggressive review acquisition strategies to break through the existing Review Moat. Success is possible, but it requires a Hero SKU approach — not a broad catalog launch. **2\. Q: Do premium products get better reviews?** A: No. The category data shows that the Budget tier ($10.10 avg price) actually achieves the highest average rating at 4.62 stars, slightly outperforming both Mid-Range (4.52) and Premium (4.60) tiers. Premium pricing does not guarantee better customer satisfaction outcomes in this category. **3\. Q: How does essence dominate with a single SKU?** A: Through Engagement Efficiency. With 413,046 reviews concentrated on a single listing, essence has created an individual product authority that is algorithmically impenetrable for organic challengers. This level of per-SKU review density far outweighs the diffused presence of brands with multiple moderate-performing listings. **4\. Q: What is the pricing "sweet spot"?** A: The Budget tier ($15 or less) with an average price of $10.10 is the dominant commercial zone, holding 70% of bestseller listings. For mass-market positioning, competing in this tier with a differentiated product and strong review velocity is the clearest path to organic search dominance. **5\. Q: What makes customers click "buy" in this category?** A: Beauty & Personal Care is fundamentally a quality and efficacy purchase backed by sensory confirmation. Quality (64,051 mentions) and Effectiveness (54,998 mentions) are the primary rational triggers, while Moisturizing outcomes and Fragrance profile close the emotional conversion loop — creating the "Add to Cart" impulse that drives the highest review velocity in the category. ### Product Matching Using TF-IDF and Cosine Similarity: A Beginner’s Guide for E-commerce Data URL: https://www.blog.datahut.co/post/product-matching-for-e-commerce-data/ Last updated: 2026-09-07T09:32:00.000Z In the world of[ ](https://business.adobe.com/blog/basics/ecommerce-definition?ref=blog.datahut.co)[e-commerce](https://business.adobe.com/blog/basics/ecommerce-definition?ref=blog.datahut.co) product matching, thousands of similar items appear under different names, prices, and brands across online shopping platforms. This creates a challenge when trying to match product listings and identify whether two items refer to the same product. For example, the same air conditioner model may appear on Amazon as “LG AI Convertible 1 Ton 5 Star Split Inverter AC.” It may appear on Flipkart as “LG Super Convertible 1 Ton 5 Star Split Dual Inverter AC.” Although both listings describe the same item, their titles are written differently, which makes product title matching for e-commerce harder with basic text comparison. To solve this, one of the simplest and most powerful methods is[ ](https://www.blog.datahut.co/post/assisted-product-matching-for-ecommerce-using-python/)[product matching](https://www.blog.datahut.co/post/assisted-product-matching-for-ecommerce-using-python/) using TF-IDF and cosine similarity. Instead of checking every word by hand, cosine similarity measures how closely two product names relate based on their meaning. It looks at the angle between them in multi-dimensional space, not their exact wording. Here, TF-IDF (Term Frequency–Inverse Document Frequency) converts product titles and brand names into numerical form so a machine can understand and compare them. After converting text into vectors, cosine similarity calculates how similar the texts are. It does this by checking the angle between the vectors. This helps find matching products across platforms like Amazon and Flipkart. This method is the main part of a real [ ](https://www.shopify.com/blog/what-is-ecommerce?ref=blog.datahut.co)[e-commerce](https://www.shopify.com/blog/what-is-ecommerce?ref=blog.datahut.co) product matching system and is also widely used in E-commerce Text Classification. It helps remove duplicate products. It also improves product discovery. It supports competitor analysis and price comparison. It is very useful to compare product prices, discounts, and reviews automatically. This helps users make better choices when picking items. Product matching also helps compare prices and discounts. It is good for choosing cheap and high-quality branded products. The following sections show a simple and beginner-friendly way to build automatic product matching for e-commerce and online shopping. We use[ scraped product data](https://www.blog.datahut.co/post/how-to-scrape-amazon-and-other-large-e-commerce-websites-at-a-large-scale/) like air conditioners and coolers from Amazon and Flipkart online shopping site. The goal is to find duplicate product listings. We match similar products across websites. We store the matched results in a database for further analysis. Each step works like a puzzle — starting from cleaning text, converting it into numbers (vectorization), applying similarity scores to the final matched output. ## What Is Cosine Similarity? A Beginner's Guide to Product Matching for E-commerce ![What Is Cosine Similarity? A Beginner's Guide to Product Matching for E-commerce](https://www.blog.datahut.co/content/images/2026/08/ChatGPT-Image-Aug-3--2026--04_58_17-PM.webp) At the heart of any product-matching system lies the ability to measure how similar two pieces of text are. [Cosine Similarity](https://memgraph.com/blog/cosine-similarity-python-scikit-learn?utm%5Fsource=chatgpt.com) is a mathematical way to find how close two text strings are. It does not compare words directly. Instead, it looks at their direction in a multi-dimensional space. When product titles are represented as numerical vectors, cosine similarity checks the angle between those vectors. If the angle is small, it means both titles point in nearly the same direction and are therefore very similar. If the angle is large, the texts differ significantly. The formula behind cosine similarity is straightforward and elegant. It measures the cosine of the angle between two vectors, expressed as: **(A · B)** **Cosine Similarity = \_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_** **(‖A‖ × ‖B‖)** Here, **A** and **B** represent two text vectors. The numerator, **A · B**, calculates how much the two vectors overlap — in other words, how similar the words or features are between the two texts. The denominator, **‖A‖ × ‖B‖**, normalizes these values based on the length (or magnitude) of each vector. This ensures that the comparison remains fair, even if one product title is longer than the other. The final cosine similarity score is always between 0 and 1\. Scores near 1 mean the titles are very similar. Scores near 0 mean the texts are quite different. To understand this intuitively, imagine two arrows drawn from the center of a circle. If both arrows point in the same direction, the angle between them is zero, and the cosine value becomes one—indicating perfect similarity. If they point in completely opposite directions, the cosine value becomes zero, meaning there is no similarity at all. What makes cosine similarity particularly effective for text is that it focuses on orientation rather than magnitude. In other words, the number of words in a title does not matter as much as the type and order of words used. This is especially useful when comparing product titles that may have different lengths but convey the same meaning. For instance, *“*[Samsung](https://www.samsung.com/in/offer/ai-just-for-you/?ref=blog.datahut.co) *1.5 Ton Split AC”* and *“Samsung Split Air Conditioner 1.5 Ton”* are written differently but represent the same product. Cosine similarity measures this relationship accurately. ## **Why We Use TF-IDF to Compare Product Titles** ![Why We Use TF-IDF to Compare Product Titles](https://www.blog.datahut.co/content/images/2026/08/ChatGPT-Image-Aug-3--2026--04_53_29-PM.webp) Every text-based matching system begins with a fundamental challenge: computers cannot understand words the way humans do. For a machine, the words *“Samsung Split AC”* and *“Split Air Conditioner by Samsung”* (product titles) are simply strings of characters. To compare such text meaningfully, these strings must first be transformed into a numerical form that preserves their meaning and context. This is where [TF-IDF Vectorization](https://www.jeremyjordan.me/identifying-related-bodies-of-text-using-tf-idf-vectorization/?ref=blog.datahut.co) plays an important role. TF-IDF stands for Term Frequency–Inverse Document Frequency. It is a statistical method that changes text into numbers based on how important each word is in a group of documents. The concept is intuitive. Words that appear often in a product title get higher weight through term frequency. Very common words that appear in almost every title—like “with,” “for,” or “and”—get lower weight through inverse document frequency. The result is a balanced representation where each word’s importance reflects both its presence and its uniqueness. When applied to product titles, TF-IDF helps highlight the distinctive terms that actually describe a product, such as its brand, model, or key features. For instance, in the product title *“*[*LG* ](https://www.lg.com/in/?srsltid=AfmBOoqK2T9YsTQqOLH2Do3PIX0u2EhSE%5FLq-%5FACFs%5FjbGfyc3gGFDHD&ref=blog.datahut.co)*1 Ton Inverter Split AC”*, words like “LG,” “Inverter,” and “Split” carry strong informational value, while filler words are considered less significant. The TF-IDF process turns the text into a vector. This vector is a list of numbers that represents the product title in a form that can be compared mathematically. Once the product titles from both datasets — [Amazon](https://www.amazon.in/?ref=blog.datahut.co) and [Flipkart](https://www.flipkart.com/?ref=blog.datahut.co)—are transformed into TF-IDF vectors, Cosine Similarity comes into action. Rather than checking whether two titles contain the exact same words, cosine similarity evaluates the angle between their corresponding TF-IDF vectors. A smaller angle, or a higher cosine value, indicates that the titles share similar patterns and vocabulary, even if the wording differs. TF-IDF and cosine similarity form a strong pair. TF-IDF changes text into numbers. Cosine similarity measures how close those numbers are. Together, they enable efficient and accurate matching of product data across large e-commerce catalogs. The process captures not just exact word overlap. It also captures how close product descriptions are in context. This helps with tasks like E-commerce Text Classification and accurate product matching. This allows the system to find true matches hidden behind different phrasing. In product-matching workflows, this approach ensures that models focus on the essence of a product rather than superficial text differences. It makes the system easy to grow, understand, and change for different product types. This works whether you compare air conditioners, smartphones, or kitchen appliances. TF-IDF uses statistical weighting. Cosine similarity measures how similar two things are using geometry. They work together.Together, they form the main analysis method in modern text-based matching systems. ## Real-World Product Matching Example: Matching Godrej AC Listings Consider two product entries taken from Amazon and Flipkart. Both describe the same Godrej air conditioner model but the titles use slightly different wording and extra phrases. The cleaned titles used for comparison are shown here: - Amazon cleaned title: [godrej](https://www.godrejenterprises.com/?ref=blog.datahut.co) 1 ton 5 star 5 in 1 convertible cooling inverter split ac copper i sense technology 2023 model ac 1t ei 12tinv5r32 gwa split ac 1t ei 12tinv5r32 rwb split white - Flipkart cleaned title: godrej 5 in 1 convertible cooling 2023 model 1 ton 5 star split inverter i sense technology with blue fin anti corrosive coating ac white gold ac 1t ei 12tinv5r32 gwa split ac 1t ei 12tinv5r32 rwb split copper condenser The calculation follows the usual steps. First, it breaks text into tokens and creates TF-IDF vectors. Then, it calculates the dot product and Euclidean norms. Finally, it uses the cosine formula. Scikit-learn’s default settings (including smooth\_idf=True and norm='l2') are used so vectors are L2-normalized. ### ***Step 1 — TF-IDF vectors and vocabulary*** The TF-IDF vectorizer builds a vocabulary of 28 tokens from the two titles. Each title is converted into a [TF-IDF vector](https://okan.cloud/posts/2022-01-16-text-vectorization-using-python-tf-idf/?utm%5Fsource=chatgpt.com); because norm='l2', both vectors have Euclidean norm 1.0\. A selection of the TF-IDF values and their per-token contributions to the dot product follows (values shown are exact as produced by scikit-learn): | Token | Amazon | Flipkart | Contribution | | ---------- | ------------------- | ------------------- | ------------------ | | 12tinv5r32 | 0.29814239699997197 | 0.25648898295269357 | 0.0764702401816010 | | 1t | 0.29814239699997197 | 0.25648898295269357 | 0.0764702401816010 | | ac | 0.4472135954999579 | 0.3847334744290404 | 0.1720580404086023 | | split | 0.4472135954999579 | 0.3847334744290404 | 0.1720580404086023 | | 2023 | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | | gwa | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | | inverter | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | | model | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | | ton | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | | white | 0.14907119849998599 | 0.12824449147634678 | 0.0191175600454003 | - Several additional shared tokens (e.g., convertible, cooling, sense, split, star, rwb, etc.) each contribute ≈ 0.01911756 or 0.0 if absent in one title (Flipkart contains some extra tokens such as blue, anti, coating, condenser, gold, with, each present only in Flipkart and thus contributing 0 to the dot product ### ***Step 2 — Dot product (numerator of cosine formula)*** The dot product sums all per-token contributions. Using the TF-IDF numbers above, the dot product equals: ``` A⋅B  =  0.8602902020430111 ``` ### ***Step 3 — Vector norms (denominator parts)*** Because scikit-learn normalized TF-IDF vectors with L2 norm, each vector has: ```       ∥A∥=1.0       ∥B∥=1.0 ``` ### ***Step 4 — Cosine similarity (title)*** Apply the cosine formula: ```   Cosine_title                      = A⋅B/∥A∥×∥B∥                      = 0.8602902020430111/(1.0×1.0)                      = 0.8602902020430111 ``` Rounded for presentation: 0.86029 (≈ 0.86). This value shows a strong similarity between the two titles after applying TF-IDF (term frequency-inverse document frequency) weighting and normalization. ### ### ***Step 5 — Brand similarity*** Both product entries use the same cleaned brand token godrej. Using TF-IDF on the brand strings (a trivial two-document corpus where both items are identical) yields identical normalized vectors, so: ``` cosine_brand=1.0 ``` ### ***Step 6 — Combined similarity with weights*** The script combines title and brand similarity using the configured weights title\_weight = 0.8 and brand\_weight = 0.2\. The weighted combination is computed as: ``` Combined_similarity = 0.8×cosine_title+0.2×cosine_brand  ``` "Combined\_similarity equals 0.8 times cosine\_title plus 0.2 times cosine\_brand" Substituting the numbers: ``` Combined_similarity =0.8×0.8602902020430111+0.2×1.0 =0.6882321616344089+0.2 = 0.8882321616344089 Rounded for reporting: 0.888 (≈ 0.89). ``` ### **Interpretation and practical note** The title cosine (≈ 0.86) shows a very close textual match between the two cleaned product titles. The brand cosine (1.0) confirms brand agreement. After applying the chosen weights, the final combined similarity is about 0.888\. This number is higher than a typical matching threshold like 0.7\. So, the pair is a strong candidate for being the same product. This example shows how small wording differences and extra descriptive phrases do not stop a successful match. This happens when TF–IDF vectorization and cosine similarity are used. The same approach scales to large catalogs: text is cleaned, converted to TF–IDF vectors, and compared with [cosine similarity](https://www.tigerdata.com/learn/understanding-cosine-similarity?utm%5Fsource=chatgpt.com) to identify likely matches. Later improvements can make the results more accurate. These include adding brand exact-match checks, giving brand tokens more weight, or using transformer embedding for better meaning capture. ## Step-by-Step Python Tutorial for Product Matching Between Amazon and Flipkart **Essential Python Libraries Used** ```python import pandas as pd import sqlite3 import re import logging from datetime import datetime from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity from tqdm import tqdm import os import sys ``` The journey starts with Pandas library, a powerful data-handling library that simplifies reading, cleaning, and transforming datasets. Here, it manages large CSV files containing product information from Amazon and Flipkart. SQLite3 follows, acting as a lightweight database that helps store matched results efficiently for future use or reporting. Before any meaningful comparison can happen, product text needs to be cleaned and standardized. The regular expression module, re, helps clean text. It removes extra characters, symbols, or spaces. This makes product titles uniform and easier to compare. Accurate tracking of the script’s progress is equally important, especially when handling large datasets. Logging is included to record the entire process, from start to finish, along with any errors or matches found. To ensure proper time-based tracking, the datetime module helps timestamp events throughout the execution. Scikit-learn offers two strong tools to find product similarity. They are TfidfVectorizer and Cosine\_similarity. The TF-IDF (term frequency-inverse document frequency) vectorizer changes text into number vectors based on word importance. Cosine similarity measures how close two vectors are, so it shows how similar two product titles are. Together, they form the core of the product matching logic. Since product comparison can be time-consuming, especially with thousands of rows, tqdm introduces progress bars that make it easier to monitor the script’s progress visually. We include os and sys for system tasks. These include file handling, managing paths, and handling unexpected interruptions smoothly. Together, these imports build a strong base for the product matching process. They combine data handling, text cleaning, similarity calculation, and result management into a clear and efficient system. ## ### **Setting Up Configuration Paths** ```python # CONFIG amazon_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/amazon-ac-v8-cleaned.csv" flipkart_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/flipkart-AC-full-v1-cleaned.csv" db_path = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/matched_products_final.db" log_file = "/home/anusha/Desktop/DATAHUT/Flip_amaz_cosine/DATA/product_match_log_final.txt" """ Configuration setup for the product matching workflow. """ ``` In this stage of the product-matching pipeline, configuration paths are defined for the main files and databases that form the core of the workflow. Setting up these configurations at the start makes the code run smoothly. It also prevents confusion later. This setup makes the code easier to manage and change when needed. The first two variables, amazon\_file and flipkart\_file, point to the cleaned datasets collected from Amazon and Flipkart. These CSV files contain structured product data such as names, specifications, and other details that will later be compared to find similar or matching products. The paths ensure that the code knows exactly where to locate these datasets on the system. Keeping the data in a cleaned and organized format at this stage is essential because it directly affects the quality and accuracy of the matching results. The db\_path variable defines the location of the SQLite database where all final matched product results will be stored. Using a database for storing results rather than saving them as flat files provides more control and flexibility. It allows for efficient querying, filtering, and retrieving specific matches without re-running the entire process. This becomes particularly valuable when working with large volumes of e-commerce data. Lastly, the log\_file variable specifies the location where all logs will be recorded. Logging plays a vital role in maintaining transparency throughout the process. Each event—whether it is a successful match, an error, or a step in progress—is automatically documented in this log file. This record helps in tracking the execution flow and debugging any unexpected issues. Together, these configuration settings form the backbone of the project. They provide structure and clarity. This makes sure that every later step works well and consistently. These steps include data cleaning, text comparison, and similarity calculation. It also removes the need for manual changes each time the code runs. ### **Setting the Right Threshold and Weights for Accurate Product Comparison** ```python # cosine similarity threshold and weights threshold = 0.7 title_weight = 0.8 brand_weight = 0.2 """ Defines the cosine similarity threshold and attribute weights for product matching. """ ``` After defining the core configuration paths, the next step in the product matching workflow involves setting up the similarity parameters. These parameters determine how closely two products need to resemble each other to be considered a match. Cosine similarity measures how similar product titles and product descriptions are. We assign weights to specific attributes. These weights control how much each attribute affects the final matching score. The threshold value, set at *0.7*, acts as a decision boundary. It represents the minimum level of similarity required between two product descriptions for them to qualify as a potential match. For instance, if the cosine similarity between two titles is 0.85, the system recognizes them as highly similar; however, if the score falls below 0.7, the match is discarded. Choosing the right threshold is crucial—it must be balanced enough to avoid both false positives (mismatches) and false negatives (missed matches). The title\_weight and brand\_weight variables help fine-tune the importance of different product attributes. In this setup, the product title carries a weight of *0.8*, while the brand description contributes *0.2* to the overall similarity score. This weighting reflects a realistic approach to e-commerce product matching, where titles often provide more detailed and descriptive information than brand names. For example, two air conditioners from the same brand can have very different specifications. But their titles mention capacity, model, or features. These product titles give a better way to compare them. By defining these parameters early in the process, the matching system is guided to make more accurate, data-driven decisions. The process keeps the comparison consistent and objective. It matches real-world product variations. This leads to reliable and meaningful matching results across large datasets. ### **Creating a Reliable Logging System for Transparent Data Workflows** ```python # LOGGING SETUP if os.path.exists(log_file): os.remove(log_file) logging.basicConfig( level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s", handlers=[logging.FileHandler(log_file), logging.StreamHandler(sys.stdout)] ) def log(msg, level="info"): getattr(logging, level)(msg) log(" Product Matching Script Started") """ Initializes and configures the logging system for the product matching script. """ ``` Once the configurations and similarity parameters are set, the next essential component is the logging setup. Logging acts as the project’s internal diary—it records every significant event, tracks progress, and helps in understanding the system’s behavior at each stage. In large data processing tasks, especially when comparing thousands of products, keeping clear and organized logs is essential. The first step ensures that any existing log file is removed before creating a new one. This keeps the log clean for every fresh run, avoiding clutter from previous executions. The logging.basicConfig() function then initializes the logging configuration with a few key parameters. The logging level is set to *INFO*, meaning the system will record all informative messages, warnings, and errors. The format specifies how each message should appear in the log, including the timestamp, log level, and the actual message text. The handlers send the output to two places at once. One goes to the log file. The other goes to the console. This lets you track in real time and keep a written record. To simplify message logging throughout the script, a helper function named log() is defined. This function dynamically calls the appropriate logging method, such as *info*, *warning*, or *error*, depending on the type of message passed to it. The first entry in the log, “ Product Matching Script Started,” marks the beginning of the process. It signals that the system has initialized successfully and is ready to begin the product matching workflow. This setup ensures that every step, from reading data to calculating similarity scores, is carefully documented. In case of unexpected errors or inconsistencies, these logs serve as a reliable trail for identifying the root cause. This organized approach improves transparency. It also helps with efficient debugging and process improvement in large data projects. ### **Cleaning and Preparing Text Data for Accurate Comparison** ### ```python # CLEANING FUNCTION def clean_text(text): if not isinstance(text, str): return "" text = text.lower() text = re.sub(r"[^a-z0-9\s]", " ", text) text = re.sub(r"\s+", " ", text).strip() return text """ Cleans and standardizes raw product text for consistent comparison. ``` In any text-based data processing task, one of the most crucial steps before analysis begins is data cleaning. Raw product data from different e-commerce platforms often has problems. It includes extra symbols, special characters, or unwanted spaces. These problems can make comparisons inaccurate. To ensure reliable results, all text must be standardized and simplified into a clean, comparable format. This is where the cleaning function plays a central role. The function named clean\_text() is designed to perform this essential preprocessing task. It begins by checking whether the input is a string. This validation step prevents errors when unexpected data types—such as numbers or missing values—appear in the dataset. If the value isn’t a string, the function simply returns an empty string, ensuring the process continues smoothly without interruptions. Once the input is confirmed as text, the function converts all characters to lowercase. This helps maintain uniformity, so words like “Samsung,” “samsung,” and “SAMSUNG” are treated as identical during similarity calculations. The next step uses regular expressions to remove unwanted characters—anything that isn’t a letter, number, or space. This eliminates punctuation marks and symbols that do not contribute meaningfully to product matching. After that, multiple spaces are replaced with a single space, and any leading or trailing spaces are removed. The end result is a clean, well-formatted string that focuses only on the meaningful content. For example, a messy product titles like *“LG 1.5-Ton Inverter A/C!!!”* becomes *“lg 1.5 ton inverter a c”* after cleaning. By applying this cleaning process consistently to all product titles and descriptions, the matching algorithm gains a clear and unbiased view of the data. It ensures that every comparison between two product titles is based purely on their relevant content, not on noise or formatting differences. This step lays the foundation for accurate similarity computation and is one of the most important parts of preparing real-world data for analysis. ### **Measuring Similarity with Jaccard Distance** ```python # JACCARD DISTANCE FUNCTION def jaccard_distance(str1, str2): set1 = set(str1.split()) set2 = set(str2.split()) if not set1 and not set2: return 0.0 intersection = set1.intersection(set2) union = set1.union(set2) return 1 - len(intersection) / len(union) """ Calculate the Jaccard distance between two strings. """ ``` ![Measuring Similarity with Jaccard Distance](https://www.blog.datahut.co/content/images/2026/08/ChatGPT-Image-Aug-3--2026--04_47_37-PM.webp) In data analysis and product matching, comparing how similar two pieces of text are is a common challenge. One practical method for this is the [Jaccard distance](https://mayurdhvajsinhjadeja.medium.com/jaccard-similarity-34e2c15fb524?ref=blog.datahut.co), which quantifies the difference between two sets of words. At its core, the Jaccard distance focuses on the overlap between the words in two strings, providing a simple yet powerful way to capture similarity. The function begins by splitting each string into a set of words. This conversion from text to sets is crucial because sets automatically remove duplicate words, which ensures that the comparison considers only unique terms. Once the sets are created, the intersection of the two sets is calculated. This intersection represents the words that appear in both strings, highlighting the shared elements between the two texts. In contrast, the union of the sets includes all unique words present in either string, capturing the total scope of vocabulary used. The [Jaccard](https://www.tandfonline.com/doi/full/10.1080/24751839.2021.1893496?ref=blog.datahut.co#d1e251) distance itself is computed as one minus the ratio of the intersection size to the union size. If two strings are exactly the same, their intersection equals the union, resulting in a distance of zero, which indicates perfect similarity. Conversely, if there are no shared words, the intersection is empty, and the distance reaches one, signaling complete dissimilarity. This simple calculation makes the Jaccard similarity easy to understand. It is useful for tasks like product matching, document comparison, or grouping similar items. This method works well when the focus is on whether terms are present or not. It gives a different view than methods like TF-IDF or cosine similarity. Using [Jaccard distance](https://datascience.stackexchange.com/questions/5121/applications-and-differences-for-jaccard-similarity-and-cosine-similarity?ref=blog.datahut.co) with other similarity measures helps us better understand text relationships. This ensures accurate and meaningful matches between datasets. The function itself is concise and efficient, handling edge cases where both strings might be empty by returning a distance of zero. This ensures robustness in real-world scenarios, where missing or incomplete data is often encountered. Jaccard similarity gives a reliable and clear way to compare text. It helps data systems find connections and patterns clearly. ### **Bringing Amazon and Flipkart Data into the Workflow** ```python # LOAD DATA try: amazon_df = pd.read_csv(amazon_file) flipkart_df = pd.read_csv(flipkart_file) log(f" Loaded Amazon: {len(amazon_df)} rows, Flipkart: {len(flipkart_df)} rows") except Exception as e: log(f" Error loading CSVs: {e}", "error") sys.exit(1) """ Loads cleaned product datasets from Amazon and Flipkart using pandas.""" ``` After defining the cleaning function, the next step in the workflow focuses on loading the datasets that will be used for product comparison. This stage acts as the entry point for the actual data that drives the entire matching process. Clean, structured data from multiple sources must be read into the program accurately before any analysis can begin. Even a small error at this point can affect every step that follows, making reliable data loading one of the most important foundations of the pipeline. In this section, the script attempts to read two CSV files—one from Amazon and the other from Flipkart—using the [pandas.read](http://pandas.read/?ref=blog.datahut.co)\_csv() function. The try-except structure is used to handle this process safely. Inside the try block, the code loads both datasets into memory as DataFrames. DataFrames have a table-like format that makes it easy to work with and analyze the data. Once the data is successfully loaded, a log message records the number of rows in each dataset. This immediate feedback confirms that the files were read correctly and helps verify that the expected amount of data has been imported. The inclusion of an exception handling block ensures the system remains stable, even if something goes wrong. For instance, if one of the file paths is incorrect or a file is missing, the program logs a clear error message and exits gracefully instead of crashing. This approach adds robustness to the script, making it capable of handling real-world scenarios where data inconsistencies or missing files are common. By the end of this step, the raw datasets from both e-commerce platforms are securely loaded and ready for the next stages of processing. This careful and organized loading process makes sure the matching system starts with reliable and accessible data. This prepares the system for accurate comparison and analysis in the next steps. ### **From Raw to Ready: Cleaning and Validating Product Data** ```python # CLEAN & VALIDATE amazon_df.columns = amazon_df.columns.str.lower() flipkart_df.columns = flipkart_df.columns.str.lower() required_cols = ["product_title", "brand"] for col in required_cols: if col not in amazon_df.columns or col not in flipkart_df.columns: log(f" Missing required column '{col}'", "error") sys.exit(1) log("🧹 Cleaning text columns...") amazon_df["clean_title"] = amazon_df["product_title"].apply(clean_text) amazon_df["clean_brand"] = amazon_df["brand"].apply(clean_text) flipkart_df["clean_title"] = flipkart_df["product_title"].apply(clean_text) flipkart_df["clean_brand"] = flipkart_df["brand"].apply(clean_text) """ Performs initial data validation and text cleaning on the loaded Amazon and Flipkart datasets. """ ``` Once the data has been successfully loaded into memory, the next essential phase is data validation and cleaning. This step makes sure that the datasets from Amazon and Flipkart have the same structure. It also ensures they have the needed information for accurate product matching. Before calculating similarity, we must standardize and check the data. Small differences in column names or missing fields can cause problems. The first step converts all column names to lowercase. This simple change removes differences caused by naming styles. For example, one dataset might use “Product\_Title” while another uses “product\_title.” By standardizing column names, the script avoids confusion and ensures smooth access to each field during processing. Next, the code checks for the presence of key columns—product\_title and brand—in both datasets. These two attributes are fundamental to the matching process. The product title provides descriptive details about the item, while the brand description helps confirm its identity. If either of these columns is missing from any dataset, the script immediately logs an error and halts execution. This safeguard prevents incomplete data from entering later stages, where it could lead to inaccurate results or unexpected failures. Once validation is complete, the focus shifts to text cleaning. Each product title and brand name is passed through the clean\_text() function defined earlier. This function removes unwanted characters, converts text to lowercase, and ensures a consistent structure across both datasets. The cleaned versions of these fields are stored in new columns—clean\_title and clean\_brand—which serve as the standardized input for all future comparisons. Through this process, both datasets are transformed into a clean, uniform, and validated state. Each entry is now prepared for reliable analysis, free from inconsistencies that could distort similarity scores. Careful attention to data quality is the foundation of any successful product matching process. It makes sure every comparison later is based on correct and meaningful information. ### **Converting Text Data into Vectors Using TF-IDF** ```python # TF-IDF FITTING log("📊 Fitting TF-IDF on combined corpus...") vectorizer = TfidfVectorizer().fit( amazon_df["clean_title"].tolist() + flipkart_df["clean_title"].tolist() ) log(" TF-IDF ready") ``` After cleaning and checking the datasets, the next important step changes the text data into a form the computer can understand and compare well. Human readers can easily see that two titles like “LG 1.5 Ton 5 Star Inverter AC” and “LG Inverter Split AC 1.5 Ton 5 Star” mean the same product. But a machine only sees both as strings of text. To enable meaningful comparisons, this textual information needs to be converted into numerical representations — and that’s where TF-IDF comes into play. TF-IDF, short for *Term Frequency–Inverse Document Frequency*, is a method that converts text into a set of numerical features based on how important each word is within a collection of texts. In simpler terms, it identifies which words carry the most value in describing a product. Common words like “air” or “ac” might appear frequently across many titles, so their importance is lower. However, more specific terms such as “inverter” or “dual cool” appear less often and thus carry greater weight when measuring similarity between two product names. In this section, the code creates a TF-IDF (term frequency-inverse document frequency) vectorizer and fits it on a combined set of cleaned product titles from both Amazon and Flipkart. By joining text data from both platforms into one group, the model learns one vocabulary. This makes sure the same words are shown the same way in both datasets. This shared representation becomes the foundation for calculating similarity scores later in the process. Once the fitting is complete, the log confirms that the TF-IDF model is ready. At this stage, every product title is prepared to be converted into its corresponding vector — a numerical form that captures the essence of the text. This change connects human language to machine calculation. Then, cosine similarity measures how closely two products relate based on their vector forms. ### **Product Matching Using TF-IDF (term frequency-inverse document frequency) and Cosine Similarity** ```python # MATCHING LOOP matches = [] log("🔍 Starting product matching...") for i, a_row in tqdm(amazon_df.iterrows(), total=len(amazon_df), desc="Matching"): try: # TF-IDF vectors for titles a_title_vec = vectorizer.transform([a_row["clean_title"]]) f_title_vecs = vectorizer.transform(flipkart_df["clean_title"]) title_sim = cosine_similarity(a_title_vec, f_title_vecs).flatten() # TF-IDF vectors for brands brand_vectorizer = TfidfVectorizer().fit( [a_row["clean_brand"]] + flipkart_df["clean_brand"].tolist() ) a_brand_vec = brand_vectorizer.transform([a_row["clean_brand"]]) f_brand_vecs = brand_vectorizer.transform(flipkart_df["clean_brand"]) brand_sim = cosine_similarity(a_brand_vec, f_brand_vecs).flatten() # Combined cosine similarity combined_sim = (title_sim * title_weight) + (brand_sim * brand_weight) best_idx = combined_sim.argmax() best_score = combined_sim[best_idx] if best_score >= threshold: # Jaccard distance for the matched pair combined_text_a = f"{a_row['clean_title']} {a_row['clean_brand']}" combined_text_f = f"{flipkart_df.iloc[best_idx]['clean_title']} {flipkart_df.iloc[best_idx]['clean_brand']}" jdl_score = round(jaccard_distance(combined_text_a, combined_text_f), 3) matches.append({ "amazon_index": i, "flipkart_index": best_idx, "cosine_similarity": round(best_score, 3), "jaccard_distance": jdl_score }) log(f" Match (Cosine {best_score:.2f} | JDL {jdl_score:.2f}): " f"{a_row['product_title'][:60]} ↔ {flipkart_df.iloc[best_idx]['product_title'][:60]}") except Exception as e: log(f" Error on Amazon row {i}: {e}", "warning") """ This loop performs product matching between the Amazon and Flipkart datasets by comparing both the product titles and brands using TF-IDF vectorization and similarity metrics. """ ``` When comparing products from different platforms, the main challenge is to find which items match between datasets. This is achieved by evaluating the similarity between product titles and brands, which are often the most descriptive identifiers. The matching loop does this task efficiently. It uses math methods to measure similarity. Each product from the Amazon dataset is processed one by one. The product title is first changed into numbers using a method called TF-IDF. TF-IDF stands for Term Frequency-Inverse Document Frequency. This technique converts text into vectors that capture the importance of each word relative to the entire dataset. A similar transformation is applied to all product titles from Flipkart, allowing a direct comparison between the two platforms. We use cosine similarity to calculate how similar two vectors are. This measures how closely two vectors point in the same direction. A high cosine similarity indicates that the product titles are closely related, which is crucial for accurate matching. Brands are treated in a similar manner. A separate TF-IDF vectorizer is fitted specifically for the brand names and brand descriptions of the current Amazon product and all Flipkart products. Once again, cosine similarity is used to evaluate how closely the brand names match. By combining the similarities from both the title and brand, a weighted score is computed. The weighting allows the title to have a stronger influence while still considering the contribution of the brand. The highest combined similarity score identifies the most likely matching product on Flipkart for the current Amazon product. In addition to cosine similarity, Jaccard distance is calculated for the matched pair. This metric provides a different perspective by evaluating the overlap of unique words between the combined title and brand texts of the two products. Cosine similarity uses vector space. Jaccard distance measures how much vocabulary is shared. Jaccard distance helps find partial matches or small text differences. The computed Jaccard distance is rounded and stored alongside the cosine similarity, providing a dual-metric assessment of how closely the products align. Each successful match is logged, showing both the cosine similarity and Jaccard distance alongside a snippet of the product titles. This not only ensures transparency in the matching process but also helps in debugging and verifying the results. The script catches and records any errors during processing, such as missing data or unexpected text formats. It does this without stopping the overall execution. This makes the script strong for large datasets. This loop forms the core of the product matching system, combining advanced text representation techniques with practical similarity measures. The system compares every product and records both cosine similarity and Jaccard distance. This method helps match products across platforms reliably. It also prepares the data for further analysis or use in e-commerce workflows. ### **Storing Amazon–Flipkart Matched Results Safely Using SQLite** ```python # DATABASE SAVE log(" Saving results to database...") conn = sqlite3.connect(db_path) if matches: matches_df = pd.DataFrame(matches) amazon_matched = [] flipkart_matched = [] combined_rows = [] all_columns = list(set(amazon_df.columns.tolist() + flipkart_df.columns.tolist())) key_columns = [ "product_url", "product_title", "company", "brand", "mrp", "sales_price", "similarity", "jaccard_distance" ] ordered_cols = key_columns + [c for c in all_columns if c not in key_columns] for _, match in matches_df.iterrows(): a_row = amazon_df.loc[match["amazon_index"]].to_dict() f_row = flipkart_df.loc[match["flipkart_index"]].to_dict() a_row["similarity"] = match["cosine_similarity"] a_row["jaccard_distance"] = match["jaccard_distance"] a_row["company"] = "Amazon" f_row["similarity"] = match["cosine_similarity"] f_row["jaccard_distance"] = match["jaccard_distance"] f_row["company"] = "Flipkart" # Fill missing columns with None for col in ordered_cols: a_row.setdefault(col, None) f_row.setdefault(col, None) amazon_matched.append(a_row) flipkart_matched.append(f_row) combined_rows.extend([a_row, f_row]) """ DATABASE SAVE SECTION This section saves the matched product data between Amazon and Flipkart into an SQLite database. """ ``` When handling product matching between two major e-commerce marketplace, the final step is often saving the results in a structured and accessible way. After we calculate similarity measures like cosine similarity and Jaccard distance, we organize the data into clear tables. This is important for future analysis or reports. The process begins by establishing a connection to a SQLite database, which serves as a lightweight and reliable storage solution for structured data. Once the database connection is ready, the script checks if there are any matched products. Each match contains references to the corresponding entries in both the Amazon and Flipkart datasets, along with the computed similarity scores. The system goes through these matches one by one. It changes the data into dictionaries. This keeps the similarity and Jaccard distance scores with the original product details. Assigning a company label to each entry clearly distinguishes between the two platforms. To maintain consistency and prevent missing values, the script ensures that every column in the final tables is filled. Any missing field is explicitly set to None, guaranteeing that the database structure remains uniform across all entries. We collect matched products from Amazon and Flipkart separately. We also create a combined dataset for full analysis. This method allows flexibility. It lets you see platform-specific matches separately. It also gives a combined view of all product comparisons. The final stage involves writing these datasets to the database. Separate tables are created for Amazon matches, Flipkart matches, and a combined table containing all matched products. We keep a special table for Jaccard distance. This table does not include the cosine similarity column. It gives another way to look at product similarity. This organized method stores the results efficiently. It also makes it easy to search and analyze the data for insights, reports, or further work. Using this method, we store matched product data in a professional and easy-to-access format. This keeps all important details and lets users explore and understand product relationships across platforms in many ways. This setup ensures that any analysis performed on the matched data is accurate, reproducible, and easy to manage. ### **Creating Separate Database Tables for Amazon, Flipkart, and Combined Matches** ```python # Convert to DataFrames amazon_df_out = pd.DataFrame(amazon_matched)[ordered_cols] flipkart_df_out = pd.DataFrame(flipkart_matched)[ordered_cols] combined_df_out = pd.DataFrame(combined_rows)[ordered_cols] # Save tables amazon_df_out.to_sql("amazon_matched_only", conn, if_exists="replace", index=False) flipkart_df_out.to_sql("flipkart_matched_only", conn, if_exists="replace", index=False) combined_df_out.to_sql("matched_combined", conn, if_exists="replace", index=False) # Save Jaccard distance table (exclude cosine similarity) jdl_df_out = combined_df_out.drop(columns=["similarity"]) jdl_df_out.to_sql("matched_jaccard", conn, if_exists="replace", index=False) log(f" amazon_matched_only: {len(amazon_df_out)} rows") log(f" flipkart_matched_only: {len(flipkart_df_out)} rows") log(f" matched_combined: {len(combined_df_out)} rows") log(f" matched_jaccard: {len(jdl_df_out)} rows") else: log(" No matches found above threshold") """ DATAFRAME CONVERSION AND DATABASE SAVE This section converts the matched product data stored in Python lists into Pandas DataFrames and saves them into an SQLite database. """ ``` Once the product matching process is complete, the results are organized and stored in a structured format for future analysis. This step ensures that all matched product information between Amazon and Flipkart is captured in a way that is easy to query, compare, and visualize. The first action involves converting the collected match lists into Pandas DataFrames. Separate DataFrames are created for Amazon products, Flipkart products, and a combined view that brings together all matched rows. This separation lets us analyze each platform alone. It also lets us see the overall matching in one dataset. Saving these DataFrames to an SQLite database provides a reliable, persistent storage mechanism. - **amazon\_matched\_only** – Stores all matched products from Amazon. - **flipkart\_matched\_only** – Stores all matched products from Flipkart. - **matched\_combined** – Contains all matched rows from both platforms. This structured approach ensures that queries can be tailored based on specific requirements, such as platform-specific insights or cross-platform comparisons. In addition to cosine similarity, the Jaccard distance between products is also stored. To focus on this metric, we create a separate table named matched\_jaccard. We make it by removing the cosine similarity column from the combined DataFrame. This separation shows the Jaccard distance clearly. It does not mix with other similarity measures. This makes it easier to understand and use for further analysis. Finally, the script logs the number of rows in each table, providing a quick overview of how many matches were successfully identified and stored. These logs serve both as a checkpoint and a record for auditing purposes. By organizing and storing the matched product data in this systematic way, the dataset becomes a powerful resource for insights, reporting, or downstream processing tasks. ## **Final Step: Safely Closing SQLite Connection and Ending the Product Matching Process** ```python conn.close() log(" Script finished successfully") log(f" Completed at {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}") """ Closes the active database connection and logs the completion of the script. """ ``` After the matched data has been successfully stored in the database, the final step gracefully concludes the entire process. The database connection, which has been open throughout the matching and saving operations, is now closed to ensure that all resources are properly released. This is an important best practice in any data-handling workflow, as keeping connections open unnecessarily can lead to issues such as memory leaks or database locks. The script closes the connection explicitly. This keeps the process efficient and stable. It shows that all operations finished safely. Once the connection is closed, a series of final log messages marks the completion of the entire workflow. These messages not only confirm that the script has run successfully but also record the exact time of completion. This acts as a clear sign for tracking how well the process runs or for fixing problems. This is especially useful when working with large datasets. This closing stage gives the entire process a sense of completion and reliability. Loading and preparing data, calculating similarities, finding matches, and saving results all help build a smooth, automated product matching system. Ending the script with clean closure and clear logs helps the pipeline run well. This makes the process professional, easy to follow, and reliable. These are important parts of good data engineering. ### **Input Data** Before starting the product matching process, we used two cleaned datasets as inputs: - **Amazon Dataset :**[ **amazon-ac.csv**](https://www.dropbox.com/scl/fi/4vat2ut66qx7mz5lwhp41/amazon-ac-v8-cleaned.csv?rlkey=do946d8v1gixb8ri3sav2lsbg&st=gkh3edyq&dl=0&ref=blog.datahut.co) - **Flipkart Dataset :** [ **flipkart-ac.csv**](https://www.dropbox.com/scl/fi/6ujb7ev4zjv09t8givg1x/flipkart-AC-full-v1-cleaned.csv?rlkey=t5nd7w9txelgo4t64wg38nf8w&st=79f5vv1h&dl=0&ref=blog.datahut.co) ### **Output Results** These files contain the matched product pairs along with their computed cosine similarity scores and related details: | **Table Name** | **What it Stores** | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [**Amazon Matched Products**](https://www.dropbox.com/scl/fi/1lkpd7vyvp3xbypqnaqlb/amazon%5Fmatched%5Fcleaned.csv?rlkey=lnrtnkp3oq0pljjwrco3wcasd&st=hvqoowkt&dl=0&ref=blog.datahut.co) | **Includes the Amazon product details that were successfully matched with Flipkart listings based on the cosine similarity threshold** | | [**Flipkart Matched Products**](https://www.dropbox.com/scl/fi/ublp7gn0x0a1s6liy0tv1/flipkart-matched%5Fcleaned.csv?rlkey=uinujk66t8f6lx60i15th75xd&st=ry2i6oe3&dl=0&ref=blog.datahut.co) | Contains only the Flipkart details of the matched items, aligned with their corresponding Amazon listings | | [**Combined Matched Products**](https://www.dropbox.com/scl/fi/b3ab7vimtcxsa63rb22j7/matched%5Fproduct%5Fcleaned.csv?rlkey=3pkpf6zcczvgu8o10fu3mhild&st=axdla4k2&dl=0&ref=blog.datahut.co) | This file consolidates all matched records from both Amazon and Flipkart, making it easier to analyze the overall matching accuracy | | [**Matched\_jaccard**](https://www.dropbox.com/scl/fi/8w6ppl8rvxqbc3gnur4ii/matched%5Fjaccard.csv?rlkey=8j0vee5fhopi317ff45digi67&st=guuypdqu&dl=0&ref=blog.datahut.co) | **Contains all matched product records from both Amazon and Flipkart, it focuses on the Jaccard distance metric. This table is useful for analyzing matches using set-based similarity rather than vector-based similarity** | ### **AUTHOR** I’m Anusha P O, Data Science Intern at Datahut. I build smart data-processing workflows and product-matching systems for e-commerce sites like Amazon and Flipkart. In this blog, I explain how we used TF-IDF vectorization, cosine similarity, and Jaccard distance to find matching products in large datasets. We turned messy, inconsistent product listings into clean, organized records ready for analysis. At [Datahut](https://www.datahut.co/?ref=blog.datahut.co), we help businesses use e-commerce data. We design strong and scalable solutions for product matching. We also work on removing duplicate items from catalogs. We gather information about competitors too. Learn more about our expertise through our [Datahut Services](https://www.datahut.co/terms-of-service?ref=blog.datahut.co) and discover[ ](https://www.datahut.co/process?ref=blog.datahut.co)[how we extract and deliver data](https://www.datahut.co/process?ref=blog.datahut.co). You can also get to know our team and mission on the[ ](https://www.datahut.co/?ref=blog.datahut.co)[About Datahut](https://www.datahut.co/?ref=blog.datahut.co) page. For more insights and tutorials, visit the[ ](https://www.blog.datahut.co/)[Blog](https://www.blog.datahut.co/) to explore all posts. If you want to use data-driven methods to improve product discovery, pricing analysis, or catalog management, contact us through the chat widget on the right side of our website. Let’s turn your e-commerce data challenges into actionable insights. ### **FAQs** **1.** What are the best practices for tuning cosine similarity thresholds in product matching tasks? The best practices include testing different threshold values on a labeled sample, analyzing false positives and false negatives, adjusting thresholds based on product category sensitivity, and using validation data to find a balanced score that maximizes accuracy while minimizing mismatches. **2\.** What are common issues in product matching implementations using cosine similarity and TF-IDF, and how can they be resolved? Common issues include mismatched titles due to noisy text, missing brand/model information, and low similarity scores for products with different wording. TF-IDF may also over-weight rare words or under perform on short titles. These issues can be resolved by cleaning text (normalization, stop-word removal), adding more attributes (brand, model, specs), using better weighting, applying thresholds carefully, and combining cosine similarity with advanced methods like embedding for improved accuracy. **3\.** What challenges does product matching address in e-commerce platforms? Cosine similarity combined with TF-IDF improves product matching. It turns product titles into numerical vectors. Then it measures how closely these vectors align. This improves search accuracy, price comparison, and overall shopping experience. **4\.** How does cosine similarity combined with TF-IDF help improve product matching in online stores? Cosine similarity combined with TF-IDF improves product matching by converting product titles into numerical vectors and measuring how closely they align. TF-IDF highlights important keywords, while cosine similarity checks their directional similarity, allowing the system to identify matching products even when titles are written differently. **5\.** Product matching helps fix problems like duplicate listings and inconsistent product titles. It also handles different naming styles across sellers. It makes it easier to tell if two products are the same. 6.Yes, cosine similarity works well. It measures how closely two text vectors align. This makes it great for matching product titles, descriptions, and other e-commerce data. It works even when the wording is different. ### Olay on Amazon: What 266 Products and 651,000 Reviews Reveal URL: https://www.blog.datahut.co/post/olay-on-amazon/ Last updated: 2026-09-07T09:21:57.000Z Amazon’s[ skincare category](https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/) is highly competitive, with thousands of brands and millions of products. Despite this, [Olay](https://www.olay.com/?ref=blog.datahut.co), a flagship Procter & Gamble brand, has established a catalog of 266 active products that consistently leads in search visibility, review volume, and buyer satisfaction. With an average sale price of **$25.82** and an average rating of **4.44 out of 5.0** from **651,038 customer reviews**, Olay demonstrates solid market performance without compromising customer satisfaction. This analysis studies Olay’s approach across 14 product forms, three price tiers, and a targeted discount strategy that is aggressive in growth areas and disciplined in established segments. ## **Key Findings from Olay's**[ **Amazon Product**](https://www.blog.datahut.co/post/discover-the-hidden-value-in-amazon-data-a-comprehensive-guide/) **Analysis** ## ![olay category overview on amazon](https://www.blog.datahut.co/content/images/2026/07/category_overview.webp) - Product Form Dominance. Creams account for 94 SKUs - more than double any other product form - giving Olay an outsized share of search visibility. - Strategic Discounting. Olay discounts most aggressively in high-growth categories like Serums (9.6% avg.) while protecting margins on Creams (4.56% avg.). - The Mid-Range Sweet Spot. The $15–$30 price tier holds the most products (117 SKUs) and earns the highest average rating, 4.51 out of 5.0. - A Distributed Portfolio. The top 10 products control only 25.32% of all reviews - a healthy, resilient spread with no single point of failure. - Sensory Drivers. Scent and skin softness are the top two review attributes - instant sensory satisfaction fuels 5-star reviews more than anything else. - Thoughtful Discounting. With an overall average discount of 9.44% across 266 products, Olay balances trial-driving promotions against brand value. ## **Olay's** [**Amazon Product**](https://www.blog.datahut.co/post/how-to-scrape-amazon-and-other-large-e-commerce-websites-at-a-large-scale/) **Portfolio: Which Product Types Dominate?** Olay's 266 active products span 14 distinct product forms, generating 651,038 customer reviews at an average rating of 4.44 - well above the broader skincare category average of roughly 4.1-4.2 on [Amazon](https://www.amazon.com/?ref=blog.datahut.co). Cream, Lotion, and Body Wash are the top three forms by review volume, confirming where Olay's[ consumer](https://www.mckinsey.com/industries/consumer-packaged-goods/our-insights?ref=blog.datahut.co) trust is most deeply established. Cream alone accounts for **363,514 reviews** \- 55% of all reviews the brand has collected - making it the single most dominant product form in Olay's [Amazon presence.](https://www.blog.datahut.co/post/how-to-scrape-product-data-from-amazon-us/) ### **Cream: The Anchor** With 94 SKUs and 363,514 reviews, Cream is the foundation of Olay's Amazon presence. Its range includes anti-aging, moisturizing, retinol, and brightening products, meeting nearly all skincare needs and customer segments. This broad coverage ensures Olay appears in most relevant face-care searches on Amazon. ### **Lotion and Body Wash: The Everyday Layer** Lotion ranks second with 96,217 reviews across 30 SKUs, demonstrating strong engagement per product due to daily use and repeat purchases. Body Wash has 78,192 reviews across 43 SKUs and is positioned in the budget segment, serving as an entry point for new customers who may later upgrade within the Olay range > The Cream Effect: Cream dominates with 94 SKUs - more than double the second-largest category, Body Wash (43). Gel (16), Wipes (13), and Bar (11) serve as targeted niche plays, while long-tail formats like Drop (6), Liquid (5), and Oil (3) capture precision buyers seeking specialized formulations. Owning 94 listings in one format isn't just inventory - it's a calculated move to own a customer's attention before a single click even happens. ![SKU count of Olay product forms across amazon](https://www.blog.datahut.co/content/images/2026/07/skucount.webp) ## **Olay's Pricing and Discount Strategy on Amazon** Olay’s discounting pattern reflects a targeted pricing strategy: the brand reduces prices in markets where it is building share and maintains pricing where it already leads. The overall category average discount is 9.44%, indicating a thoughtful rather than reactive approach. Serums have the highest average discount at **9.6%**, reflecting Olay’s investment in expanding this competitive, higher-margin segment by encouraging initial purchases. Liquid (7.8%), Lotion (7.69%), and Drop (7.33%) also receive above-average discounts in categories with greater competition. Body Wash has a moderate 5.48% discount, which is typical for a mature, everyday-use category where established habits lessen the need for promotions. > Margin protection guides Olay’s approach: Cream, the brand’s largest SKU category, has the lowest discount at 4.56%. Wipes are discounted even less, at 2.71%. Olay does not rely on discounts in categories where it already leads; the largest discounts are reserved for segments still seeking market share. ![Olay discount strategy on Amazon](https://www.blog.datahut.co/content/images/2026/07/av_discount.webp) ### **Price Tier Analysis: Identifying the Mid-Range Advantage** The distribution of Olay's 266 active products across various price tiers demonstrates the brand's comprehensive understanding of its customer segments, ranging from budget-conscious consumers purchasing Body Wash to premium buyers investing in Serums. The **Budget segment (≤$15)** comprises 78 SKUs with an average price of $10.56 and a strong average rating of 4.46, primarily driven by Body Wash and entry-level Creams. The **Mid-Range segment ($15–$30)** represents the core of Olay's portfolio, featuring 117 SKUs, the highest average rating in the catalog (4.51), and the most robust assortment of Creams and Lotions. The **Premium segment ($30+)** includes 71 SKUs but records the lowest average rating at 4.29, indicating that higher price points may lead to increased consumer scrutiny, especially among Cream and Serum buyers comparing these products to luxury alternatives. All three price bands fall within a narrow 0.22-point range in average ratings, indicating that pricing alone does not determine customer satisfaction. The mid-range tier excels in both product volume and customer satisfaction, establishing it as the foundation of Olay's product portfolio. ![Price brands across Amazon - Olay products](https://www.blog.datahut.co/content/images/2026/07/price-band.webp) ### **Portfolio Resilience: No Single Hero Product** How distributed are Olay's reviews across its catalog? This is important because brands that depend on one or two hero products are vulnerable. A single bad batch or discontinuation can quickly reduce visibility. In contrast to winner-take-all categories where the top 10 products account for 70% or more of reviews, Olay's top 10 represent only **25.32%**. The remaining 75% of reviews are distributed across more than 256 other products. This indicates that Olay shoppers engage with a broad range of formulations and use cases, rather than focusing on a few bestsellers. > Why This Matters: No single product dominates demand. If a top SKU is discontinued or faces a quality issue, the brand can absorb the impact without a significant loss in visibility. New Olay launches also have the opportunity to gain review share without needing to replace an established bestseller. ![Review concentration of olay products on amazon](https://www.blog.datahut.co/content/images/2026/07/tier.webp) ## **Voice of the Customer: Engagement and Ratings by Format** The average number of reviews per SKU by product form highlights where consumers are most engaged and where Olay has established its strongest communities of loyal, repeat buyers. Dissolving Pads lead with an average of 4.7K **reviews per listing,** the highest engagement rate in the catalog. This is driven by loyal repeat buyers of these water-activated cleansing pads, despite the format's niche size. Clay (3.8K), Cream (3.7K), and Lotion Pack of 2 (3.6K) follow closely. Body Wash, Olay's second-largest category by SKU count, has the lowest average engagement at 1.8K reviews per listing, reflecting a commoditized segment where brand switching is more frequent. Ratings provide a complementary perspective. Across all product forms, Olay maintains a consistent satisfaction score, with only a 0.46-point range between the lowest and highest among nine categories. ![Average review of Olay products on Amazon](https://www.blog.datahut.co/content/images/2026/07/avgreviewcount.webp) Average review of Olay products on Amazon Liquid (4.74) and Bar (4.62) lead in quality perception, reflecting strong performance in smaller-volume categories. Gel (4.51) and Cream (4.48) also perform well, confirming that Olay’s high-volume formats deliver consistent quality at scale. Wipes (4.28) rank lowest, making them the most significant opportunity for product or packaging improvement. ![Average rating of Olay products on Amazon](https://www.blog.datahut.co/content/images/2026/07/avg_rating.webp) ## **What Actually Drives Purchase Decisions: Review Text Mining** In addition to star ratings, customer reviews highlight the emotional and functional factors that drive purchases, repeat business, and brand loyalty. Analysis of Olay’s reviews identifies seven key attributes: six positive and one risk indicator. Smell and scent are mentioned most frequently, with 1,321 references. Amazon skincare buyers prioritize immediate sensory experiences, and a pleasant scent is the strongest driver of 5-star reviews and repeat purchases. Quality (1,311 mentions) and Skin Softness (1,198) complete the top three, reflecting the core expectations that Olay products are effective and pleasant from the first use. Hydration and moisturizing, with 1,157 mentions, are the main functional reasons customers choose Olay. Value for money (818 mentions) is also important. With an average price of $25.82, Olay is often compared to premium brands priced at $80 to $200, and customers reward Olay when its results justify the price difference. > Texture and greasiness, with 479 mostly negative mentions, represent the most significant risk identified in the data. Products that feel sticky or greasy can negate other positive attributes, highlighting a key area for future product development. ![what makes people buy olay products](https://www.blog.datahut.co/content/images/2026/07/mover.webp) ![emotional drivers driving the purchase](https://www.blog.datahut.co/content/images/2026/07/mover2.webp) ## **What This All Means: The Olay Formula** Olay’s Amazon strategy is based on three concurrent commitments rather than a single approach. **Anchor: Establish strong presence in a single format. With 94 Cream SKUs,** Olay appears in nearly every relevant face-care search on Amazon. **Balance: Offer products** across budget, mid-range, and premium tiers to avoid reliance on a single price point. Data indicates that mid-range products achieve the highest volume and satisfaction. **Protect and Invest: Maintain margins in leading categories, such as Creams with a 4.56% average discount, and allocate promotional spending to areas where the brand is still growing, such as Serums at 9.6%**. > Olay’s primary advantage is not a single standout product or formula, but rather 651,038 reviews distributed across 266 SKUs. This distribution ensures that no single product failure significantly impacts the brand’s Amazon presence. Building this level of resilience requires years of consistent performance and robust data. Review data provides important human insights. Shoppers choose Olay not just for availability or price, but because it offers a pleasant scent, immediate skin softness, and good value. The 1,321 scent mentions and 1,198 references to skin softness demonstrate that sensory and emotional validation are as important as clinical results. The 479 complaints about texture and greasiness highlight that a single unresolved sensory issue can outweigh other positive attributes. ## **Turn These Insights Into Growth** Structured, timely, and accurate data forms the basis of effective e-commerce decisions. Brands using category-level intelligence not only know what customers are saying, but also understand purchase drivers, loyalty risks, and emerging growth opportunities. If you want to: - Track competitors' pricing and discounts in near real-time - Benchmark your brand's share of reviews and visibility against category leaders - Identify hidden revenue leaks, assortment blind spots, or emerging customer trends - Build dashboards your marketing, sales, and merchandising teams can act on instantly - Power AI models and forecast demand with clean, reliable marketplace data [Datahut](https://www.datahut.co/contact?ref=blog.datahut.co) provides fully structured, automatically refreshed data tailored to your needs. Request a free 10-minute data audit, consult with a Datahut strategist about category monitoring or pricing intelligence, or launch a pilot within days. No infrastructure or internal engineering resources are required. ## **Frequently Asked Questions** ### **Why does Olay have so many cream products on Amazon?** Cream is Olay's primary skincare format, with 94 SKUs and over 363,000 reviews covering anti-aging, hydration, brightening, and retinol treatments. This variety maximizes Olay's visibility in Amazon's face-care search results and addresses diverse skincare needs. ### **Are Olay products on Amazon considered affordable or premium?** Most Olay products are priced in the mid-range segment, averaging $25.82\. This positions the brand between budget and luxury skincare, offering premium results without premium pricing.ssarily. While premium products exist, the mid-range band ($15–$30) shows the strongest combination of high ratings and product availability, suggesting customers feel the greatest value and satisfaction in that range. ### **How many reviews do Olay products have on Amazon?** Olay products have received over 651,000 customer reviews, providing a robust and statistically meaningful base for evaluating product performance and quality. ### **Why do Olay products maintain such high ratings?** Olay's average rating of 4.44 out of 5 is notable for a catalog of this size. Reviews consistently highlight scent, product quality, skin softness, and hydration as key drivers of positive experiences. ### **What price range has the most Olay products?** The mid-range tier ($15–$30) includes the largest portion of the catalog, with 117 SKUs. This approach allows Olay to reach a broad audience while maintaining strong performance and satisfaction. ### **What do customers care about most when buying Olay products?** Scent, product quality, skin softness, and moisturizing performance are the primary decision drivers. These factors deliver immediate sensory and functional benefits, strongly influencing positive reviews and repeat purchases. ### How to Scrape Product Data from Lazada Using Python? URL: https://www.blog.datahut.co/post/how-to-scrape-product-data-from-lazada-using-python/ Last updated: 2026-09-07T09:32:04.000Z Let’s step into the world of online shopping for a moment. Every day, websites like Lazada are updated with thousands of products—new listings, changing prices, customer reviews, and more. Now imagine trying to gather all that information by visiting each product page one by one. Sounds exhausting, right? That’s exactly where web scraping comes in to save the day. [Web scraping](https://www.blog.datahut.co/post/what-is-web-scraping/) is like teaching a computer how to browse a website and collect specific pieces of information—just like you would, but much faster. It helps turn scattered data from websites into clean, structured formats like spreadsheets. This is especially useful in fields like price tracking, market research, or keeping an eye on competitors. For this project, we chose Lazada’s Skincare – Face Masks & Packs category. This section alone has a huge variety of products. Collecting details like prices, names, discounts, and seller info manually would take forever. So, we automated it. Using a script, we first collected all the product links by scrolling through the pages. Then, we visited each of those links to extract useful details- such as product name, price, any available offers, ratings, number of reviews, and seller information. With everything neatly gathered, the data will be ready to analyze—whether that’s for understanding [pricing trends](https://www.blog.datahut.co/post/competitor-price-monitoring/), finding popular products, or studying how different sellers are performing. ## Product data from Lazada: Product Links Scraping Importing Required Libraries ``` import logging import random from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup import sqlite3 import time ``` The script begins by loading all the tools it needs to do the job. Think of it like gathering all your ingredients before starting to cook—it helps everything run smoothly. First, we bring in the logging library. This helps us keep track of what’s happening in the script—whether everything is going fine or if something goes wrong. Then we use random to add variety in things like selecting user-agents (these help us mimic different browsers) and setting small delays between actions. The time module is used to pause the script briefly between steps so we don’t overload the website with requests. Next, we load sync\_playwright from [Playwright](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch/). This is the tool that helps us handle web pages that load content dynamically—just like when you scroll down an online store and more products appear. [Playwright](https://playwright.dev/python/?ref=blog.datahut.co) lets us control a browser behind the scenes, just like a person would. We also use [BeautifulSoup](https://www.blog.datahut.co/post/how-to-scrape-h-m-product-data-using-python/) to read and understand the structure of the web page so we can find and extract the product links easily. And finally, sqlite3 is included so we can save those links into a database for later use. Starting with these imports ensures the script has everything it needs before it begins. Playwright is especially important here, because most online stores use JavaScript to show content—and Playwright helps us handle that smoothly. Altogether, these libraries form a strong setup for collecting and storing product links efficiently. ### Configuring Logging ``` # Configure logging logging.basicConfig( filename="lazada_scraper.log", filemode="a", format="%(asctime)s - %(levelname)s - %(message)s", level=logging.INFO ) ``` Once the tools are set up, the script moves on to something very important—**logging**. This is like keeping a diary of everything the script does. It helps us know what went right, what went wrong, and when things happened. The script sets up a log file called lazada\_scraper.log. As the scraper runs, it keeps writing updates into this file. Each log entry includes the time, the type of message (like info, warning, or error), and a short message about what happened. This is really helpful when you want to go back and check if something went wrong—or just see how the script performed. The logging level is set to **INFO**, which means it records general updates, any warnings, and errors—but skips very detailed debugging messages. This keeps the log clean and easy to read. Logging becomes especially valuable in web scraping. If a page fails to load, if a user-agent is missing, or if there’s a problem saving data to the database, the log tells you where and why it happened. It’s like having a map to find where things went off track. And when the scraper is running through many pages, structured logging ensures you can spot and fix issues without guessing. ### Loading User Agents ### ``` def load_user_agents(file_path): """ Load a list of user agents from a text file for browser spoofing. Args: file_path (str): Path to the text file containing user agents, one per line Returns: list: List of user agent strings. Returns empty list if file reading fails. Raises: FileNotFoundError: If the specified file path doesn't exist IOError: If there are issues reading the file Notes: - Empty lines in the file are skipped - Logs a warning if no user agents are found in the file - File should contain one user agent string per line """ try: with open(file_path, 'r') as file: user_agents = [line.strip() for line in file.readlines() if line.strip()] if not user_agents: logging.warning("No user agents found in the file.") return user_agents except Exception as e: logging.error(f"Failed to load user agents: {e}") return [] ``` The next part of the script deals with something websites are very cautious about—**bots**. Many websites try to block scrapers, especially if they detect that the same “browser” is sending too many requests. To avoid this, we use something called a **user-agent**. A user-agent is just a short message that tells the website what kind of browser or device is being used—like Chrome on Windows or Safari on an iPhone. The function load\_user\_agents(file\_path) helps the scraper read a list of these user-agent strings from a text file. This way, the scraper can pretend to be a different browser each time it makes a request. Here’s how it works: the function tries to open the file and read the user-agents line by line, storing them in a list. If the file is missing or empty, it doesn’t break the script—instead, it logs a warning and returns an empty list. That way, the script continues running, and you can check the logs to see what went wrong. By rotating user-agents, the scraper blends in more naturally with normal web traffic, making it less likely to get blocked. And with proper error handling, it stays stable even when something unexpected happens. ```python def initialize_browser(user_agents): """ Initialize a Playwright browser session with a randomly selected user agent. Args: user_agents (list): List of user agent strings to choose from Returns: tuple: (playwright instance, browser instance, page instance) - playwright: The Playwright context manager - browser: The browser instance (Chromium) - page: The initialized page object Raises: Exception: If browser initialization fails IndexError: If user_agents list is empty Notes: - Launches browser in non-headless mode (visible browser window) - Randomly selects a user agent from the provided list - Logs the selected user agent for debugging purposes """ try: user_agent = random.choice(user_agents) logging.info(f"Using user agent: {user_agent}") playwright = sync_playwright().start() browser = playwright.chromium.launch(headless=False) context = browser.new_context( user_agent=user_agent ) page = context.new_page() return playwright, browser, page except Exception as e: logging.error(f"Failed to initialize browser: {e}") raise ``` Next, the script sets up the browser—the heart of the scraping process. The initialize\_browser(user\_agents) function uses Playwright to open a browser window that the script can control, just like a real user browsing the site. To make things more realistic, it randomly picks one user-agent from the list we loaded earlier. This user-agent is logged so you know which identity the scraper is using for that session. Then, it opens a **Chromium** browser—not in headless mode, which means you can actually see the browser on your screen as it loads pages. This is especially helpful during testing or when something isn’t working and you want to watch what’s happening. The function gives back three things: the Playwright instance (which manages everything), the browser object, and the page object (used to load and interact with websites). If something goes wrong—like the browser fails to start—the function doesn’t just stop silently. It logs the issue and raises an error, making sure the script doesn’t continue with a broken setup. This way, the scraper acts more like a real person visiting the site, which helps avoid detection and makes debugging easier when needed. ### Initializing the Database ```python def initialize_database(): """ Initialize SQLite database and create the required table structure. Returns: sqlite3.Connection: Database connection object Raises: sqlite3.Error: If database initialization or table creation fails Notes: - Creates a new database file 'lazada.db' if it doesn't exist - Creates a table 'lazada_products' with columns: * id (INTEGER PRIMARY KEY AUTOINCREMENT) * link (TEXT UNIQUE) - The UNIQUE constraint prevents duplicate product links """ try: conn = sqlite3.connect("lazada.db") cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS lazada_products ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT UNIQUE ) """) conn.commit() return conn except sqlite3.Error as e: logging.error(f"Database initialization failed: {e}") raise ``` The initialize\_database() function takes care of setting up a place to store the product links we’re collecting. It connects to an **SQLite database**, which is a simple, file-based database that works perfectly for small to medium projects like this. Inside the database, it checks if a table called lazada\_products exists. If not, it creates one with two columns: an id that automatically increases for each entry, and a link column to store the product URL. The link column is set to be **unique**, which means duplicate URLs won’t be added more than once. Using a database like this is much more reliable than saving links in plain text files. It allows you to search, filter, and organize the data easily—especially helpful when the project grows. And if something goes wrong while setting up the database, the script doesn’t just fail quietly—it logs the problem so you’ll know exactly what needs fixing. ### Fetching Page Content ```python def fetch_page_content(page, url): """ Navigate to a URL and retrieve the page content using Playwright. Args: page (playwright.sync_api.Page): Initialized Playwright page object url (str): The URL to fetch content from Returns: str or None: HTML content of the page if successful, None if failed Notes: - Sets a high timeout (20 minutes) to handle slow-loading pages - Includes a 2-second delay after navigation to allow dynamic content to load - Logs errors if page fetch fails Rate Limiting: - Includes a 2-second fixed delay to prevent overwhelming the server """ try: page.goto(url, timeout=1200000) # Set a timeout to avoid indefinite waits time.sleep(2) # Allow time for the page to load return page.content() except Exception as e: logging.error(f"Failed to fetch content from {url}: {e}") return None ``` The fetch\_page\_content(page, url) function handles the job of opening a web page and getting its full HTML content—just like what you'd see if you visited the page in your own browser. It uses Playwright to open the given URL and sets a long timeout (up to 20 minutes) to make sure even slow-loading pages have enough time to load. After the page loads, the script waits an extra two seconds. This short pause is important because many e-commerce sites use JavaScript to load content after the page first appears. Without this wait, we might miss out on important product details. Unlike simple HTTP requests that only grab the raw page source, Playwright gives us the fully rendered version, just like what an actual user would see. And if something goes wrong—like a slow network or a page that fails to load—an error is recorded in the logs. This way, the script won’t crash unexpectedly, and you’ll know where the problem occurred. ### Parsing Product Links ```python def parse_product_links(html_content): """ Extract product links from Lazada's HTML content using BeautifulSoup. Args: html_content (str): Raw HTML content of the page Returns: list: List of product URLs found on the page Notes: - Uses CSS selector to find product card elements - Handles relative URLs by prepending 'https:' when necessary - Returns empty list if parsing fails or no links found """ try: soup = BeautifulSoup(html_content, "html.parser") product_links = [] product_cards = soup.select("#root > div > div.ant-row.FrEdP.css-1bkhbmc.app > div:nth-child(1) > div > div.ant-col.ant-col-20.ant-col-push-4.Jv5R8.css-1bkhbmc.app > div._17mcb > div > div > div > div.ICdUp > div > a") # Adjust selector for product links for card in product_cards: link = card.get("href") if link and not link.startswith("https:"): link = "https:" + link if link: product_links.append(link) return product_links except Exception as e: logging.error(f"Error parsing HTML content: {e}") return [] ``` The parse\_product\_links(html\_content) function is where the actual scraping happens. It takes the full HTML of a page and uses [**BeautifulSoup**](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co) to look through the content and find product links. It searches for specific parts of the page—called product cards—using a CSS selector. From each product card, it grabs the href value, which is the link to the individual product. Sometimes these links are incomplete (starting with just a /), so the function adds https: at the beginning to make them complete and usable. This step helps us pull out only the links we care about from the rest of the page’s structure. Since websites often change their design, the CSS selector used here should be reviewed from time to time to make sure it's still accurate. If anything goes wrong—like the expected structure isn’t found—the function won’t break the whole script. Instead, it logs an error and returns an empty list. All collected links are cleaned and standardized before storing them in the database, keeping everything consistent and easy to work with later. ### Saving Links to the Database ```python def save_links_to_database(conn, links): """ Save extracted product links to SQLite database. Args: conn (sqlite3.Connection): Database connection object links (list): List of product URLs to save Notes: - Uses INSERT OR IGNORE to handle duplicate links gracefully - Commits transaction after all insertions - Logs any errors during insertion process - Does not close the database connection (handled by caller) Database Schema: Table: lazada_products Columns: - id (INTEGER PRIMARY KEY AUTOINCREMENT) - link (TEXT UNIQUE) """ cursor = conn.cursor() for link in links: try: cursor.execute("INSERT OR IGNORE INTO lazada_products (link) VALUES (?)", (link,)) except sqlite3.Error as e: logging.error(f"Error saving link {link}: {e}") conn.commit() ``` The save\_links\_to\_database(conn, links) function takes the list of product links we’ve scraped and saves them into the SQLite database. It uses an SQL command called INSERT OR IGNORE, which is helpful because it avoids saving the same link more than once. If a link is already in the database, the command simply skips it instead of causing an error. This keeps the database clean and free of duplicates. Each time a link is added, the function commits the change—basically telling the database, “Save this now.” This way, even if the script suddenly stops running, the links that were already inserted won’t be lost. If there’s any trouble while inserting a link—maybe due to a database issue—it gets logged. That way, we know something went wrong without the whole script crashing. In the bigger picture, this function makes sure every product URL we collect gets safely stored and is ready for any later steps, like analysis or further scraping. ### Implementing a Random Delay ```python def random_delay(min_delay=2, max_delay=4): """ Implement a random delay between web requests for rate limiting. Args: min_delay (float): Minimum delay in seconds (default: 2) max_delay (float): Maximum delay in seconds (default: 4) Notes: - Uses uniform distribution for randomization - Logs the actual delay duration for monitoring - Helps prevent detection of automated scraping - Recommended to adjust delay range based on website's terms of service """ delay = random.uniform(min_delay, max_delay) logging.info(f"Sleeping for {delay:.2f} seconds.") time.sleep(delay) ``` The random\_delay(min\_delay, max\_delay) function adds a randomly introduced delay between requests, simulating human web browsing behavior. It records the actual delay period and suspends execution with time.sleep(). Rate limiting is crucial in web scraping to avoid detection and IP blocks. By incorporating small, random pauses, the scraper makes it less likely to activate anti-bot measures. Varying the delay range can assist in finding a balance between efficiency and stealth, based on the target site's terms of use. The fact that randomized delays are included helps the scraping process seem more natural, making it less likely to be blocked. ### Scraping Lazada Product Links ```python def scrape_lazada_links(base_url, start_page=1, end_page=102, user_agents_file="user_agents.txt"): """ Main scraping function to extract product links from multiple Lazada pages. Args: base_url (str): URL template with {page_no} placeholder for pagination start_page (int): First page number to scrape (default: 1) end_page (int): Last page number to scrape (default: 102) user_agents_file (str): Path to file containing user agents (default: "user_agents.txt") Notes: - Initializes browser with random user agent rotation - Creates/connects to SQLite database for storing links - Implements rate limiting with random delays between requests - Handles errors gracefully and logs all activities - Ensures proper cleanup of resources in all scenarios Process Flow: 1. Load user agents and initialize browser 2. Set up database connection 3. Iterate through pages and extract product links 4. Save links to database 5. Clean up resources Rate Limiting: - Implements random delays between requests - Adjustable through random_delay() function parameters """ user_agents = load_user_agents(user_agents_file) if not user_agents: logging.critical("No user agents available. Exiting the scraper.") return playwright, browser, page = initialize_browser(user_agents) conn = initialize_database() try: for page_no in range(start_page, end_page+1): current_url = base_url.format(page_no=page_no) logging.info(f"Scraping page: {current_url}") html_content = fetch_page_content(page, current_url) if html_content: product_links = parse_product_links(html_content) if product_links: save_links_to_database(conn, product_links) else: logging.warning(f"No product links found on page {current_url}.") else: logging.error(f"Failed to fetch or parse content from page {current_url}. Skipping.") # Add a random delay between requests random_delay() except Exception as e: logging.critical(f"Critical error during scraping: {e}") finally: conn.close() browser.close() playwright.stop() logging.info("Scraping process completed.") ``` The scrape\_lazada\_links(base\_url, start\_page, end\_page, user\_agents\_file) function is the heart of the scraper. It runs the entire scraping process from start to finish. Here's what it does in order: - Loads user agents from a file. - Starts the Playwright browser and SQLite database. - Loops through the specified page range. - For each page, it: - Builds the URL. - Loads and waits for content. - Extracts product links. - Saves them to the database. - Waits randomly before the next request (to avoid detection). - Handles any errors during scraping so the script doesn’t crash midway. - Cleans up at the end by closing the browser, database, and stopping Playwright. If no user agents are found, it logs a fatal error and stops early. This function ties all parts together—link extraction, storage, and browser automation—into a controlled and reusable workflow. Its modular structure means it’s easy to adjust, extend, or debug when needed. ### Running the Scraper The script includes a main execution block: ```python if __name__ == "__main__": BASE_URL = "https://www.lazada.sg/skincare/?page={page_no}&spm=a2o42.tm80139881.cate_4.1.7d83yXEryXErvj" try: scrape_lazada_links(BASE_URL, start_page=1, end_page=102) logging.info("Scraping finished successfully.") except Exception as e: logging.critical(f"Scraper failed: {e}") ``` This final block ensures that the script only runs when you execute it directly—like running it from a terminal or command line. It acts like a gatekeeper, making sure the scraping process doesn't accidentally start when the file is imported somewhere else. The BASE\_URL here is a template that includes {page\_no}, which gets replaced by actual page numbers as the scraper loops through them. In this case, it scrapes from page 1 to 102\. Throughout the process, the script keeps an eye out for serious issues—logging any major failures that happen. If all goes well, it logs a success message once everything is done. This setup makes the script clean and manageable. It starts only when you want it to and helps catch problems early. Now that we’ve collected all the product links, the next big step is to extract detailed information from each of those product pages. In the upcoming sections, we’ll walk through how the script approaches that, breaking it down into clear steps to make the process easy to follow. ## Product Data Scraping ```python import logging import sqlite3 import time import random from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError from bs4 import BeautifulSoup # Configure logging logging.basicConfig( filename="lazada_scraper_data.log", filemode="a", format="%(asctime)s - %(levelname)s - %(message)s", level=logging.INFO ) ``` Just like before, we start by importing the required libraries—only this time, they’re geared toward pulling detailed information from each product page. We reuse the same tools, but now they focus on extracting rich product data instead of just the links. As expected, logging is set up again to track the script’s activity, helping us catch issues and understand what’s happening behind the scenes. ### **initialize\_browser Function** ```python def initialize_browser(user_agents): """ Initialize a Playwright browser instance with a randomly selected user agent. Args: user_agents (list): List of user agent strings to choose from Returns: tuple: (playwright instance, browser instance, page instance) - playwright: The Playwright context manager - browser: The Chromium browser instance - page: The initialized page object Raises: Exception: If browser initialization fails IndexError: If user_agents list is empty Notes: - Launches browser in non-headless mode for debugging - Uses random user agent for each session to avoid detection - Creates a new browser context for clean session management """ try: user_agent = random.choice(user_agents) playwright = sync_playwright().start() browser = playwright.chromium.launch(headless=False) context = browser.new_context(user_agent=user_agent) page = context.new_page() return playwright, browser, page except Exception as e: logging.error(f"Failed to initialize browser: {e}") raise ``` This method sets up the browser environment where the scraper will run. It starts by picking a random user agent from the provided list. This user agent mimics real browsers and helps the scraper avoid detection by the website. Randomizing it reduces the chance of getting blocked. Next, it initializes Playwright and opens a Chromium browser in **non-headless mode**, which means the browser window is visible. This is helpful for debugging, though in production it can be switched to **headless mode** to run in the background. It then sets the browser context with the selected user agent and opens a new page for navigation. If anything goes wrong—like Playwright fails to launch or no user agent is available—it logs the error and raises an exception. This makes sure the scraping only continues when everything is properly set up. ### **load\_user\_agents Function** ```python def load_user_agents(file_path): """ Load user agent strings from a text file. Args: file_path (str): Path to the text file containing user agents Returns: list: List of user agent strings Raises: FileNotFoundError: If the specified file doesn't exist IOError: If there are issues reading the file Notes: - Each user agent should be on a separate line in the file - Empty lines are filtered out - Logs warning if file is empty """ try: with open(file_path, "r") as f: user_agents = f.read().splitlines() if user_agents: return user_agents else: logging.warning("No user agents found in the file.") return [] except FileNotFoundError as e: logging.error(f"User agents file not found: {e}") raise ``` The load\_user\_agents function loads the user-agent strings that help the scraper mimic different browsers and devices. These strings are used to make each request look like it’s coming from a real user, reducing the chance of getting blocked. It reads from a text file given by file\_path, where each line is a user agent. It filters out any empty lines and returns a clean list. If the file is empty or missing, it logs a warning or error. If something goes wrong during reading, it logs the issue and raises an exception. This setup ensures the scraper always has a pool of user agents to rotate through for safer scraping. ### **initialize\_database Function** ```python def initialize_database(): """ Initialize SQLite database and set up the required tables structure. Returns: sqlite3.Connection: Database connection object Raises: sqlite3.Error: If database operations fail Database Schema: Table: lazada_products - Existing table with added 'scraped' column (INTEGER DEFAULT 0) Table: lazada_product_details - id (INTEGER PRIMARY KEY AUTOINCREMENT) - link (TEXT UNIQUE) - title (TEXT) - brand (TEXT) - sale_price (TEXT) - price (TEXT) - discount (TEXT) - reviews (TEXT) Notes: - Checks for and adds 'scraped' column if missing - Creates product details table if it doesn't exist - Uses transactions for data consistency """ try: conn = sqlite3.connect("lazada.db") cursor = conn.cursor() # Ensure `scraped` column exists in `lazada_products` table cursor.execute("PRAGMA table_info(lazada_products)") columns = [column[1] for column in cursor.fetchall()] if "scraped" not in columns: cursor.execute("ALTER TABLE lazada_products ADD COLUMN scraped INTEGER DEFAULT 0") logging.info("Added 'scraped' column to lazada_products table.") # Create new table for scraped data cursor.execute(""" CREATE TABLE IF NOT EXISTS lazada_product_details ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT UNIQUE, title TEXT, brand TEXT, sale_price TEXT, price TEXT, discount TEXT, reviews TEXT ) """) logging.info("Initialized lazada_product_details table.") conn.commit() return conn except sqlite3.Error as e: logging.error(f"Database initialization failed: {e}") raise ``` The initialize\_database function prepares the SQLite database that will store all the scraped product data. It opens a database file named lazada.db, creating it if it doesn't already exist. This database is where all links and product details will be safely stored for later analysis. First, the function checks if the scraped column exists in the lazada\_products table. It does this using a special SQL command (PRAGMA table\_info) that looks at the table’s structure. If the scraped column isn’t found, it adds it and sets its default value to 0, meaning the product hasn’t been scraped yet. This step ensures the table is ready for the scraping process. Next, it sets up a new table called lazada\_product\_details, where the actual product data will go. This table includes columns like id, link, title, brand, sale\_price, price, discount, and reviews. These fields are important for tracking product information over time or generating reports. If any errors happen while setting up the tables—like SQL mistakes or connection issues—they're logged so you can easily find and fix the problem. This function makes sure the database is properly structured before the scraper starts collecting product details. ### **fetch\_unscraped\_links Function** ```python def fetch_unscraped_links(conn): """ Retrieve all unscraped product links from the database. Args: conn (sqlite3.Connection): Database connection object Returns: list: List of tuples containing unscraped product links Raises: sqlite3.Error: If database query fails Notes: - Returns links where 'scraped' column is 0 - Returns empty list if no unscraped links found """ try: cursor = conn.cursor() cursor.execute("SELECT link FROM lazada_products WHERE scraped = 0") return cursor.fetchall() except sqlite3.Error as e: logging.error(f"Failed to fetch unscraped links: {e}") raise ``` The fetch\_unscraped\_links function is responsible for getting the list of product links that haven’t been scraped yet. It looks in the lazada\_products table and selects all rows where the scraped column is set to 0\. This means the product link is still new and hasn’t been processed. By doing this, the scraper avoids repeating the same work or putting extra load on the database. It ensures that only fresh, unprocessed links are used during the scraping run. The function returns the result as a list of tuples, which the main scraping loop goes through one by one. If something goes wrong—like a database error or a bad SQL query—the issue is logged, and an exception is raised so it doesn’t go unnoticed. This function helps keep the workflow clean and efficient by focusing only on products that are yet to be scraped. ### **fetch\_page\_content Function** ``` def fetch_page_content(page, url): """ Navigate to a product page and retrieve its content. Args: page (playwright.Page): Playwright page object url (str): Product page URL to fetch Returns: str or None: HTML content of the page if successful, None if failed Raises: PlaywrightTimeoutError: If page load exceeds timeout Exception: For other navigation errors Notes: - Sets 60-second timeout for initial page load - Includes 30-second delay for dynamic content loading - Returns None if any errors occur during fetch """ try: page.goto(url, timeout=60000) time.sleep(30) # Wait for the page to load return page.content() except PlaywrightTimeoutError: logging.error(f"Page timeout while fetching URL: {url}") return None except Exception as e: logging.error(f"Failed to fetch content from {url}: {e}") return None ``` The fetch\_page\_content function is used to load the HTML content of a product page. It uses Playwright’s page.goto method to open the provided URL and waits until the page has fully loaded. A timeout is set to 60 seconds—so if the page takes longer than that, the script stops trying and logs an error. After the page finishes loading, the script waits an extra 30 seconds using time.sleep(30). This pause is important because many online stores load extra content—like images, prices, or reviews—after the initial page load using JavaScript. Waiting gives everything time to appear before we start scraping. If anything goes wrong—like the page doesn’t load in time or another error occurs—it’s caught and logged so you can understand what happened. If everything works smoothly, the function returns the full HTML content of the page as a string. This content will later be parsed to pull out detailed product information. This step is key to making sure the scraper is actually seeing the same content a real user would. ### **parse\_title, parse\_brand, parse\_prices, and parse\_ratings Functions** ``` def parse_title(soup): """ Extract product title from the page content. Args: soup (BeautifulSoup): Parsed HTML content Returns: str: Product title or 'N/A' if not found Notes: - Uses CSS selector: 'h1.pdp-mod-product-badge-title' - Returns 'N/A' if title element not found """ title = soup.select_one("h1.pdp-mod-product-badge-title") return title.get_text() if title else "N/A" def parse_brand(soup): """ Extract product brand from the page content. Args: soup (BeautifulSoup): Parsed HTML content Returns: str: Brand name or 'N/A' if not found Notes: - Uses CSS selector for brand link element - Returns 'N/A' if brand element not found """ brand = soup.select_one("a.pdp-link.pdp-link_size_s.pdp-link_theme_blue.pdp-product-brand__brand-link") return brand.get_text() if brand else "N/A" def parse_prices(soup): """ Extract price information from the product page. Args: soup (BeautifulSoup): Parsed HTML content Returns: tuple: (sale_price, original_price, discount) - sale_price (str): Current sale price or 'N/A' - original_price (str): Original price or 'N/A' - discount (str): Discount percentage or 'N/A' Notes: - Extracts three price-related elements using specific CSS selectors - Returns 'N/A' for any missing price information """ sale_price = soup.select_one("#module_product_price_1 > div > div > span") price=soup.select_one("#module_product_price_1 > div > div > div > span.notranslate.pdp-price.pdp-price_type_deleted.pdp-price_color_lightgray.pdp-price_size_xs") discount=soup.select_one("#module_product_price_1 > div > div > div > span.pdp-product-price__discount") return sale_price.get_text() if sale_price else "N/A",price.get_text() if price else "N/A",discount.get_text() if discount else "N/A" def parse_ratings(soup): """ Extract product review information. Args: soup (BeautifulSoup): Parsed HTML content Returns: str: Review information or 'N/A' if not found Notes: - Uses CSS selector for review element - Returns 'N/A' if review information not found """ reviews=soup.select_one("#module_product_review_star_1 > div > a") return reviews.get_text() if reviews else "N/A" ``` These functions are designed to pull specific details—like title, brand, price, and ratings—from each product page. They all use **BeautifulSoup**, a Python library that helps read and navigate the HTML content of a web page. The parse\_title function looks for the product’s name by targeting a specific HTML element: h1.pdp-mod-product-badge-title. If that element is found, it returns the text inside it. If not, it simply returns "N/A" so that the script doesn’t fail. Similarly, the parse\_brand function searches for the product’s brand using the brand link element. If it can’t find the brand, it also returns "N/A". The parse\_prices function collects three price-related values: the sale price, the original price, and the discount percentage. It uses CSS selectors to locate each of these. If any piece of this price data is missing, it substitutes "N/A" just for that part, allowing the rest of the data to be saved. Lastly, the parse\_ratings function tries to find review or rating information from the page. Again, if the expected element is missing, it returns "N/A" instead of crashing. Together, these functions make sure the scraper collects all the available product details in a structured and dependable way. And when something is missing on the page, the default values ensure the scraper keeps running smoothly. ### **save\_scraped\_data Function** ``` def save_scraped_data(conn, link, data): """ Save extracted product data to the database. Args: conn (sqlite3.Connection): Database connection object link (str): Product URL data (dict): Dictionary containing product details: - title: Product title - brand: Product brand - sale_price: Current sale price - price: Original price - discount: Discount percentage - reviews: Review information Raises: sqlite3.Error: If database insertion fails Notes: - Uses parameterized query to prevent SQL injection - Logs successful data saves and errors """ try: cursor = conn.cursor() cursor.execute(""" INSERT INTO lazada_product_details (link, title, brand,sale_price, price, discount, reviews) VALUES (?, ?, ?, ?, ?, ?, ?) """, (link, data["title"], data["brand"], data["sale_price"], data["price"], data["discount"], data["reviews"])) conn.commit() logging.info(f"Saved scraped data for link: {link}.") except sqlite3.Error as e: logging.error(f"Error saving data for link {link}: {e}") ``` The save\_scraped\_data function takes the detailed product information we’ve collected and saves it into the SQLite database. It receives the data as a dictionary and uses a **parameterized SQL query** to insert it into the lazada\_product\_details table. Using this method helps keep the database secure by preventing SQL injection. Inside the function, the data is passed into cursor.execute() for insertion. After that, conn.commit() is called to make sure the changes are saved permanently in the database. If something goes wrong—like trying to insert a duplicate entry or if there’s an issue with the SQL query—it’s caught and logged without stopping the script. This function ensures that every product’s data is stored reliably, so it can be used later for analysis, reporting, or any other tasks. ### **mark\_link\_as\_scraped Function** ``` def mark_link_as_scraped(conn, link): """ Update the scraped status of a product link in the database. Args: conn (sqlite3.Connection): Database connection object link (str): Product URL to mark as scraped Raises: sqlite3.Error: If database update fails Notes: - Sets 'scraped' column to 1 for the specified link - Logs successful updates and errors """ try: cursor = conn.cursor() cursor.execute("UPDATE lazada_products SET scraped = 1 WHERE link = ?", (link,)) conn.commit() logging.info(f"Marked link as scraped: {link}.") except sqlite3.Error as e: logging.error(f"Error marking link as scraped: {e}") ``` ### **scrape\_product\_data Function** ``` def scrape_product_data(): """ Main function to orchestrate the product data scraping process. Process Flow: 1. Load user agents and initialize browser 2. Set up database connection 3. Fetch unscraped product links 4. For each link: - Fetch page content - Parse product details - Save data to database - Mark link as scraped - Apply rate limiting delay 5. Clean up resources Error Handling: - Handles browser initialization failures - Manages database connection errors - Catches parsing and scraping exceptions - Ensures proper resource cleanup Notes: - Uses random user agents for each session - Implements rate limiting between requests - Logs all major operations and errors - Ensures proper cleanup in case of failures """ user_agents = load_user_agents("user_agents.txt") if not user_agents: logging.critical("No user agents available. Exiting scraper.") return playwright, browser, page = None, None, None conn = None try: playwright, browser, page = initialize_browser(user_agents) conn = initialize_database() links = fetch_unscraped_links(conn) for (link,) in links: logging.info(f"Scraping link: {link}") html_content = fetch_page_content(page, link) if html_content: soup = BeautifulSoup(html_content, "html.parser") data = { "title": parse_title(soup), "brand": parse_brand(soup), "sale_price": parse_prices(soup)[0], "price": parse_prices(soup)[1], "discount": parse_prices(soup)[2], "reviews": parse_ratings(soup) } if data: save_scraped_data(conn, link, data) mark_link_as_scraped(conn, link) else: logging.warning(f"No data found for link: {link}.") else: logging.error(f"Failed to fetch content for link: {link}.") random_delay() except Exception as e: logging.critical(f"Critical error during scraping: {e}") finally: if conn: conn.close() if browser: browser.close() if playwright: playwright.stop() logging.info("Scraping process completed.") ``` The scrape\_product\_data function is the core of the entire scraping workflow. It brings all the pieces together—from setting up the tools to collecting and saving the data. It starts by loading the user agents and setting up both the browser and the database. Then, it pulls a list of product links that haven’t been scraped yet. For each link in that list, it follows these steps: 1. Loads the page using fetch\_page\_content. 2. Parses the HTML with BeautifulSoup. 3. Extracts product details like title, brand, price, and ratings using the parsing functions. 4. Saves the extracted data into the database using save\_scraped\_data. 5. Marks the link as scraped using mark\_link\_as\_scraped. To avoid getting blocked by the site, the function also adds a random delay between each request using random\_delay. If something goes wrong while processing a link—like a missing page or an unexpected structure—it logs the error and moves on to the next link. This setup makes the scraper dependable and efficient, allowing it to continue running smoothly even if a few links fail along the way. ### Running the Scraper ```python if __name__ == "__main__": try: scrape_product_data() logging.info("Scraping finished successfully.") except Exception as e: logging.critical(f"Scraper failed: {e}") ``` The if **name** \== "\_\_main\_\_": block acts as the script’s starting point. It ensures that the scrape\_product\_data function runs **only** when the script is executed directly—not when it's imported into another script. Inside this block, the script tries to run the scraping process using a try block. If everything goes well, a success message is logged to confirm that the scraping finished without issues. But if something goes wrong, the except block catches the error and logs it as a **critical failure**, along with the full error details. This setup makes the script more reliable. It prevents crashes from going unnoticed, helps with debugging, and keeps the scraping process easy to manage and maintain. ## Conclusion In summary, this Lazada web scraping script offers a well-organized and reliable way to automate the collection of product links and detailed data. It smartly uses tools like **Playwright** to handle dynamic websites and **BeautifulSoup** to extract key information from the page—such as product titles, prices, brands, and reviews. All collected data is stored neatly in an **SQLite database**, making it easy to access later for analysis or reporting. The script is built with thoughtful features like **logging**, **error handling**, and **random delays** to stay under the radar and avoid detection by the website. What makes this setup strong is its attention to both functionality and stability. It’s designed to handle changes, prevent crashes, and run responsibly—respecting ethical scraping practices like pacing requests and managing system resources properly. Altogether, the script creates a smooth, flexible, and beginner-friendly scraping experience that’s not just effective—but also dependable. **AUTHOR** I’m Shahana, a Data Engineer at Datahut, where I focus on building clean and scalable data pipelines that turn unstructured web content into valuable datasets—especially for use cases in e-commerce, retail intelligence, and product tracking. At Datahut, we work closely with clients to automate data extraction from websites, even those that use dynamic elements like JavaScript and infinite scroll. In this blog, I shared a real-world example of how we scraped product data from Lazada using Playwright, BeautifulSoup, and SQLite. The solution was built to handle dynamic page content, structured storage, and responsible scraping practices—while keeping the code modular and easy to understand, especially for those just getting started. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. ## Frequently Asked Questions ### 1\. Is it legal to scrape product pricing data from Lazada? Yes, scraping publicly available product information can be legal when done responsibly and in compliance with a website's Terms of Service and applicable laws. Always review the site's policies before collecting data at scale. If you're unsure, read our guide on **Is Web Scraping Legal?**:[https://www.blog.datahut.co/post/is-web-scraping-legal](https://www.blog.datahut.co/post/is-web-scraping-legal/) ### 2\. Why does this tutorial use Playwright instead of Requests? Many modern e-commerce websites, including Lazada, load content dynamically with JavaScript. Playwright renders the page like a real browser, making it easier to extract complete product information. You can learn more in the official Playwright documentation:[https://playwright.dev/python/](https://playwright.dev/python/?ref=blog.datahut.co) ### 3\. Can I scrape other e-commerce websites using the same approach? Yes. The same workflow—collecting product URLs first and then extracting product details—works for many online stores. You'll only need to update the CSS selectors and pagination logic to match the target website. ### 4\. How can I avoid getting blocked while scraping websites? Some best practices include rotating user agents, adding random delays between requests, respecting rate limits, and avoiding unnecessary requests. For enterprise-scale projects, these techniques are often combined with proxy rotation and browser fingerprint management. ### 5\. What can businesses do with scraped pricing data? Pricing data can be used for competitor price monitoring, assortment analysis, dynamic pricing, market research, and trend analysis. Learn more about **competitor price monitoring** here:[https://www.blog.datahut.co/post/monitor-competitor-prices-automatically-for-free-no-code-web-scraping-with-n8n](https://www.blog.datahut.co/post/monitor-competitor-prices-automatically-for-free-no-code-web-scraping-with-n8n/) ### 6\. Why is SQLite used in this tutorial? SQLite is lightweight, easy to set up, and ideal for learning or small scraping projects. As your scraper grows, you can migrate to databases such as PostgreSQL or MySQL for improved scalability and performance. ### 7\. Can I scrape thousands of product pages with this script? Yes, but you'll need additional features such as proxy rotation, retry mechanisms, browser session management, and distributed scraping infrastructure. If you're collecting data at enterprise scale, consider using a professional **web scraping service**:[https://www.datahut.co](https://www.datahut.co/web-scraping-services?ref=blog.datahut.co) ### 8\. Where can I learn more about Python web scraping? If you're just getting started, check out our beginner-friendly tutorials on **How to Build a Web Crawler in Python from** **Scratch**:[https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratchand/)[ ](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratchand/)[and](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratchand/) **Top Python Web Scraping Libraries**:[https://www.blog.datahut.co/post/top-python-web-scraping-libraries](https://www.blog.datahut.co/post/top-python-web-scraping-libraries/) These guides explain the core concepts you'll use in projects like this one. ### Pure Abstention: What BCG's European Fashion Data Means for Brands URL: https://www.blog.datahut.co/post/eeuropean-fashion-spending-decline-data/ Last updated: 2026-09-07T09:40:44.000Z When European consumers tighten their belts, fashion is where the scissors come out first. [BCG's 2026 European Consumer Sentiment Survey](https://www.bcg.com/press/9june2026-europeans-worried-about-personal-finances-spending-cutbacks?ref=blog.datahut.co) \- covering more than 20,000 consumers across 11 countries - makes this explicit. Across 12 consumer categories tracked, [fashion](https://www.mckinsey.com/industries/retail/our-insights/state-of-fashion?ref=blog.datahut.co) placed last. BCG's own characterisation is blunt: "pure abstention." Groceries and pet care were the only categories to post positive net spending and even that, BCG noted, is largely because of increasing prices rather than greater volume. The story gets sharper at the subcategory level. Functional purchases - footwear, casualwear, sportswear are holding up. Aspirational ones - handbags, accessories, formal wear are in freefall. Consumers aren't cutting fashion evenly. They're cutting what feels optional. ## European Fashion Spending Declines Most in Handbags, Accessories, and Formal Wear This is the third consecutive year of rising pessimism, and BCG's assessment is clear: value-seeking is no longer a stress response. It's a default mode. For fashion brands, the old playbook - seasonal campaigns, brand equity investment, loyalty programme nudges - is aimed at consumers who've already decided to buy less and buy cheaper. The brands that hold ground from here are the ones that replace assumptions with data. Here's where that intelligence comes from. Learn how leading companies build continuous market intelligence systems instead of relying on intuition in our guide on [**Competitive Market Intelligence**](https://www.blog.datahut.co/post/competitive-market-intelligence/)[.](https://www.blog.datahut.co/post/competitive-market-intelligence/) ![](https://www.blog.datahut.co/content/images/2026/07/graph1.webp) Fashion sub category spending ## The Discount Trap - And How to Navigate It With Price Intelligence 73% of fashion consumers will only buy at a discount. The instinctive response - more promotions, deeper discounts, longer sale windows - is exactly what erodes margin and trains customers to wait. The smarter response is to promote more precisely. Instead of tracking only a few competitors, many retailers now build [**category-wide price indexes**](https://www.blog.datahut.co/post/category-price-index/) to understand whether pricing changes reflect isolated promotions or broader market shifts. The chart below tells you something important: discounts drive conversion, but brand reputation barely registers. Most brands are investing in the wrong thing. ### What Drives Brand Switching in Fashion? Discounts Matter More Than Brand Loyalty What most brands are still missing is real-time visibility into what competitors are actually doing - not what's in the trade press, but what's live on site today. Automated [competitor price monitoring](https://www.blog.datahut.co/post/dynamic-pricing/) closes this gap. Tracking rival promotional structures -which subcategories, which price points, which timing, which discount depths , tells you where the gaps are before your numbers tell you what you missed. If a competitor is holding full price on casualwear while discounting occasion wear hard, that's a signal. Act on it before the next quarterly review. ![](https://www.blog.datahut.co/content/images/2026/07/graph2.webp) brand switching triggers ## Brand Switching Is Structural. Shelf Data Tells You Who's Winning It. Brand loyalty in European fashion is at a multi-year low which means when a consumer opens Zalando looking for a winter coat, the brand they buy is often whoever shows up best at the right price. The switching is happening regardless. The question is whether it flows toward you or away from you. Most fashion brands have limited visibility into how they're actually showing up across the retail landscape. They know their DTC metrics. They don't know where their listings rank when a consumer searches "women's coat under €100" on a major European retailer, whether a private label alternative is sitting above them, or whether their product content is competitive with the brands winning that search. E-commerce shelf scraping answers these questions at scale - tracking product rankings, listing quality, pricing, and availability across Zalando, ASOS, About You, and other major European platforms simultaneously. Showing up better on the digital shelf is one of the highest-leverage things a fashion brand can do in this market. Digital shelf visibility is increasingly becoming part of broader [competitive intelligence strategies](https://www.blog.datahut.co/post/competitive-market-intelligence/) across retail. ## The Middle Is Getting Squeezed From Both Ends The BCG data reveals two distinct consumer segments in European fashion: a price-sensitive group that has cut sharply, and a more resilient segment still spending in luxury. The consumers cutting back hardest are aspirational shoppers who used to stretch for a premium product and are now abstaining entirely. ### How Europe's Fashion Market Is Splitting Between Luxury and Budget Shoppers The secondhand data makes this even more concrete. Nearly half of European consumers now buy second-hand and as the chart below shows, the motivation is overwhelmingly economic, not ethical. Brands positioning themselves on sustainability to win back the value-seeking consumer are solving the wrong problem. ![](https://www.blog.datahut.co/content/images/2026/07/graph3.webp) Splitting Between Luxury and Budget Shoppers ### Why European Consumers Are Buying More Second-Hand Fashion Data helps you figure out which segment you're actually winning and losing. Review mining across your own products and competitors' surfaces what price-sensitive customers are walking away from versus what premium-leaning customers are seeking. Assortment tracking shows where your product sits in the price hierarchy and whether you're being positioned as value or premium regardless of your intent. Looking only at competitors often misses broader pricing movements across the category. A category price index provides much stronger market context. ![](https://www.blog.datahut.co/content/images/2026/07/second-hand-purchase-motivation.webp) second hand purchase motivation ## Gen Z Discovery Has Changed - Most Fashion Brands Haven't Caught Up Discovery has structurally shifted toward social and AI, while most fashion brands are still measuring success on Google rankings and email open rates. The gap between where younger consumers find products and where brands are investing attention is widening fast. ### How Gen Z Discovers Fashion Brands Compared to Older Consumers Three places where data changes the game: **Creator intelligence over follower counts.** A micro-creator with 40,000 engaged followers in sustainable fashion in Germany is more valuable for the right brand than a macro-influencer with a million generalist followers. Scraping creator performance signals - engagement rates, comment sentiment, brand mention frequency - lets you identify and activate these creators before competitors do and before rates go up. **Language that converts, not just content that gets views.** Creator captions and product reviews contain the exact vocabulary that resonates with your target audience right now. Scraping this content at scale surfaces how consumers actually talk about products in your category - the framing, the claims, the use cases — and feeds directly into product copy and creator briefs. **AI discovery as the new SEO.** [AI visibility increasingly depends on structured, publicly available information collected from across the web.](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/) When a 22-year-old asks ChatGPT "best sustainable denim brands in Europe," an answer comes back drawn from public web content — reviews, editorial, Q&A platforms. Fashion brands showing up in those answers have a significant discovery advantage. Optimising for AI discovery means ensuring your differentiators appear clearly across the content AI models draw from. Very few fashion brands have started this yet. That's the opportunity. ![](https://www.blog.datahut.co/content/images/2026/07/genz.webp) Discovery channels by generation ## What This Adds Up To Fashion placed last across every spending category in Europe. Pure abstention. The consumer who was stretching for an aspirational purchase has stopped stretching. The switching consumer is going wherever the price is better. The brands that navigate this aren't the ones with the biggest marketing budgets. They're the ones with the clearest view of what's actually happening in competitor pricing, on the digital shelf, in the content signals that drive Gen Z discovery. That intelligence exists in public data. The question is whether you're collecting it. [ *Datahut works with fashion retailers*](https://www.blog.datahut.co/post/data-for-fashion-retailers-the-four-problems/) *and consumer brands across Europe, providing managed web scraping and market intelligence data. If you want visibility into competitor pricing, shelf position, or consumer content signals in your category,* [*let's talk*](https://datahut.co/?ref=blog.datahut.co)*.* ## Frequently Asked Questions ## 1\. Why is European fashion spending declining in 2026? European fashion spending has declined because consumers are prioritizing essential purchases over discretionary ones. According to BCG's 2026 European Consumer Sentiment Survey, fashion ranked last among major spending categories, with many shoppers postponing or avoiding non-essential purchases due to ongoing economic uncertainty and higher living costs. ## 2\. What does "Pure Abstention" mean in the European fashion market? "Pure Abstention" is BCG's description of consumers choosing not to purchase fashion products rather than simply switching brands or delaying purchases. It reflects a broader shift toward reduced discretionary spending across Europe. ## 3\. Why are discounts becoming more important for fashion retailers? More consumers are actively waiting for promotions before making purchases. This makes competitor price monitoring increasingly important, helping retailers identify pricing opportunities without relying on excessive discounting that can reduce profit margins. ## 4\. What is digital shelf intelligence in fashion retail? Digital shelf intelligence is the process of monitoring how products appear across online marketplaces and retailer websites. It includes tracking search rankings, product availability, pricing, reviews, content quality, and competitor visibility to improve online performance. ## 5\. How does competitor pricing help fashion brands? Competitor pricing data allows fashion brands to monitor promotions, pricing changes, and assortment strategies across competing retailers. This enables faster pricing decisions, protects margins, and helps brands respond to market changes in real time. ## 6\. Why is second-hand fashion growing across Europe? Second-hand fashion continues to grow because consumers are looking for more affordable alternatives to buying new products. While sustainability plays a role, affordability remains the primary reason many European shoppers purchase pre-owned clothing. ## 7\. How is Gen Z changing fashion discovery? Gen Z increasingly discovers fashion products through social media platforms, creators, online communities, and AI-powered search experiences instead of relying solely on traditional search engines. Brands that optimize for these channels improve their chances of reaching younger consumers. ## 8\. How can web scraping help fashion brands make better decisions? Web scraping enables fashion brands to collect publicly available data on competitor pricing, product assortments, digital shelf performance, reviews, and consumer trends. These insights support pricing strategies, assortment planning, market intelligence, and competitive analysis at scale. ### L'Oréal Paris on Amazon: What 300 Products and 2.8 Million Reviews Reveal URL: https://www.blog.datahut.co/post/l-oreal-paris-on-amazon/ Last updated: 2026-09-07T09:42:50.000Z The beauty and personal care category on Amazon is among the most competitive digital shelves in e-commerce. And few brands manage it with the deliberateness of [L'Oréal Paris](https://www.loreal.com/en/usa/?ref=blog.datahut.co). With 300 active products, an average sale price of $15.41, and an overall rating of 4.44 out of 5 across more than 2.8 million reviews, the brand has built something most competitors just cannot replicate: scale that doesn't sacrifice satisfaction. This [analysis](https://www.blog.datahut.co/post/how-to-conduct-amazon-product-research-using-data/) breaks down how L'Oréal does it - across 27 product forms, multiple price tiers, and a discount strategy that is aggressive where it needs to be and disciplined everywhere else. The picture that emerges is strikingly different from [clinical-focused competitors like Cetaphil.](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/) Where those brands win on focus, L'Oréal wins on coverage. ![Thisanalysisbreaks down how L'Oréal does it - across 27 product forms, multiple price tiers, and a discount strategy that is...](https://www.blog.datahut.co/content/images/2026/07/chart_1-3-2.webp) ## Six Patterns That Define L'Oréal's Amazon Playbook Before diving into each area, here are the six core findings that shape everything that follows: ![Six Patterns That Define L'Oréal's Amazon Playbook](https://www.blog.datahut.co/content/images/2026/07/img-302.png-1.webp) ## L'Oréal Paris on Amazon: A Catalog Built for Every Search Query L'Oréal's 300 active SKUs span 27 distinct product forms, making it one of the broadest catalogs in mass beauty. The strategy is simple in theory, complex in execution: have a product ready for nearly every beauty-[related search on Amazon.](https://www.blog.datahut.co/post/how-to-scrape-product-data-from-amazon-us/) Liquid is the largest form by both SKU count and review volume. Cream sits second, showing the brand's deep heritage in moisturizers and anti-aging [skincare](https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/). Aerosol products - mainly hairsprays and dry shampoos - punch well above their weight in terms of engagement per SKU. ### Top Product Forms by Review Volume ![Top Product Forms by Review Volume](https://www.blog.datahut.co/content/images/2026/07/chart_2-1-2.webp) ## The "Mass Variety" Strategy: How L'Oréal Owns the Digital Shelf The SKU [distribution](https://www.scribd.com/document/472712576/Loreal-logistics-management?ref=blog.datahut.co) reveals a deliberate "Mass Variety" approach - one built to ensure a L'Oréal product appears for every possible beauty search query on Amazon. Liquid (97 SKUs) and Cream (85 SKUs) together represent 60.6% of the entire portfolio. This concentration ensures L'Oréal owns the basic categories of both makeup and skincare. But the long tail matters just as much: with 27 forms including Serums (14 SKUs), Pencils (10 SKUs), and Powders (8 SKUs), the brand leaves no meaningful sub-category underserved. Smaller categories like Mousse, Kits, and Masks allow the brand to test niche trends and capture specialized consumer needs without losing focus on core volume drivers. For competitors trying to carve out space, this breadth creates a near-permanent first-mover advantage in search rankings. For competitors trying to carve out space, this breadth creates a near-permanent first-mover advantage in search rankings. ## Pricing Strategy: How L'Oréal Positions Across Three Tiers ![Pricing Strategy: How L'Oréal Positions Across Three Tiers](https://www.blog.datahut.co/content/images/2026/07/chart_3-1-2.webp) L'Oréal's pricing data tells a nuanced story about how a mass-market brand can maintain quality perception across vastly different price points. The three tiers each play a distinct strategic role. Unlike clinical competitors who often see ratings deteriorate in the premium tier, L'Oréal's top-end products hold steady at 4.42 - a validation of the brand's "masstige" (mass market + prestige) positioning that has defined its global strategy for decades. ## Discount Strategy: Surgical, Not Blanket Perhaps the most revealing part of the data is how L'Oréal deploys discounts. With an overall average discount of 15.23%, the brand is far from aggressive on its core catalog. Still beneath that average lies a sharply differentiated approach. ![Perhaps the most revealing part of the data is how L'Oréal deploys discounts. With an overall average discount of 15.23%,...](https://www.blog.datahut.co/content/images/2026/07/chart_4-1-2.webp) The logic here is clear and disciplined. Deep discounts on Masks (55%) and Kits (36%) serve as customer acquisition tools - they pull new shoppers into niche product categories and set them up to buy into L'Oréal's wider ecosystem. Meanwhile, core bestsellers in the Liquid and Cream categories are protected with moderate discounting, preserving margins on the products that actually carry the catalog. ## Portfolio Resilience: Why No Single Product Can Make or Break the Brand One of the most striking findings in the entire 2,824,778-review dataset is how evenly distributed engagement is across the catalog. ![One of the most striking findings in the entire 2,824,778-review dataset is how evenly distributed engagement is across the...](https://www.blog.datahut.co/content/images/2026/07/chart_5-2-2.webp) For competitors, this distributed structure is the most intimidating aspect of L'Oréal's Amazon presence. Displacing a brand that relies on one blockbuster product is relatively achievable - you out-market it or out-formulate it in one category. Displacing a brand with hundreds of "mini-blockbusters" requires competing across the entire catalog simultaneously. ## Ratings Consistency: The 0.16-Point Range That Defines the Brand Across all 27 product forms, L'Oréal's ratings fall within a remarkably tight band - 4.44 to 4.60\. That 0.16-point spread is one of the most striking data points in the entire analysis. Masks and Towelettes lead at 4.60, consistent with their role as high-intensity, visible-results products. Roll-on, Pencil, and Clay formats cluster around 4.50–4.55\. Core formats - Cream, Liquid, and Serum - the most formulation-complex products in the catalog - hold steady between 4.44 and 4.48. This kind of consistency doesn't happen by accident. It reflects decades of formulation investment, a quality control apparatus that scales across hundreds of SKUs, and an extensive understanding of what mass-market beauty consumers expect at each price point. ## What Actually Drives Purchase Decisions: Review Text Mining Studying the language across millions of reviews reveals three distinct clusters of purchase motivation and tells a story about how L'Oréal competes across both utilitarian and affective dimensions. ![Studying the language across millions of reviews reveals three distinct clusters of purchase motivation and tells a story...](https://www.blog.datahut.co/content/images/2026/07/chart_6-1-2.webp) Three Layers of the Purchase Decision The aesthetic hook - Color (35K+ mentions). Color is the [#1](https://www.blog.datahut.co/blog/hashtags/1) driver of conversions across the catalog, reflecting L'Oréal's dominance in hair dye, cosmetics, and tinted skincare. Customers are buying a visible transformation - a specific look - rather than just a functional treatment. This is the "glow factor" that the brand has built its marketing identity around. The sensory moat - Fragrance and Feel (19K+ mentions). While clinical rivals often win on "fragrance-free" safety messaging, L'Oréal wins on sensory pleasure. Fragrance functions as a loyalty driver - it creates a brand association that is difficult for clinical-positioned products to replicate. This confirms the brand's long-standing "salon experience at home" positioning. The performance baseline - Effectiveness and Quality (65K+ combined). The most important insight here is that aesthetic results and functional performance are not substitutes for each other - they stack. Products that deliver visible color results and clinical effectiveness generate the most reviews and the highest ratings. Neither dimension alone is sufficient. ## What This All Means: The L'Oréal Formula L'Oréal's Amazon playbook is not about hero products, premium pricing, or any single strategic lever. It is built on three simultaneous commitments: Coverage. Show up for nearly every beauty search query across 27 product forms. Leave no niche so underserved that a competitor can build an alternative brand around it. Consistency. Hold ratings above 4.40 across almost every product form, at every price tier. The 0.16-point rating range across 300 SKUs and 27 forms is proof that quality at scale is achievable - and that L'Oréal has systematized it. Selective aggression. Use deep discounts surgically on specialty SKUs - Masks at 55%, Kits at 36% - to pull shoppers into the wider catalog, while protecting margins on the core bestsellers that actually pay the bills. [![Selective aggression. Use deep discounts surgically on specialty SKUs - Masks at 55%, Kits at 36% - to pull shoppers into...](https://www.blog.datahut.co/content/images/2026/07/chart_7-2.webp)](https://www.datahut.co/contact?ref=blog.datahut.co) ## Frequently Asked Questions What is L'Oréal Paris's primary strategy for dominating Amazon? L'Oréal dominates through "Mass Variety" and "Portfolio Resilience" — prioritizing unmatched breadth and volume with 300 active products across 27 distinct product forms. The goal is to have a L'Oréal product present for every possible beauty search query. Which product forms drive the most customer engagement for L'Oréal on Amazon? Liquid and Cream forms are the primary drivers — Liquid leads with over 851,000 reviews across 97 SKUs, followed by Cream with over 728,000 reviews across 85 SKUs. However, Aerosol formats punch above their weight with roughly 30,700 reviews per SKU. How does L'Oréal use discounting on Amazon? L'Oréal takes a surgical approach: the overall average discount is 15.23%, but specialty items such as Masks (55% discount) and Kits (36% discount) receive deep cuts to drive trial and pull shoppers into the wider catalog. Core bestsellers are discounted more moderately to protect margins. What price range offers the best-rated L'Oréal products on Amazon? The mid-range price band ($15–$30) holds the highest average rating at 4.48\. Customers perceive L'Oréal's more advanced serums and treatments in this tier as the best equilibrium of performance and value. Does L'Oréal depend on a single hero product? No. The top single product holds just 3.49% of all reviews, and the top 10 products combined account for only 24.09% of total engagement. This distributed review structure makes the brand extremely resilient and difficult to displace through category-level competition. What are the top three drivers of purchase decisions, based on review analysis? Color (35,436 mentions), Effectiveness (32,967 mentions), and Quality (32,754 mentions). Color is the primary conversion hook — customers are buying a visual transformation. Effectiveness and Quality act as essential validation that those aesthetic results are backed by real performance. ### How to Bypass Browser Fingerprinting While Web Scraping URL: https://www.blog.datahut.co/post/bypass-browser-fingerprinting/ Last updated: 2026-07-23T07:48:29.000Z Ever wondered how a website seems to know you're a bot, even when you've changed your IP and rotate your user-agent? The answer is usually browser fingerprinting a set of techniques websites use to quietly collect dozens of details about your browser and device (your browser type, operating system, time zone, language, screen size, graphics hardware, and more), combine them, and turn them into a single identifying signature. That signature is your "fingerprint," and it sticks to you across requests even when cookies are cleared. For anyone scraping at scale, fingerprinting is the main thing standing in the way. A scraper that looks even slightly off gets flagged, and once it's flagged, the blocks follow. This guide breaks down how fingerprinting works, the common techniques behind it, and the practical ways to get past each one. ## How Browser Fingerprinting Works The idea is simple. When you visit a site, scripts on the page (and checks on the server) gather a long list of attributes about your setup things like your user-agent string, installed fonts, screen resolution, and how your device renders graphics. Individually, none of these is unique. But bundle enough of them together and hash the result, and you get a signature that's distinct enough to tell one visitor from another. Websites compare that signature against what they'd expect from a normal human visitor. If something doesn't add up say, a browser claiming to be Chrome on Windows while showing signals that look nothing like it the site flags the visitor as automated and starts blocking. The key thing to understand is that fingerprinting works in layers, and the layers check each other. It isn't enough to fake one detail convincingly. Every detail has to agree with every other one. That's what makes modern fingerprinting so hard to beat, and it's the thread running through everything below. Modern sites check your scraper across several independent layers. Get every layer perfect but one, and the contradiction is what gives you away. ## Common Techniques Used in Browser Fingerprinting and How Bypass Browser Fingerprinting works ![Common Techniques Used in Browser Fingerprinting and How Bypass Browser Fingerprinting works](https://www.blog.datahut.co/content/images/2026/07/chart_1-1-2.webp) ### The User-Agent and HTTP Headers The user-agent is the simplest signal of all , a short text string your browser sends with every request that announces what browser and operating system you're using. It's the first thing a site looks at, and an odd or outdated user-agent is an easy way to get flagged. How to handle it: Set your user-agent to match a common, current browser, and rotate through a pool of realistic ones rather than reusing a single string. But here's the part most guides skip rotating user-agents alone barely helps anymore. A user-agent that claims to be Chrome means nothing if the rest of your fingerprint contradicts it. Treat it as table stakes, not a solution. ### Canvas Fingerprinting Canvas fingerprinting is one of the cleverest techniques. The site asks your browser to draw a hidden image or piece of text, then reads back the exact pixels it produced. Because the result depends on your graphics card, drivers, and how your operating system renders fonts, different machines produce subtly different output a bit like handwriting. Visit again and your "handwriting" should look the same, and that consistency is what identifies you. How to handle it: The instinct is to scramble your canvas output so it looks different every time. Resist it that backfires. A real person's machine produces the same canvas result on every visit, so constant randomness actually makes you stand out more, not less. The better approach is a consistent, realistic profile where your canvas output matches the device you're claiming to be. Specialised stealth browsers like Camoufox handle this automatically. ### WebGL Fingerprinting WebGL is closely related to canvas, but instead of measuring how your browser draws, it asks your graphics hardware about itself directly the make and model of your GPU, its capabilities, and its limits. Since these vary widely across devices, they make for a strong identifier. How to handle it: The same rule applies. Don't feed it fake or random values; report hardware details that are consistent with the rest of your profile. A scraper that claims a high-end GPU while everything else looks like a budget laptop is easy to catch. ### TLS Fingerprinting (the Handshake) This is the layer most scrapers overlook, and it's the one that gets them blocked before the page even loads. Before any content is exchanged, your connection performs a quick technical "handshake," and real browsers do this in a very specific, recognizable way. Security systems read that handshake using methods known as JA3 and JA4 and compare it to what your scraper claims to be. The trap: if you build your scraper with a standard tool like Python's requests library, its handshake looks unmistakably automated, even while your user-agent insists you're a browser. That single contradiction is enough. How to handle it: You can't fix this by changing headers, because the handshake happens beneath your code. The practical fix is a tool built to imitate a real browser's handshake the most popular being curl\_cffi, which copies the exact way browsers like Chrome and Firefox introduce themselves, without the overhead of running a full browser. For most projects, this is the single highest-impact change you can make. It's also refreshingly simple in practice a single impersonate argument does the work: ``` from curl_cffi import requests # Introduce the connection exactly like a real Chrome browser would response = requests.get( "https://example.com", impersonate="chrome120" ) print(response.status_code) ``` Run that same request with a standard library and the handshake gives you away instantly; run it with the line above and it matches a genuine Chrome connection at the network level. ### Behavioural Fingerprinting Once the technical details check out, sites watch how you actually behave. Humans move a mouse in loose, looping curves, scroll at uneven speeds, and pause for slightly different lengths each time. Bots tend to move in straight lines and act with perfect, repetitive timing. Scripts record all of this and flag anything that looks too mechanical to be human. ![Once the technical details check out, sites watch how you actually behave. Humans move a mouse in loose, looping curves,...](https://www.blog.datahut.co/content/images/2026/07/chart_3-4.webp) How to handle it: Build human-like behavior into your automation — add varied delays between actions, move the cursor along natural curves rather than straight lines, and avoid perfectly regular timing. The goal is to look a little imperfect, the way real people are. ### Browser Leaks (WebRTC and DNS) Even a flawless disguise fails if your real location slips out a side door. Two leaks cause most of the trouble. WebRTC, a browser feature meant for video calls, can quietly expose your true IP address even when you're behind a proxy. A DNS leak happens when your requests to look up website addresses get routed through your real internet provider instead of your proxy, revealing where you actually are. How to handle it: Make sure WebRTC and DNS traffic are forced through your proxy, and that your apparent time zone and language match the location your proxy is exiting from. One inconsistency here gives away everything else. (Note: simply switching WebRTC off entirely is itself a small giveaway, since normal browsers leave it on — routing it through the proxy is cleaner.) ## Which System Are You Up Against? Worth a quick mention: the major anti-bot providers Akamai, Cloudflare, DataDome, Kasada, PerimeterX (now HUMAN), and F5 Shape Security each leave their own telltale signs, and each calls for a different approach. Some focus on behavior, some on the handshake, some rotate their defenses every minute. Identifying which one is guarding a site, and tailoring the strategy to it, is a big part of what separates a scraper that works from one that gets blocked and it's the kind of know-how we keep in-house and bring to client work rather than spell out for competitors. ## Best Practices to Bypass Browser Fingerprinting Pulling it all together, here's a practical checklist: ![Pulling it all together, here's a practical checklist:](https://www.blog.datahut.co/content/images/2026/07/chart_4-4.webp) 1. Match, don't just rotate. Your user-agent, headers, handshake, and hardware signals all need to describe one consistent, believable device. Consistency beats variety. 2. Use a browser-grade connection. For API-style scraping, replace standard libraries with a tool like curl\_cffi so your handshake matches a real browser. 3. Use realistic profiles, not random noise. Lean on stealth tools (such as Camoufox) that generate consistent, plausible device fingerprints, instead of scrambling your canvas or WebGL output. 4. Choose good proxies and bind everything to them. Route traffic through quality residential or mobile proxies, with time zone, language, and DNS all tied to the proxy's location. 5. Plug WebRTC and DNS leaks so your real IP and location never slip out. 6. Behave like a human. Add natural delays, irregular timing, and non-linear mouse movement on any site that tracks behavior. 7. Don't over-disable. Turning features off entirely (WebRTC, JavaScript) can be as suspicious as leaving them misconfigured. Aim for normal, not absent. 8. Test before you deploy. Free tools like Browserleaks, FingerprintJS, and CreepJS let you see your scraper's fingerprint the way a website would, so you can catch leaks early. They're not as sharp as commercial anti-bot systems, but they'll catch the obvious mistakes. 9. Keep updating. Browsers and anti-bot systems change constantly. A setup that's invisible today can be flagged next month, so treat this as ongoing maintenance, not a one-time job. ## Wrapping Up Browser fingerprinting is a moving target, and no single trick defeats it. The websites worth scraping run layered defenses that cross-check each other, and the only reliable way through is a setup where every signal handshake, hardware, behavior, and network tells the same consistent story, maintained as the rules keep shifting. That last part is the real challenge. Keeping a scraping operation invisible isn't a configuration you finish; it's a continuous game of cat and mouse against vendors who update their defenses weekly. It's also why a lot of teams decide it isn't worth doing in-house. That's the gap Datahut fills. We run managed web scraping as a service — our team handles the handshakes, the realistic profiles, the proxy and network hygiene, and the human-like behavior, and we keep all of it current as anti-bot systems evolve. You just receive clean, structured data on a schedule, without your engineers spending their week fighting blocks. If that sounds better than maintaining it yourself, [let's talk](https://www.datahut.co/?ref=blog.datahut.co) tell us what data you need and at what scale, and we'll handle the part that keeps breaking. ### FAQ What is browser fingerprinting? 1. Can browser fingerprinting be bypassed? Yes, but not by hiding a single detail. Because the layers cross-check each other, the only reliable approach is to present a fully consistent profile — matching handshake, hardware, behavior, and network signals — and keep it consistent over time. That's why it takes ongoing engineering rather than a one-time setting. 1. Does clearing cookies stop browser fingerprinting?No. Fingerprinting was designed specifically to work without cookies, measuring intrinsic properties of your browser and device. Clearing cookies, using incognito mode, or blocking tracking scripts has little effect on the fingerprint itself. 2. Does rotating user agents or IP addresses prevent fingerprinting?On their own, not really — and this is the most common mistake. Rotating user-agents and IPs was enough years ago, but modern systems check far deeper signals like TLS and hardware. A rotated user-agent that doesn't match the rest of your fingerprint can actually make you easier to spot. 3. What's the best tool to avoid fingerprinting when web scraping?It depends on the layer. For API-style scraping, curl\_cffi mimics a real browser's TLS and HTTP/2 handshake without the overhead of a full browser. For JavaScript-heavy sites, stealth browsers like Camoufox generate consistent, realistic device profiles. Most serious setups combine both with quality residential proxies. 4. Is bypassing browser fingerprinting to scrape data legal?The tools themselves are legal and widely used. Whether a given scrape is lawful depends on what you collect and how the site's terms, its robots directives, and applicable data-protection laws like GDPR not the tool you use. Collecting publicly available data responsibly, without harvesting personal data or overloading servers, is the standard to aim for. (This isn't legal advice.) Datahut is a managed web scraping and data-as-a-service company that delivers clean, structured web data to teams who would rather analyze data than fight to collect it. ### Zepto Fruits & Vegetables Data Analysis: 52% of Listings Are Out of Stock Across 6 Cities URL: https://www.blog.datahut.co/post/zepto-fruits-vegetables-data/ Last updated: 2026-09-07T09:42:52.000Z We scraped Zepto's entire fruits and vegetables catalogue across Delhi, Mumbai, Bangalore, Chennai, Kolkata, and Kochi - 1,035 listings covering 417 unique products. The top findings: - 52% of all listings were out of stock at the time of collection. Guava (83%), dragon fruit (77%), and kiwi (71%) were the worst-hit categories. - Delhi has the worst availability of any city , 71% of listings out of stock. Chennai has the deepest catalogue and the best availability. - Mango is Zepto's largest fresh produce category with 127 SKUs, but variety counts range from 21 in Bangalore down to 13 in Kolkata. - The average discount is 27%, rising to 41% on stone fruits , a sign of clearance pricing, not strategy. - Imported produce sells at a 61% price premium (₹157 average vs ₹98 for local). - "Immunity" is the most-used health tag (185 products), followed by "Gut Health" (148). India's quick commerce sector has crossed [$10 billion in GMV with over 30 million monthly users](https://redseer.com/articles/quick-commerce-indias-retail-darling-or-profit-mirage/?ref=blog.datahut.co) — yet our data shows half its fresh produce shelf is empty at any given moment. For sellers, each of these numbers is an actionable gap: competitor stockouts are open demand windows, under-assorted cities are expansion opportunities, and deep discounting signals weak demand forecasting you can out-execute. Here is the full breakdown. This is part of our ongoing series of catalogue-level data analyses — see our [five-day study of Carrefour UAE's vegetables category](https://www.blog.datahut.co/post/carrefour-uae-web-scraping-case-study/) and our [full-catalogue analysis of Cetaphil on Amazon](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/). ## How much of Zepto's fresh produce catalogue is out of stock? ![How much of Zepto's fresh produce catalogue is out of stock?](https://www.blog.datahut.co/content/images/2026/07/chart_1-4-2.webp) About 52% of Zepto's fruit and vegetable listings were out of stock across the six cities we analysed. This is not a minor fulfilment issue — it is a structural pattern that repeats across categories. Vegetables are far more stable — banana sits at 24% OOS, apple at 31%, root vegetables at 40%. What this means for sellers: Every time a competitor goes out of stock, there is an open window for your listing to capture demand. But you can only act on that window if you know it exists. Manually checking competitor listings across six cities every day is not realistic. Automated stock monitoring is. Equally important: your own OOS events are invisible to you unless you are tracking them. If your Alphonso listing went out of stock last Tuesday afternoon in Delhi, do you know how long it stayed that way? Do you know whether your competitors were in stock during that window? ## Which city has the best and worst stock availability on Zepto? Chennai has the best availability (122 of 186 listings in stock) and the widest assortment. Delhi has the worst only 50 of 171 listings were available, a 71% out-of-stock rate in India's largest city. ![Chennai has the best availability (122 of 186 listings in stock) and the widest assortment. Delhi has the worst only 50 of...](https://www.blog.datahut.co/content/images/2026/07/image.png-1.webp) For a seller operating in Delhi, this is both a problem and an opportunity. The market has demand that is consistently unmet. A seller who can maintain reliable availability in Delhi even at a slight price premium has a structural advantage over everyone else on the shelf. What this means for sellers: City-level data is operationally critical. Your procurement, pricing, and listing strategy should differ by city based on local demand patterns, competitor density, and stock availability norms. A one-size-fits-all approach across six cities is leaving money on the table. ## Which city lists the most mango varieties on Zepto? Bangalore leads with 21 mango varieties, followed by Chennai and Mumbai at 20 each. Kolkata trails with 13\. Mango is the single largest category on Zepto's fresh produce section 127 SKUs across 6 cities, more than any other category. During peak season, the platform stocks everything from Alphonso and Kesar to Lalbagh, Kalapadi, Raspuri, and Jawahar Pasand. ![Bangalore leads with 21 mango varieties, followed by Chennai and Mumbai at 20 each. Kolkata trails with 13. Mango is the...](https://www.blog.datahut.co/content/images/2026/07/chart_2-3-2.webp) What this means for sellers: Assortment decisions should be driven by what is actually selling in your city — and what gaps competitors have not filled yet. A seller in Kolkata who adds 3 to 4 varieties that are doing well in Bangalore, before any local competitor does, owns that demand. The deeper question is: which varieties are going OOS fastest in your city? That tells you where unmet demand actually sits. ## How deep are discounts on Zepto's fresh produce? The average discount across the catalogue is 27% but stone fruits are discounted at 41%, dragon fruit at 38%, and avocado at 35%. ![The average discount across the catalogue is 27% but stone fruits are discounted at 41%, dragon fruit at 38%, and avocado at...](https://www.blog.datahut.co/content/images/2026/07/chart_3-3-2.webp) A 41% average discount on peach, plum, litchi, and cherry suggests sellers are discounting aggressively to move perishable seasonal stock before it expires. This is not a pricing strategy. It is a clearance problem. What this means for sellers: The sellers who discount the least while maintaining availability are the ones with better demand forecasting. If you know a litchi season typically creates a demand spike in week three and a surplus by week five, you can plan procurement and pricing accordingly. That knowledge comes from historical data - not intuition. ## How much more does imported produce cost on Zepto? Imported produce sells at a 61% premium: ₹157 average sale price versus ₹98 for locally sourced produce, across 78 imported listings. The top-priced items in the entire catalogue: ![The top-priced items in the entire catalogue:](https://www.blog.datahut.co/content/images/2026/07/chart_4-2-2.webp) The premium segment is real and it is being served. But imported and premium listings also carry higher OOS risk they are harder to replenish quickly, and a stockout on a ₹1,000+ product is a significant lost sale. What this means for sellers: If you are in the premium segment, pricing intelligence matters more, not less. Are you priced correctly relative to what the market is bearing this week? Is a competitor undercutting you on Alphonso in Mumbai while you are priced at a premium in Chennai? Only [systematic competitor price tracking](https://www.blog.datahut.co/post/how-to-leverage-web-scraping-to-create-a-competitor-price-monitoring-strategy/) can answer these questions. ## What health tags does Zepto use on fresh produce? "Immunity" is the most common health tag (185 products), followed by "Gut Health" (148) and "Heart Health" (83). !["Immunity" is the most common health tag (185 products), followed by "Gut Health" (148) and "Heart Health" (83).](https://www.blog.datahut.co/content/images/2026/07/chart_5-3-2.webp) This hierarchy is not accidental - it reflects consumer demand signals Zepto has already acted on in its catalogue curation. What this means for sellers: If your listings are not tagged with the right health attributes, you are losing discoverability. Consumers filtering by "Immunity" or "Gut Health" will not find your product even if it belongs there. Catalogue optimisation - ensuring your product attributes, tags, and descriptions match how the platform surfaces results - is a competitive lever most sellers underestimate. ## What kind of data would actually help you compete? The insights above come from a single scrape of Zepto's catalogue. A one-time snapshot is interesting. What sellers actually need is this data running continuously: - [Daily price tracking](https://www.datahut.co/solutions/ecommerce-web-scraping?ref=blog.datahut.co) across competitors and cities, so you know when to hold your price and when to adjust - Stock availability monitoring to catch competitor OOS windows in near real-time - Assortment gap analysis to identify categories where demand exists but supply is thin - Historical trend data to anticipate seasonal demand spikes before they happen rather than reacting after the fact This is exactly what Datahut does. We are a managed data services company that builds custom data pipelines from platforms like Zepto, Blinkit, Swiggy Instamart, and others - delivering clean, structured, ready-to-use data directly to your operations team. You do not need to scrape anything yourself. You do not need engineers to build or maintain the pipeline. You tell us what data you need, from which platforms and cities, at what frequency and we deliver it. ## The Bottom Line Quick commerce is not slow commerce made faster. With the channel [projected to grow from $4 billion to over $25 billion in GMV by 2030](https://redseer.com/reports/reinventing-packaged-fb-with-quick-commerce/?ref=blog.datahut.co), the pace at which prices change, stock moves, and new SKUs appear means sellers running on gut feel and weekly spreadsheet reviews are operating blind. The 52% out-of-stock rate in this data is not just a Zepto operations problem. It is a signal that many sellers on the platform do not yet have the data infrastructure to match supply to demand in real time. That gap is a competitive advantage for the sellers who close it first. Interested in a custom data feed for your Zepto operations? Get in touch with the Datahut team to discuss what a data pipeline would look like for your business. ## Frequently Asked Questions 1. What percentage of Zepto's fruits and vegetables are out of stock? Based on a Datahut analysis of 1,035 listings across 6 Indian cities, approximately 52% of Zepto's fruit and vegetable listings were out of stock, with exotic fruits like guava (83%) and dragon fruit (77%) hit hardest. 1. Which Indian city has the worst stock availability on Zepto? Delhi, with a 71% out-of-stock rate, only 50 of 171 fresh produce listings were available. Chennai had the best availability with 122 of 186 listings in stock. 1. How much does imported produce cost on Zepto compared to local produce? Imported produce averages ₹157 per listing versus ₹98 for local produce — a 61% price premium. The most expensive item found was a Ratnagiri Alphonso mango combo at ₹1,402. 1. How can Zepto sellers monitor competitor prices and stock levels? Manually checking listings across cities daily is impractical at scale. Managed data services like Datahut deliver automated daily feeds of competitor prices, stock status, and assortment changes from Zepto, Blinkit, and Swiggy Instamart without sellers needing in-house scraping infrastructure. About the data: This analysis is based on a scrape of Zepto's fruits and vegetables catalogue across Delhi, Mumbai, Bangalore, Chennai, Kolkata, and Kochi. Data was collected from the platform's [public-facing product pages](https://www.blog.datahut.co/post/is-web-scraping-legal/) and reflects listing information at the time of collection, including product names, prices, stock status, seller details, and product attributes. ### How to Scrape Oakley's Eye-wear Data? URL: https://www.blog.datahut.co/post/how-to-scrape-oakley-s-eye-wear-data/ Last updated: 2026-09-07T09:42:55.000Z Think of having a super-fast assistant who can explore websites and pick out only the things you need. That’s what [web scraping](https://www.blog.datahut.co/post/web-scraping-vs-web-crawling-2026/) does - it saves you from the boring job of copying and pasting information one by one. It’s a handy skill, especially when you want to collect a lot of data quickly and without stress. In this blog, we’ll walk through a fun and useful project [where we collect data](https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/) about glasses from[ Oakley’s official website](https://www.oakley.com/en-us?ref=blog.datahut.co). You’ll see how we set everything up, how we go step by step to collect the data, and how we clean it to make it ready for use. By the end, you’ll understand how a real web scraping project works from start to finish. Now, Oakley’s website isn’t your typical web page where everything shows up right away. Instead, the website loads a lot of its content on the go, which means it uses background processes (called JavaScript, though we won’t dive deep into that here). Because of this, scraping Oakley’s website is a bit more challenging than usual. Also, the website doesn’t really like bots or automated tools grabbing its data, so it has a few roadblocks to stop that from happening. But don’t worry - we’ll show you how to work around these challenges in a smart way using the right tools. We’ll break the whole task into two clear steps: 1. First, we’ll use a tool called [Playwright](https://www.blog.datahut.co/post/scraping-amazon-reviews-playwright-python/) to go through Oakley’s main pages and collect links to each product. Think of it like making a list of all the glasses available on the site. 2. Next, we’ll visit each of those links. For every single product page, [we’ll again use Playwright to open the page](https://www.blog.datahut.co/post/scraping-decathlon-using-playwright-in-python/). Then, we’ll use another tool called [Beautiful Soup](https://www.blog.datahut.co/post/top-5-open-source-web-scraping-frameworks-and-libraries/) to pull out all the important details—like the name, price, and other information about the glasses. In the next sections, we’ll break down each part of the process - Scrape Oakley's Eye-wear. You’ll see the code, how it works, and some tips to handle common issues. ## Scrape Oakley's Eye-wear Products Urls [The Python script we’ve built](https://www.blog.datahut.co/post/how-to-scrape-zepto-s-fruits-and-vegetables-data-using-python/) is designed to collect product details from Oakley’s website. These include sunglasses, prescription sunglasses, and eyeglasses. It carefully goes through each category, grabs the links for all the products, and saves those links into a small local database called SQLite (just think of it as a digital notebook where data is saved neatly). Since Oakley’s website loads some content in the background and doesn’t show everything all at once, we’ve used a few smart tricks to handle that. The script behaves in a way that’s similar to how a real person would browse the site—slow and steady—so it doesn’t raise any red flags. It also has backup steps in case something goes wrong, helping it retry without crashing. To keep things running quickly and smoothly, the script uses a method that allows it to do many small tasks at the same time. It also uses a browser tool that lets it "click around" and "scroll" on web pages, just like you would. The script follows a clear plan—it visits one category at a time, gathers the data, and moves on to the next. It takes short breaks between actions and changes the way it looks to the website (by switching user agents, which basically means it pretends to be a different browser or device) to avoid getting blocked. All of this helps make the scraping process smoother and safer. ### Imports ``` import asyncio from playwright.async_api import async_playwright from bs4 import BeautifulSoup import random import time import sqlite3 ``` Before our scraper can start collecting data, we need to bring in a few important tools—just like preparing everything before starting a task. Here’s what each tool does: - asyncio helps the script run more smoothly by allowing it to handle many small tasks at once, instead of waiting for one to finish before starting another. - [playwright](https://playwright.dev/python/?ref=blog.datahut.co) is what lets our script open and control a web browser. It can click on things, scroll through pages, and even wait for content to load—just like a person would. - [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co) helps us read the web page’s content and find the exact details we want, like product names or prices. - random and time are used to make our scraper act more naturally. We use them to take small breaks and mix up its actions so it doesn’t look like a robot. - sqlite3 gives us a way to store the data we collect in a simple, organized format. It keeps everything saved and easy to use later. When we put all these tools together, our scraper can explore websites, gather the data we need, and save it safely—without standing out or getting blocked. ### User Agent Handling Functions ``` def load_user_agents(file_path): """ Load user agents from a file. This function reads a text file containing user agent strings, one per line. It filters out any empty lines and returns a list of valid user agent strings. Using multiple user agents helps in mimicking different browsers and potentially avoiding detection as a bot, which is crucial for web scraping tasks. Args: file_path (str): Path to the file containing user agents. Returns: list: A list of user agent strings. """ with open(file_path, 'r') as file: user_agents = [line.strip() for line in file.readlines() if line.strip()] return user_agents ``` Websites can sometimes tell when a tool or script—like our scraper—is visiting instead of a real person. One way they do this is by checking something called a user agent. This is a small piece of information your browser shares when it visits a site. It tells the website what kind of device and browser you're using. To avoid being noticed, our scraper changes its user agent regularly. That’s where the load\_user\_agents function comes in. It opens a file that contains many different user agent strings—these are just text versions of different browser identities. The function reads them all and keeps them ready. ``` def get_random_user_agent(user_agents): """ Select a random user agent from the provided list. This function is used to randomise the user agent for each request. By using different user agents, the script can simulate requests coming from various browsers and devices. This randomization helps in distributing the requests and potentially reducing the chances of being blocked by the website due to suspicious activity patterns. Args: user_agents (list): A list of user agent strings. Returns: str: A randomly selected user agent string. """ return random.choice(user_agents) ``` The get\_random\_user\_agent function is the one that actually picks a user agent from the list we loaded earlier. Every time our scraper visits the website, it chooses one of those user agents at random. This means that each visit can look slightly different—sometimes like it's coming from a phone, other times from a different browser or device. This simple trick helps our scraper stay under the radar and avoid being blocked. By using all these functions together, our scraper doesn't stand out. It behaves more like regular web traffic, making it harder for websites to tell that it's an automated tool collecting data. ### Main Scraping Function ``` async def scrape_product_urls(category_url, category_name, user_agents, retries=3): """ Scrape product URLs from a given category page. This function is the core of the scraping process. It uses Playwright to navigate to the category URL and interact with the page. The function simulates scrolling and clicking the 'Load More' button to ensure all products are loaded. It then uses BeautifulSoup to parse the HTML and extract product URLs. The function includes error handling and retry logic to deal with potential network issues or anti-scraping measures. It's designed to be resilient and can retry the scraping process multiple times if errors occur. Args: category_url (str): The URL of the category page to scrape. category_name (str): The name of the category being scraped. user_agents (list): A list of user agent strings to use for requests. retries (int, optional): Number of retry attempts in case of failure. Defaults to 3. Returns: list: A list of dictionaries containing product URLs and their categories. """ product_urls = [] for attempt in range(retries): try: async with async_playwright() as p: browser = await p.chromium.launch(headless=False, args=["--disable-http2"]) user_agent = get_random_user_agent(user_agents) context = await browser.new_context(user_agent=user_agent) page = await context.new_page() print(f"Navigating to {category_url} with User-Agent: {user_agent}") await page.goto(category_url, timeout=120000) while True: await page.evaluate('window.scrollTo(0, document.body.scrollHeight)') await page.wait_for_timeout(5000) load_more_button = await page.query_selector( '#skipToMainContent > div.replaceProducts > div > div > div.lazy-load-pagination > div > a' ) if load_more_button: is_visible = await load_more_button.is_visible() if is_visible: await load_more_button.click() print("Clicked 'Load More' button.") await page.wait_for_timeout(10000) else: print("'Load More' button is not visible.") break else: print("No 'Load More' button found.") break html = await page.content() soup = BeautifulSoup(html, 'html.parser') base_url = "https://www.oakley.com/" footer_divs = soup.find_all("div", class_="prod-tile_footer") product_links = [ {"url": base_url + div['data-href'], "category": category_name} for div in footer_divs if 'data-href' in div.attrs ] await browser.close() return product_links except Exception as e: print(f"Error on {category_url}: {e}") if "ERR_HTTP2_PROTOCOL_ERROR" in str(e): print("HTTP/2 Protocol Error: Retrying with random delay...") time.sleep(random.uniform(5, 15)) else: print(f"Retrying ({attempt+1}/{retries})...") time.sleep(5) return product_urls ``` This part of the scraper is where most of the action happens. Here's what it does: First, it picks a random user agent to stay hidden and then visits a product category page on the website. Once there, it scrolls down the page—just like a real person looking through all the items. While scrolling, it watches for a "Load More" button. If it sees one, it clicks it to load more products. It keeps doing this until the page shows everything and there’s nothing left to load. Once the page is fully loaded, the scraper grabs all the content and uses BeautifulSoup to look through it and collect the links for each product. If something goes wrong—maybe the page didn’t load properly or there was an error—it doesn’t quit right away. The function will try again, up to three times, to make sure it gets the job done. ### Category Scraping Coordinator ``` async def scrape_all_categories_sequentially(categories, user_agents): """ Scrape product URLs from all specified categories sequentially. This function acts as a coordinator for the scraping process. It iterates through each category in the provided list and calls the scrape_product_urls function for each one. The sequential approach ensures that categories are scraped one at a time, which can help in managing resources and avoiding overwhelming the target website with simultaneous requests. This function aggregates the results from all categories into a single list, providing a comprehensive collection of all scraped product URLs across all specified categories. Args: categories (list): A list of dictionaries containing category URLs and names. user_agents (list): A list of user agent strings to use for requests. Returns: list: A list of dictionaries containing all scraped product URLs and their categories. """ all_products = [] for category in categories: print(f"Scraping category: {category['name']}") category_products = await scrape_product_urls(category['url'], category['name'], user_agents) all_products.extend(category_products) return all_products ``` This function helps the scraper visit different sections of the website—one at a time. Think of it like following a plan that lists all the product categories we want to check, such as sunglasses, prescription glasses, and more. It goes through each category in order. For every category, it calls the main function that does the scraping, waits for it to finish, and then moves on to the next one. By doing this step-by-step, we make sure we don’t overload the website or appear suspicious. Once all categories have been visited, it collects every product link found from each one and combines them into one big list. This gives us a full list of all the products we’re interested in—across all sections of the site ### Database Functions These functions are all about organizing and saving the information we collect. ``` def init_database(): """ Initialise the SQLite database and create the products table if it doesn't exist. This function sets up the SQLite database for storing the scraped product data. It creates a new database file if it doesn't exist, or connects to an existing one. The function also ensures that the necessary table structure is in place by executing a CREATE TABLE IF NOT EXISTS query. This approach allows the script to be run multiple times without duplicating the table structure. The table is designed with an auto-incrementing primary key, a category field, and a unique URL field to prevent duplicate entries. Returns: sqlite3.Connection: A connection object to the SQLite database. """ conn = sqlite3.connect('oakley_products.db') cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS products ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT UNIQUE ) ''') conn.commit() return conn ``` The init\_database function gets everything ready for storing data. First, it checks if the database (where we’ll store everything) already exists. If it doesn’t, the function creates one. Then it checks if the right kind of storage space—called a table—is available in the database. If it’s not there, the function creates it. This makes sure we have the right setup before we begin saving product information. ``` def save_to_database(data, conn): """ Save the scraped product data to the SQLite database. This function is responsible for persisting the scraped data into the SQLite database. It iterates through the list of scraped products and attempts to insert each one into the database. The function uses parameterized queries to prevent SQL injection vulnerabilities. It also handles potential IntegrityError exceptions, which could occur if a duplicate URL is encountered (due to the UNIQUE constraint on the url field). This approach ensures that the database remains consistent and free of duplicates, even if the scraping the process is run multiple times or encounters duplicate products. Args: data (list): A list of dictionaries containing product URLs and their categories. conn (sqlite3.Connection): A connection object to the SQLite database. """ cursor = conn.cursor() for product in data: try: cursor.execute(''' INSERT INTO products (category, url) VALUES (?, ?) ''', (product['category'], product['url'])) except sqlite3.IntegrityError: print(f"Duplicate URL found: {product['url']}") conn.commit() print("Data saved successfully to the database.") ``` The save\_to\_database function handles the saving part. It takes all the product details we’ve collected and stores them neatly in the database. It’s smart enough to check if a product is already saved. If it finds that the same item is already there, it skips it—so we don’t end up saving the same thing twice. Together with the setup functions, this makes sure everything our scraper finds is stored safely, without any confusion or repetition. This way, we can easily come back and use the data whenever we need it. ### Main Function ``` async def main(): """ Main function to orchestrate the scraping process and database operations. This function serves as the entry point and coordinator for the entire scraping operation. It begins by loading the list of user agents, which will be used throughout the scraping process to vary the requests. It then defines the categories to be scraped, each with a specific URL and name. The function initiates the scraping process by calling scrape_all_categories_sequentially, which handles the actual data collection. Once the scraping is complete, the function initializes the database connection and saves the collected data. Finally, it ensures proper closure of the database connection. This structured approach allows for clear separation of concerns between data collection, storage, and overall process management. """ user_agents = load_user_agents('useragents.txt') categories = [ {"url": "https://www.oakley.com/en-us/category/sunglasses", "name": "Sunglasses"}, {"url": "https://www.oakley.com/en-us/category/prescription/sunglasses", "name": "Prescription Sunglasses"}, {"url": "https://www.oakley.com/en-us/category/prescription/eyeglasses", "name": "Prescription Eyeglasses"} ] all_products = await scrape_all_categories_sequentially(categories, user_agents) conn = init_database() save_to_database(all_products, conn) conn.close() ``` The main function is the part that brings everything together. It doesn’t do the scraping itself, but it tells all the other parts when to start and what to do. First, it loads the list of user agents—these are used to help our scraper move around the website without being noticed. Then it sets up the list of categories we want to scrape. Once everything is ready, it starts the scraping process by calling the function that goes through each category. While the scraping is happening, the main function just waits and watches. After all the data is collected, the main function takes over again. It tells the program to save the data into the database. Finally, it makes sure everything is closed properly, like turning off the lights after the job is done. ### Script Execution ``` if __name__ == "__main__": asyncio.run(main()) ``` Now we come to the last and simplest part—but also one of the most important. This is the "on switch" that kicks off everything we've built so far. The line if name == "\_\_main\_\_": is just a smart check. It asks, "Are we running this script directly, or is it being used somewhere else?" If the answer is yes—meaning we’ve run the script directly—it moves ahead and starts everything. What comes next is [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(main()). This is like hitting the play button for the whole operation. It tells [Python](https://www.python.org/?ref=blog.datahut.co), "Let’s start our main function and keep things moving efficiently." Since our scraper works on multiple tasks at once (like visiting pages, scrolling, loading content, and saving data), [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)() acts like a smart stage manager. It keeps all the moving parts in sync and ensures everything happens in the right order—without wasting time. So with this one simple line, we turn on our entire scraping system. From disguising itself to gathering product info to saving everything neatly in a database—this is where it all begins. ## Scraping Products Data Now that we've gathered all the product URLs from Oakley's sunglasses, prescription sunglasses, and eyeglasses sections, it's time to dive deeper. This next script is responsible for visiting each collected URL and extracting full product details such as name, price, description, features, and more. ### Import Section ``` import asyncio import sqlite3 from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` As we've explained before, we start by importing all the essential tools for our web scraper: asyncio, sqlite3, playwright, and BeautifulSoup. ### Database Setup Function ``` def setup_database(db_name): """ Set up the SQLite database for storing product information. This function performs the following tasks: 1. Connects to the specified SQLite database. 2. Checks if the 'scraped' column exists in the 'products' table and adds it if not present. 3. Creates a new 'scraped_products' table if it doesn't already exist. Args: db_name (str): The name of the SQLite database file. Returns: None """ conn = sqlite3.connect(db_name) cursor = conn.cursor() # Check if 'scraped' column exists in products table cursor.execute("PRAGMA table_info(products)") columns = [column[1] for column in cursor.fetchall()] # Alter products table to add scraped column if it doesn't exist if 'scraped' not in columns: cursor.execute(''' ALTER TABLE products ADD COLUMN scraped INTEGER DEFAULT 0 ''') # Create scraped_products table cursor.execute(''' CREATE TABLE IF NOT EXISTS scraped_products ( id INTEGER PRIMARY KEY, url TEXT, collection TEXT, title TEXT, sale_price TEXT, original_price TEXT, discount TEXT, number_of_colors TEXT, size TEXT, fit TEXT, bridge TEXT, light_transmission TEXT, light_conditions TEXT, information_notice TEXT, product_code TEXT, size_ TEXT, description TEXT, content TEXT ) ''') conn.commit() conn.close() ``` The setup\_database function is the foundation for storing all the product data we scrape. Before any scraping begins, this function prepares our database to make sure everything is in place. One smart thing it does is check whether the 'scraped' column already exists in the database. If it doesn’t, the function adds it. This makes the setup flexible and backward-compatible—meaning, if you’ve used the database before, it won’t break or lose data when the structure changes. Instead, it smoothly updates the schema while keeping existing data safe. The most important part of this function is creating the scraped\_products table. This table acts like a blueprint, organizing how we store product information. It defines specific fields (or columns) for each attribute we plan to collect—such as name, URL, category, and so on. This structured setup allows us to scrape a wide range of product details and store them neatly in one place. In the end, this careful setup makes it easy to search, analyze, or even visualize the data once the scraping is done. ### Retrieve Unscraped URLs Function ``` def get_unscraped_urls(db_name): """ Retrieve URLs of products that have not been scrapped yet. Args: db_name (str): The name of the SQLite database file. Returns: list: A list of tuples containing (id, url) for unscraped products. """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute("SELECT id, url FROM products WHERE scraped = 0") urls = cursor.fetchall() conn.close() return urls ``` The get\_unscraped\_urls function plays a key role in making our scraping process efficient. Its job is to find out which products still need to be scraped. It does this by checking the database for any products where the 'scraped' flag is set to 0, meaning they haven’t been processed yet. This helps the scraper focus only on new or pending work, instead of going over the same products again and again. It saves time and system resources by skipping what’s already done. The function returns a list of tuples, with each tuple containing two important things: the product's ID and its URL. The URL tells the scraper where to go, and the ID helps us update the database later to mark the product as "scraped" once we're done. This setup is simple but very useful. It gives the scraper a clear checklist of what’s left to do, while also making it easy to keep track of progress and avoid duplication. ### Update Scraped Status Function ``` def update_scraped_status(db_name, product_id): """ Update the 'scraped' status of a product in the database. Args: db_name (str): The name of the SQLite database file. product_id (int): The ID of the product to update. Returns: None """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute("UPDATE products SET scraped = 1 WHERE id = ?", (product_id,)) conn.commit() conn.close() ``` The update\_scraped\_status function acts like an accountant for our scraper. After a product has been successfully scraped, this function updates the database to mark that product as "done." It does this by setting the 'scraped' flag to 1 for that specific product ID. This small action plays a big role in keeping the scraping process clean and organized—ensuring we don’t scrape the same product more than once. This update helps the scraper remember what’s already been covered, which is especially useful if the process is paused and resumed later. Even after multiple runs or interruptions, the scraper can pick up exactly where it left off without repeating any work. In short, it’s a simple but essential function for building a reliable and efficient scraping system that can scale to handle lots of data without getting confused or duplicating effort. ### Save Scraped Product Function ``` def save_scraped_product(db_name, product_data): """ Save the scraped product details to the 'scraped_products' table. Args: db_name (str): The name of the SQLite database file. product_data (dict): A dictionary containing the scraped product information. Returns: None """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(''' INSERT INTO scraped_products ( url, collection, title, sale_price, original_price, discount, number_of_colors, size, fit, bridge, light_transmission, light_conditions, information_notice, product_code, size_, description, content ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) ''', ( product_data['url'], product_data['collection'], product_data['title'], product_data['sale_price'], product_data['original_price'], product_data['discount'], product_data['number_of_colors'], product_data['size'], product_data['fit'], product_data['bridge'], product_data['light_transmission'], product_data['light_conditions'], product_data['information_notice'], product_data['product_code'], product_data['size_'], product_data['description'], product_data['content'] )) conn.commit() conn.close() ``` The save\_scraped\_product function is like the librarian of our scraper. Once the scraper has gathered all the details about a product, this function makes sure everything is neatly stored in the database. It uses something called a parameterized SQL query—which is a safe way to insert data. This prevents security issues like SQL injection and also makes sure each piece of data is stored in the right format. The function handles a wide range of product details—from basic information like the product’s URL and title, to very specific measurements like bridge width or light transmission. This shows how thorough our scraper is: it collects not just surface-level info, but deep product specs too. By storing all of this in a well-structured way, the function allows us to analyze, filter, or even visualize the data later with precision. It's an essential step that turns raw scraped data into organized, usable information. ### parse\_collection Function ``` def parse_collection(soup): """ Parse the collection name from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product pa+++++++++++++ge. Returns: str: The collection name if found, 'N/A' otherwise. """ collection_tag = soup.select_one("#pdhero > div.wrapper > form > div > div.pdp-area.pdp-sidebar.oo-w-100 > div > div.items-backBadge.slider-gallery-items__backBadge > div > p") if collection_tag: return collection_tag.get_text(strip=True) return 'N/A' ``` The parse\_collection function is like a label reader for our scraper. Its job is to find the name of the product’s collection—basically, the group or series the product belongs to. It does this by looking for a specific part of the product page using a CSS selector. If it finds that part, it grabs the text, trims any extra spaces from the beginning or end, and returns it. If the scraper doesn’t find the collection info on the page, it doesn’t panic—it just returns "N/A" to let us know that the collection name wasn’t available. This makes the scraper flexible and robust, even when some product pages are missing expected details. Next, we use the same method for product tittle, sale price, original price, discount, number of colors, size, the bridge, fit, light transmissions, light conditions, notice of information, product code, description and content. Take a look at the functions: ### parse\_product\_title Function ``` def parse_product_title(soup): """ Parse the product title from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The product title if found, 'N/A' otherwise. """ title_tag = soup.select_one('#pdhero > div.wrapper > form > div > div.pdp-sidebar-wrapper__mobile > div > h1 > span') if title_tag: return title_tag.get_text(strip=True) return "N/A" ``` ### parse\_sale\_price Function ``` def parse_sale_price(soup): """ Parse the sale price from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The sale price if found, 'N/A' otherwise. """ sale_price_tag = soup.select_one('[data-test="sale-price"]') if sale_price_tag: return sale_price_tag.get_text(strip=True) return "N/A" ``` ### parse\_original\_price Function ``` def parse_original_price(soup): """ Parse the original price from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The original price if found, the sale price if original price is not found. """ original_price_tag = soup.select_one('[data-test="original-price"]') if original_price_tag and original_price_tag.get_text(strip=True): return original_price_tag.get_text(strip=True) return parse_sale_price(soup) ``` ### parse\_discount Function ``` def parse_discount(soup): """ Parse the discount percentage from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The discount percentage if found, '0' otherwise. """ discount_tag = soup.select_one('[data-test="percentage-off"]') if discount_tag and discount_tag.get_text(strip=True): return discount_tag.get_text(strip=True) return "0". ``` ### parse\_number\_of\_colors Function ``` def parse_number_of_colors(soup): """ Parse the number of available colours from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The number of colours if found, 'N/A' otherwise. """ colors_tag = soup.select_one("#pdhero > div.wrapper > form > div > div.pdp-area.pdp-sidebar.oo-w-100 > div > div.pdp-filter-by-technology.pdp-filter-by-technology__abtest.nosize > div.oo-w-100.genericActionBox.o21_bg.thumbnails > div > div > div > h2 > span.oo-text.bold.o21_text-color2.colorLabel") if colors_tag: return colors_tag.get_text(strip=True) return "N/A" ``` ### parse\_size Function ``` def parse_size(soup): """ Parse the size information from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The size information if found, 'N/A' otherwise. """ size_tag = soup.select_one("#pdhero > div.wrapper > form > div > div.pdp-area.pdp-sidebar.oo-w-100 > div > div.product-cart-wrapper > div.sizeSelectorWrapper__mobile > div > div > div:nth-child(1) > span") if size_tag: return size_tag.get_text(strip=True) return "N/A" ``` ### parse\_fit Function ``` def parse_fit(soup): """ Parse the fit information from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The fit information if found, 'N/A' otherwise. """ fit_tag = soup.select_one("#pdhero > div.wrapper > form > div > div.pdp-area.pdp-sidebar.oo-w-100 > div > div.product-cart-wrapper > div.sizeSelectorWrapper__mobile > div > div > h2 > span.o21_text-normal") if fit_tag: return fit_tag.get_text(strip=True) return "N/A" ``` ### parse\_bridge Function ``` def parse_bridge(soup): """ Parse the bridge size information from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The bridge size if found, 'N/A' otherwise. """ bridge_tag = soup.select_one("#size-guide-panel > div.size-selection-box.form-field.select-size.hidden-md-down > div.fit-description > span:nth-child(2)") if bridge_tag: return bridge_tag.get_text(strip=True) return 'N/A' ``` ### parse\_light\_transmission Function ``` def parse_light_transmission(soup): """ Parse the light transmission percentage from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The light transmission percentage if found, 'N/A' otherwise. """ transmission_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-featandtechAccordion > div > div > div.lensDetailsAccordion.oo-flex > div.lensDetailsAccordion-section > div > ul > li:nth-child(1) > span.o21_text-bold", {'data-field': 'lightTransmission percentage'}) if transmission_tag: return transmission_tag.get_text(strip=True) return 'N/A' ``` ### parse\_light\_conditions Function ``` def parse_light_conditions(soup): """ Parse the recommended light conditions from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The recommended light conditions if found, 'N/A' otherwise. """ conditions_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-featandtechAccordion > div > div > div.lensDetailsAccordion.oo-flex > div.lensDetailsAccordion-section > div > ul > li:nth-child(2) > span.o21_text-bold", {'data-field': 'lightTransmission lightingCondition'}) if conditions_tag: return conditions_tag.get_text(strip=True) return "N/A" ``` ### parse\_information\_notice Function ``` def parse_information_notice(soup): """ Parse the information notice from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The information notice if found, 'N/A' otherwise. """ notice_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-featandtechAccordion > div > div > div.lensDetailsAccordion.oo-flex > div.lensDetailsAccordion-section > div > ul > li:nth-child(4) > span.o21_text-bold") if notice_tag: return notice_tag.get_text(strip=True) return "N/A" ``` ### parse\_product\_code Function ``` def parse_product_code(soup): """ Parse the product code from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The product code if found, 'N/A' otherwise. """ code_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-productInfoAccordion > div > div > div > h2.pdpBottom-productInfo-code.o21_text-color2.o21_text8.o21_text-medium.text-uppercase > span.o21_text-color1") if code_tag: return code_tag.get_text(strip=True) return "N/A" ``` ### parse\_description Function ``` def parse_description(soup): """ Parse the product description from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The product description if found, 'N/A' otherwise. """ description_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-productInfoAccordion > div > div > div > div.pdpBottom-productInfo-description.o21_text-color2.o21_text7.o21_text-medium") if description_tag: return description_tag.get_text(strip=True) return "N/A" ``` ### parse\_content Function ``` def parse_content(soup): """ Parse additional content from the product page. Args: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The additional content if found, 'N/A' otherwise. """ content_tag = soup.select_one("#pdhero > div.pdpBottom > div.pdpBottom-productInfoAccordion > div > div > div > div.accordionBox.oo-w-100.pdpAccordion.pdpBottom-productInfo-readmore.o21_bg > div > div") if content_tag: return content_tag.get_text(strip=True) return "N/A" ``` ### scrape\_product\_details Function ``` async def scrape_product_details(page, url): """ Scrape product details from a given URL using Playwright and BeautifulSoup. This function navigates to the product page, waits for the content to load, and then extracts various product details using BeautifulSoup. Args: page (Page): A Playwright page object. url (str): The URL of the product page to scrape. Returns: dict: A dictionary containing the scraped product details. """ await page.goto(url, timeout=60000) await page.wait_for_timeout(2000) # Wait for 2 seconds to ensure page content is loaded content = await page.content() soup = BeautifulSoup(content, 'html.parser') return { 'url': url, 'collection': parse_collection(soup), 'title': parse_product_title(soup), 'sale_price': parse_sale_price(soup), 'original_price': parse_original_price(soup), 'discount': parse_discount(soup), 'number_of_colors': parse_number_of_colors(soup), 'size': parse_size(soup), 'fit': parse_fit(soup), 'bridge': parse_bridge(soup), 'light_transmission': parse_light_transmission(soup), 'light_conditions': parse_light_conditions(soup), 'information_notice': parse_information_notice(soup), 'product_code': parse_product_code(soup), 'size_': parse_size_(soup), 'description': parse_description(soup), 'content': parse_content(soup) } ``` The scrape\_product\_details function is the core engine of the scraping process—it’s where the real work of collecting product information happens. This function is asynchronous, which means it can handle many tasks efficiently without waiting for one to finish before starting the next. It uses Playwright to control a web browser and visit a specific product page. Once it receives a browser page object and the product URL, it: 1. Navigates to the product page using page.goto(), with a generous timeout in case the page takes a while to load. 2. Waits for 2 seconds to give JavaScript content enough time to fully appear on the screen. 3. Extracts the page’s HTML content and passes it to BeautifulSoup, which helps break down the HTML into something Python can easily work with. Then, the function uses a set of parsing helper functions like parse\_title, parse\_collection, and others. Each one is in charge of pulling out a specific piece of data—like the product’s name, category, price, or size. All this data is then packed into a dictionary, which makes it easy to store and access each product's details later on. This structure keeps the scraping logic modular and clean: if one part of the webpage changes, you only need to update the relevant parsing function—not the entire scraping process. ### main Function ``` async def main(): """ Main entry point for the web scraping process. This function performs the following tasks: 1. Set up the database. 2. Retrieves unscraped URLs from the database. 3. Launches a Playwright browser instance. 4. Iterates through unscraped URLs, scraping product details for each. 5. Saves scraped data to the database and updates the scraped status. 6. Handles any errors that occur during scraping. 7. Closes the browser instance after scraping is complete. Returns: None """ db_name = "oakley_products.db" # Set up the database setup_database(db_name) # Get unscraped URLs from the database urls_to_scrape = get_unscraped_urls(db_name) async with async_playwright() as playwright: browser = await playwright.chromium.launch(headless=False) page = await browser.new_page() for product_id, url in urls_to_scrape: try: # Scrape product details product_details = await scrape_product_details(page, url) # Save scraped product details save_scraped_product(db_name, product_details) # Update scraped status update_scraped_status(db_name, product_id) print(f"Successfully scraped and saved product {product_id}") except Exception as e: print(f"Error scraping product {product_id} ({url}): {e}") await browser.close() ``` The main function acts as the conductor of the entire scraping process—it brings everything together and makes sure each part of the scraper performs its role correctly. Here’s what it does step-by-step: 1. Sets up the database using a helper function. It ensures that everything is ready to store the scraped data. 2. Gets a list of product URLs that haven’t been scraped yet. This smart step helps the scraper resume from where it left off—no wasting time on already-processed products. Next, it moves into the core scraping loop: 1. It launches a Playwright browser using a context manager (async with). This ensures that the browser and system resources are cleaned up properly when scraping is done. 2. For each unscraped product URL:It calls the scrape\_product\_details() function to extract information.Then it saves that data to the database.And finally, it marks the product as “scraped” to avoid reprocessing it later. This entire process is wrapped in a try-except block, meaning if one product fails to scrape (due to a timeout or broken page), it will log the error and move on without stopping the whole program. That’s important when you're dealing with a large list of products—you want the script to be robust and resilient. Finally, once everything is done, the function closes the browser and finishes the process. ### Script Execution ``` # Entry point for running the script if __name__ == "__main__": asyncio.run(main()) ``` The final section of the script is conditional code to check if the script is being run as a main program which is the same as we did in scraping urls part above which is the entry point the code. ## Conclusion Web scraping Oakley eyewear products highlights how automation can greatly simplify the extraction of structured information from online stores. By using Playwright, we were able to navigate dynamic web pages—ensuring that all product content was fully loaded—before passing the HTML to Beautiful Soup for data extraction. This automated approach eliminates the need for manual data collection and allows for efficient gathering of product names, prices, and specifications. Such a technique proves highly valuable for tasks like price comparison, market research, and inventory tracking, where access to large volumes of product data in a usable format is essential. However, web scraping also comes with important responsibilities. Since each website has its own structure and terms of service, it’s critical to [respect ](https://www.rfc-editor.org/info/rfc9309/?ref=blog.datahut.co)[robots.txt directives](https://www.rfc-editor.org/info/rfc9309/?ref=blog.datahut.co) and adhere to ethical scraping practices. This includes setting appropriate delays between requests, handling retries gracefully, and avoiding excessive strain on servers. This project demonstrates how Playwright and Beautiful Soup can be effectively combined to automate product data collection. It enables more accessible and scalable data-driven research, especially in the dynamic and competitive world of e-commerce. AUTHOR I’m Shahana, a Data Engineer at Datahut. I focus on designing intelligent data pipelines that transform complex, dynamic web data into clean, structured insights—empowering brands to make smarter decisions in fashion, eyewear, and e-commerce. At Datahut, we’ve spent over a decade helping companies leverage automation to streamline product tracking, competitor monitoring, and market research. In this blog, I guide you through how we used Playwright and Beautiful Soup to scrape Oakley eyewear products—extracting specifications, pricing, and availability from dynamic web pages with precision. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. ### Frequently Asked Questions (FAQs) ### 1\. Why did this project use both Playwright and Beautiful Soup instead of just one tool? Playwright and Beautiful Soup serve different purposes. Playwright is used to render JavaScript-heavy pages and interact with dynamic content, such as scrolling and clicking "Load More" buttons. Once the page is fully loaded, Beautiful Soup efficiently parses the HTML and extracts structured data like product names, prices, descriptions, and specifications. Combining both tools provides a reliable and scalable scraping workflow. ### 2\. Why can't Oakley's website be scraped using simple HTTP requests? Oakley's website relies heavily on JavaScript to load product listings and details dynamically. A simple HTTP request only retrieves the initial HTML and often misses content generated after the page loads. Browser automation tools like Playwright can execute JavaScript, making them ideal for scraping modern e-commerce websites. ### 3\. What product information can be extracted from Oakley's eyewear pages? The scraper can collect a wide range of product details, including: - Product name and collection - Sale and original prices - Discount percentage - Available colors - Size and fit information - Bridge width - Lens light transmission details - Product descriptions - Product codes - Additional technical specifications This data can be stored in a database for analysis, monitoring, or reporting. ### 4\. How does the scraper avoid getting blocked while collecting data? The scraper uses several techniques to mimic normal user behavior, including: - Rotating user agents - Adding delays between actions - Browsing categories sequentially - Handling retries when requests fail - Using a real browser through Playwright These practices help reduce the chances of triggering anti-bot systems while maintaining responsible scraping behavior. ### 5\. What are the practical uses of scraping Oakley eyewear data? Scraping eyewear product data can support various business and research applications, such as: - Competitor price monitoring - Product catalog analysis - Market research and trend tracking - Inventory and assortment analysis - E-commerce intelligence dashboards - AI and data science projects requiring structured product datasets By automating data collection, businesses can gain insights much faster than through manual research. ### How to Scrape Zepto’s Fruits and Vegetables Data Using Python? URL: https://www.blog.datahut.co/post/how-to-scrape-zepto-s-fruits-and-vegetables-data-using-python/ Last updated: 2026-09-07T09:42:57.000Z Have you ever tried copying [product details](https://www.blog.datahut.co/post/why-retailers-should-invest-in-web-scraping-product-matching-and-bi/) from a website by hand—one item at a time? If so, you probably know how slow and frustrating it can be. Now, imagine if a robot could do that for you - fast, accurately, and without complaining. That’s exactly what web scraping does. It’s like having a personal assistant that reads through web pages and pulls out the useful bits for you, turning messy website code into neat, usable data. In this blog, we’re going to explore a real [web scraping](https://www.blog.datahut.co/post/web-scraping-in-python/) project. Our goal? To collect product details from the fruits and vegetables section of Zepto, a popular online grocery delivery app in India. If you’ve been looking for a hands-on example of how to scrape a modern website, you’re in the right place. ## Why did we pick Zepto? Zepto is growing fast in the Indian market, and it offers a smooth online shopping experience. But there’s a twist—its website runs on JavaScript, which makes scraping a bit trickier. That’s actually a good thing here because it gives us the chance to learn how to [handle such sites](https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/) using the right tools. ### Our Plan to Scrape Zepto : Two Simple Steps We’ll break down the scraping process into two clear parts: 1. Link Collection – First, we’ll collect the URLs of all the product pages from the category listings. 2. Data Collection – Then, we’ll visit each product page and grab the details we need—like the name, price, discount, description, and whether it’s in stock. By following this step-by-step approach, our code will stay organized, and we’ll have an easier time spotting and fixing errors if anything goes wrong. ### The Tools We’ll Use Here’s what we’ll be working with: - [Playwright](https://playwright.dev/python/docs/intro?ref=blog.datahut.co) – Think of this as a tool that opens a browser and clicks around the website just like a real person. It’s especially helpful when dealing with websites that load content using JavaScript. - [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co) – This is a Python library that helps us read and pull out specific pieces of data from a web page once it's fully loaded. - [SQLite](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co) – A simple, file-based database where we’ll neatly [store all our scraped data](https://www.blog.datahut.co/post/how-to-scrape-data-from-booking-com/). It’s easy to use and doesn’t require any server setup. In the next sections, I’ll walk you through how we combine these tools to build a working scraping system. Along the way, I’ll share what worked, what didn’t, and what I learned—so you can avoid common mistakes and get better at this, one project at a time. ## Links Collection In this part of the project, we built a simple Python scraper to collect product links from Zepto’s fruits and vegetables section. As we told before, to make this work we used a few helpful tools: Playwright for loading the website like a real browser, BeautifulSoup for reading the page’s content, and SQLite for storing the links we collect. Here’s what the scraper actually does behind the scenes:Zepto loads more products as you scroll down the page. So instead of grabbing just what’s visible at first, our scraper scrolls through the page automatically—just like you would if you were browsing. As it scrolls, it collects all the product links one by one and saves them into a local database. These links will come in handy later when we need to visit each product page to collect detailed information. The code is structured in a neat and organized way. Each task—like fetching the page content, picking out the product links, saving them, handling errors, and even printing updates—is handled by its own function. This makes the code easier to follow, fix if anything goes wrong, and reuse in future projects. Now, let’s break the whole process down step by step and take a closer look at how each part works. ### Library Imports and Initial Setup ``` import sqlite3 import logging from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup import time from datetime import datetime # Configure logging logging.basicConfig(filename="scraper.log", level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") BASE_URL = "https://www.zeptonow.com" ``` This piece of code initializes everything that we need for our scraper to work. We begin by importing a couple of very useful Python librarie : sqlite3,logging captures events and errors when our scraper runs, Playwright, BeautifulSoup ,time which add pauses between actions to avoid overloading the website and date-time module will provide our data with time stamps. We also set up a logging system that can write detail to a file called scraper.log; each entry will include the date / time, along with timestamps, log type, and message — this will make it easier to trace our movements and troubleshoot if things go astray. Finally we defined a constant called BASE\_URL, in which we store an address for the website. We will use this to build the full product URLs by appending the shorter paths that we scrape from the website. ### Page Content Fetching ``` def fetch_page_content(url): """ Fetches the full page content using Playwright with incremental scrolling to load all dynamic content. This function launches a Chromium browser, navigates to the specified URL, and performs scrolling operations to ensure all lazy-loaded content is rendered before returning the page's HTML. Args: url (str): The URL to fetch content from. Returns: str or None: The HTML content of the page if successful, None otherwise. Raises: Various exceptions from Playwright which are caught and logged. Note: The function uses a non-headless browser (visible) which might not be suitable for production environments. Change headless=False to headless=True for invisible operation. """ try: with sync_playwright() as p: browser = p.chromium.launch(headless=False) page = browser.new_page() page.goto(url, timeout=60000) scroll_position = 0 # Start from the top while True: # Scroll down in increments scroll_position += 800 page.evaluate(f"window.scrollTo(0, {scroll_position})") time.sleep(2) # Wait for new content to load # Get new page height after scrolling new_height = page.evaluate("document.body.scrollHeight") # Stop if we can't scroll further if scroll_position >= new_height: break content = page.content() browser.close() logging.info(f"Successfully fetched content from {url}") return content except Exception as e: logging.error(f"Error fetching page content from {url}: {e}") return None ``` This function is designed to handle one of the trickiest parts of scraping modern websites—[lazy loading](https://en.wikipedia.org/wiki/Lazy%5Floading?ref=blog.datahut.co). On sites like Zepto, the full list of products doesn’t appear all at once. Instead, more items load only when you scroll down. So, if we want to grab all the product details, we need a way to scroll automatically, just like a human would. To do that, we use Playwright, which launches a real browser. We run it in visible mode (headless=False) so you can actually watch how the scraping works in action. The browser is set to wait up to 60 seconds for the page to fully load—this helps on slower internet connections. The main trick here is the simulated scrolling. The scraper scrolls down the page bit by bit—about 800 pixels at a time—and pauses for 2 seconds between each scroll. That short pause gives the website time to load more content in the background using JavaScript. In most cases, 2 seconds is enough for this to happen smoothly. As it scrolls, the scraper checks whether it has reached the bottom of the page by comparing how far it's scrolled with the total height of the content. Once all products are loaded and there's nothing more to scroll, the function captures the full HTML content of the page. After collecting the data, the browser closes, and the HTML content is returned so we can parse it later. If anything goes wrong—like if the browser crashes or the network fails—the function catches the error, logs what happened, and returns None instead of breaking the whole script. This helps the scraper run smoothly and makes it easier to troubleshoot when needed. ### Link Extraction ``` def parse_links(html_content): """ Parses the HTML content and extracts all product links. This function uses BeautifulSoup to parse the HTML and extract links to product pages, specifically targeting elements that match the product card selector. Args: html_content (str): The HTML content to parse. Returns: list: A list of product URLs with the base URL prepended. Raises: Exceptions during parsing are caught and logged. Note: The function specifically targets div elements with data-testid="product-card" which should be updated if the website structure changes. """ try: soup = BeautifulSoup(html_content, 'html.parser') links = [BASE_URL + a['href'] for a in soup.select('div.w-full > div.grow > div > div > a[data-testid="product-card"]', href=True)] logging.info(f"Extracted {len(links)} links.") return links except Exception as e: logging.error(f"Error parsing links: {e}") return [] ``` Now that we’ve got the full HTML content of the page, the next step is to pull out the useful parts—in this case, the links to individual product pages. This is where BeautifulSoup comes in. BeautifulSoup takes the raw HTML and turns it into a format that’s much easier to navigate—almost like a family tree of elements. This lets us zoom in on exactly what we need without digging through all the messy code manually. To find the product links, we use a CSS selector, which is basically a way to tell BeautifulSoup, “Look here!” The selector we use is: 'div.w-full > div.grow > div > div > a\[data-testid="product-card"\]'. This points directly to the a tags (which are HTML links) that have the attribute data-testid="product-card". These are the clickable product cards on Zepto’s site. It’s a reliable way to identify them, since this structure stays consistent across the page. However, the links we collect are relative URLs—they’re not full links yet. So we use a little Python trick called a list comprehension to combine each relative URL with Zepto’s base URL. This gives us a full set of proper product links we can actually use later. Before the function finishes, it logs how many links it found. That gives us a quick confirmation that the scraping worked. And if something goes wrong during the parsing, it won’t crash the script—it’ll just log the error and return an empty list, so the rest of the code can still keep running smoothly. ### Database Storage ``` def save_links_to_db(links, category, db_name="scraped_links.db"): """ Saves the extracted links with category information to an SQLite database. This function creates a database (if it doesn't exist) and adds the extracted links along with their category, the current date, and a default 'scraped' status of 0 (unscraped). It handles duplicate links by ignoring them rather than raising an error. Args: links (list): A list of URLs to save. category (str): The category label for these links (e.g., "fruits", "vegetables"). db_name (str, optional): The name of the SQLite database file. Defaults to "scraped_links.db". Returns: None Raises: Various database exceptions which are caught and logged. Database Schema: - id: Primary key, auto-incrementing integer - url: The product URL (unique) - category: The product category - scraped_date: The date when the link was added to the database - scraped: Integer flag (0=unscraped, 1=scraped) """ try: conn = sqlite3.connect(db_name) cursor = conn.cursor() # Create table with a 'scraped' column (default 0) if not exists cursor.execute(""" CREATE TABLE IF NOT EXISTS links ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, category TEXT, scraped_date TEXT, scraped INTEGER DEFAULT 0 ) """) current_date = datetime.now().strftime("%Y-%m-%d") # Get today's date in YYYY-MM-DD format for link in links: try: cursor.execute("INSERT INTO links (url, category, scraped_date, scraped) VALUES (?, ?, ?, ?)", (link, category, current_date, 0)) except sqlite3.IntegrityError: logging.warning(f"Duplicate link ignored: {link}") conn.commit() conn.close() logging.info("Links saved successfully to the database.") except Exception as e: logging.error(f"Error saving links to database: {e}") ``` Once we’ve collected the product links, we need a safe place to store them—somewhere they won’t get lost and can be reused later. That’s where this function comes in. It creates an SQLite database to save all the scraped links in an organized way. The first time this function runs, it creates a new .db file (a simple database file on your computer), along with a table to hold our product link data. Here’s what each column in the table does: - id – A unique number for each entry. It helps us keep track of the rows. - url – The actual product link. This is marked as UNIQUE, so we don’t accidentally save the same link more than once. - category – The product type, like "fruits" or "vegetables". - date – The date when the link was scraped, in a standard YYYY-MM-DD format. - scraped – A flag that’s either 0 or 1\. It tells us if this link has already been scraped for product details or not. That last column, scraped, is especially useful when we’re splitting our scraping into two steps: first, collecting links, and later, visiting each link to get more information. When links are first added, scraped is set to 0\. After we process them in the second step, we update it to 1 so we don’t scrape the same page twice. Sometimes, we might try to add a link that’s already in the database. When this happens, an IntegrityError is raised because of the UNIQUE setting on the URL column. That’s expected, so we handle it gently using a try-except block. Instead of stopping the whole program, we simply log a message and move on. We use [logging.info](http://logging.info/?ref=blog.datahut.co) here, since it’s just a normal part of the scraping process—not something to worry about. Finally, once all the links are added, the function saves the changes and closes the connection to the database properly. This ensures our data is stored safely and ready for the next step. ### Main Execution Logic ``` def main(): """ Main function that orchestrates the scraping process. This function defines the target URLs to scrape, along with their corresponding categories. For each URL, it: 1. Fetches the page content 2. Parses the content to extract product links 3. Saves the extracted links to the database Returns: None Note: The URLs are hardcoded in this function. For more flexibility, consider loading them from a configuration file or command-line arguments. """ urls = { "https://www.zeptonow.com/cn/fruits-vegetables/fresh-vegetables/cid/64374cfe-d06f-4a01-898e-c07c46462c36/scid/b4827798-fcb6-4520-ba5b-0f2bd9bd7208": "vegetables", "https://www.zeptonow.com/cn/fruits-vegetables/fresh-fruits/cid/64374cfe-d06f-4a01-898e-c07c46462c36/scid/09e63c15-e5f7-4712-9ff8-513250b79942": "fruits" } for url, category in urls.items(): html_content = fetch_page_content(url) if html_content: links = parse_links(html_content) if links: save_links_to_db(links, category) print("Links saved successfully.") if __name__ == "__main__": main() ``` The main() function is like the brain of our scraper—it controls how everything runs from start to finish. It starts by creating a simple dictionary that maps each product category (like fruits or vegetables) to its matching URL. This setup makes it easy to add new categories or rename existing ones later, without having to change the rest of the code. Once the setup is ready, the function goes through each category one by one and follows these steps: 1. It first uses the scrolling function to load the full page content. 2. If the content loads correctly, it moves on to extract product links from the page using our parsing function. 3. If product links are found, they’re saved into the database, along with the category they belong to. After each step, we use if conditions to check whether things worked as expected before moving forward. This helps avoid a chain reaction of errors—if one part fails, the scraper simply skips the next step instead of crashing. There’s also a small but important detail: the main() function only runs when this script is executed directly, not when it's imported into another file. This is a best practice in Python that helps keep your code clean, modular, and reusable. Finally, when everything’s done, the function prints a simple message to confirm that the scraping process has been successfully completed. It’s a neat way to wrap things up and know your data is safely collected. ## Data Collection After collecting all the product links from Zepto’s fruits and vegetables section, the next step is to visit each of those links and pull out detailed product information. This part of the scraper is designed to do just that—it opens each product page, grabs the important details, and saves them neatly into our database. The scraper works by going through one product at a time, taking short pauses between each request. These small delays help us scrape responsibly, without putting too much pressure on the website’s servers. The product scraper is built using a [modular design](https://www.blog.datahut.co/post/web-scraping-a-dynamic-ecommerce-website-using-python/), which means each task is handled by its own function. Some functions connect to the database, others fetch the web page, some parse the product information, and a few manage the overall workflow. This makes the code easier to read, test, and update later if needed. In the sections that follow, we’ll break down each part of this scraper and walk through how it works—from grabbing the page to saving the final results. ### Setup and Configuration ``` import sqlite3 import logging from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup import time import json import random from datetime import datetime # Configure logging logging.basicConfig(filename="product_scraper.log", level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") DB_NAME = "scraped_links.db" ``` Before we can start building the product scraper, Here also we need to bring in the same set of Python libraries we used in the first part. We also set up logging, which writes messages and errors into a file called "product\_scraper.log". Since this part of the scraper runs separately from the link collector, logging gives us a way to monitor what’s going on and spot issues if they come up. Lastly, we define a constant called DB\_NAME. This tells our scraper which database file to use, so all parts of the code stay in sync and work with the same data throughout the process. ### Database Operations for Link Retrieval ``` def get_unscraped_links(): """ Fetch all links from the database where scraped = 0. Queries the 'links' table to retrieve records that haven't been processed yet. Returns: list: A list of tuples containing (id, url, category) for each unscraped link. Raises: Various database exceptions which are caught and logged. Note: This function relies on the database structure created by the link scraper. The 'links' table should have columns for id, url, category, and scraped status. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("SELECT id, url, category FROM links WHERE scraped = 0") links = cursor.fetchall() # List of tuples: (id, url, category) conn.close() logging.info(f"Fetched {len(links)} unscraped links from the database.") return links except Exception as e: logging.error(f"Error fetching unscraped links: {e}") return [] ``` This part of the code is all about getting the list of product links that still need to be scraped. To do that, we connect to our SQLite database and fetch every product entry where the scraped value is set to 0—which means the data hasn’t been collected from those pages yet. For each of these entries, we pull out three things: - The ID, which helps us later mark the product as "done" once it’s scraped - The URL, which points to the product’s page - The category, which tells us whether it’s a fruit, vegetable, or something else This setup acts like a simple to-do list inside our database. Rather than storing links in a file or a long list in memory, the database keeps track of what’s left to do. So if the scraper stops midway—say due to a network issue—we can run it again later and it’ll simply continue from where it left off. It’s a practical way to make sure no work is repeated, and every product gets scraped exactly once. ### Content Fetching ``` def fetch_page_content(url): """ Fetches the full product page content using Playwright. Launches a headless browser to render JavaScript and fetch the complete HTML content of a product page, ensuring all dynamic elements are loaded. Args: url (str): The URL of the product page to fetch. Returns: str or None: The HTML content of the page if successful, None otherwise. Raises: Various exceptions from Playwright which are caught and logged. Note: Includes a 5-second delay to allow JavaScript content to fully load. """ try: with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto(url, timeout=60000) time.sleep(5) # Give some time for JavaScript content to load content = page.content() browser.close() logging.info(f"Successfully fetched page content for {url}") return content except Exception as e: logging.error(f"Error fetching page content from {url}: {e}") return None ``` This function is responsible for loading individual product pages in the background so we can scrape the details we need. It uses a headless browser, which means it runs without opening a visible browser window. This helps things run faster and more efficiently, especially since we don’t need to scroll or interact with the page—we just need to wait for it to fully load. Once the page opens, the scraper waits for 5 seconds to give the website enough time to load everything, including any dynamic content like the product’s price or description. We also set a 60-second timeout, just in case the page is slow or your internet connection takes a little longer. If something goes wrong—like the page fails to load or times out—the function doesn’t crash the whole program. Instead, it simply logs the error and moves on to the next product. This way, a few failed pages won’t stop the entire scraping process. It’s a smart way to keep things running smoothly, even when there are occasional hiccups. ### Parsing Functions ``` def parse_product_name(soup): """ Extracts the product name from the page. Args: soup (BeautifulSoup): The parsed HTML of the product page. Returns: str: The product name if found, 'N/A' otherwise. Note: Uses a specific CSS selector that may need updating if the website structure changes. """ try: return soup.select_one("#product-features-wrapper > div:nth-child(1) > div > div.mt-2.flex.items-center.justify-between.gap-6 > h1").text.strip() except AttributeError: logging.warning("Product name not found.") return "N/A" ``` In this part of the scraper, we’ve created a group of small helper functions, where each one focuses on pulling out just one piece of product information—like the name, price, or stock status. This approach keeps the code organized and easy to manage. If something breaks—say the website layout changes slightly—we only need to fix the specific function related to that data point, without touching the rest of the scraper. Each function follows the same basic steps: 1. It looks for a specific element on the page using a CSS selector. 2. It grabs the text inside that element. 3. It removes any extra spaces around the text. 4. It returns the final value. If the function can’t find what it’s looking for—maybe the element is missing or the layout changed—it won’t crash the program. Instead, it logs a warning and returns "N/A". This way, the scraper keeps going and collects everything else that’s still available. The CSS selectors used in these functions were carefully picked by inspecting Zepto’s product page layout using browser tools like “Inspect Element.” For example, one function grabs the product name, while others pull out the net quantity, discounted price, original price, and availability status—all in the same reliable way. By handling each piece of data separately, we make the scraper much easier to read, test, and update when needed. ### Complex Data Parsing ``` def parse_product_highlights(soup): """ Extracts the product highlights section with key-value pairs. This function identifies the product highlights section and extracts all key-value pairs found within it, returning them as a JSON string. Args: soup (BeautifulSoup): The parsed HTML of the product page. Returns: str: A JSON string containing key-value pairs of product highlights. Returns '{}' (empty JSON object) if no highlights are found or in case of errors. Raises: Exceptions during parsing are caught, logged, and an empty JSON object is returned. Note: The function expects a specific HTML structure with h3 elements for keys and p elements for values within each highlight div. """ try: highlights_section = soup.select("#productHighlights > div > div > div.flex.flex-col.gap-8 > div.flex.items-start.gap-3") if not highlights_section: logging.warning("No product highlights found in the given HTML.") return json.dumps({}) highlights = {} for div in highlights_section: try: key_element = div.find("h3") value_element = div.find("p") if key_element and value_element: key = key_element.get_text(strip=True) value = value_element.get_text(strip=True) highlights[key] = value else: logging.warning("Missing key-value pair in a product highlight div.") except Exception as e: logging.error(f"Error processing a highlight div: {e}") continue return json.dumps(highlights, indent=4) except Exception as e: logging.error(f"Error parsing product highlights: {e}") return json.dumps({}) ``` For more detailed information—like product highlights—our scraper needs to go beyond just grabbing plain text. These details often come in pairs, like a heading and a description, so we use a special method to collect them as key-value pairs, and then turn them into a JSON string. This makes the data easier to store, read, and use later on. Here’s how it works: The function starts by finding the highlights section on the page using a CSS selector. Once found, it loops through each item in that section—grabbing the heading (which becomes the key) and the description (which becomes the value). Each pair is added to a dictionary. After collecting everything, the dictionary is neatly converted into a formatted JSON string. This structure is really helpful when the product has multiple features or specifications. Rather than trying to force everything into one long string, we get clean, organized data that’s easy to work with. The function also uses nested try/except blocks to handle any errors. The outer block checks whether the entire highlights section exists. Inside that, each individual item is handled carefully—so if one part is missing or doesn’t load properly, it won’t break the entire process. If the first part can’t be read, it simply skips the rest and keeps going without crashing. We use the same method for another section on the page too—the product info section. This keeps everything consistent and helps us capture more structured data where needed. ### Combined Product Details Extraction ``` def parse_product_details(html_content): """ Parses all product details from the HTML content. This function serves as the main parser that coordinates the extraction of all product details by calling individual parsing functions for each data point. Args: html_content (str): The HTML content of the product page. Returns: dict or None: A dictionary containing all extracted product details if successful, None otherwise. Raises: Exceptions during parsing are caught, logged, and None is returned. Note: The returned dictionary contains keys for name, net_quantity, sale_price, product_price, in_stock, highlights, and information. """ try: soup = BeautifulSoup(html_content, 'html.parser') product_data = { "name": parse_product_name(soup), "net_quantity": parse_net_quantity(soup), "sale_price": parse_sale_price(soup), "product_price": parse_product_price(soup), "in_stock":parse_stock(soup), "highlights":parse_product_highlights(soup), "information":parse_product_info(soup) } logging.info(f"Extracted product details: {product_data}") return product_data except Exception as e: logging.error(f"Error parsing product details: {e}") return None ``` The main parsing function is the part that brings everything together. Its job is to take the raw HTML from a product page and coordinate all the smaller functions to collect complete product information. It starts by turning the raw HTML into a BeautifulSoup object, which makes it much easier to search and extract specific elements from the page. Once the page is ready, the function calls each of our smaller helper functions—like the ones that get the product name, prices, highlights, and availability. Each of these functions pulls out one piece of information, and their results are all combined into a single, well-structured dictionary. This dictionary holds everything we’ve scraped for that product in one place. To make sure things run smoothly, the function includes error handling. If something unexpected happens while parsing the page, the error is caught and logged, so the scraper doesn’t break. Before returning the final result, the function also logs all the details it collected. This is really helpful when you want to double-check your results or troubleshoot if something doesn’t look right. It gives you a clear view of what was successfully scraped for each product. ### Database Operations for Product Storage ``` def save_product_data(product_data, category, scraped_date, url): """ Saves product details into the database. Creates a 'products' table if it doesn't exist and inserts the extracted product data along with metadata such as category and scrape date. Args: product_data (dict): Dictionary containing product details. category (str): The product category. scraped_date (str): The date when the product was scraped (YYYY-MM-DD format). url (str): The URL of the product page. Returns: bool: True if data was successfully saved, False otherwise. Raises: Various database exceptions which are caught and logged. Note: The return value is used to determine whether to mark the link as scraped in the 'links' table. If saving fails, the link remains unscraped so it can be retried later. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() # Create table if it doesn't exist cursor.execute(""" CREATE TABLE IF NOT EXISTS products ( url TEXT, name TEXT, net_quantity TEXT, sale_price TEXT, product_price TEXT, in_stock TEXT, category TEXT, highlights TEXT, information TEXT, scraped_date TEXT ) """) cursor.execute(""" INSERT INTO products (url, name, net_quantity, sale_price, product_price, in_stock, category, highlights, information, scraped_date) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?) """, (url, product_data["name"], product_data["net_quantity"], product_data["sale_price"], product_data["product_price"], product_data["in_stock"], category, product_data["highlights"], product_data["information"], scraped_date)) conn.commit() conn.close() logging.info(f"Product data saved successfully for {url}") return True # Indicate success except sqlite3.IntegrityError: logging.warning(f"Duplicate product ignored: {url}") return False # Do not mark as scraped except Exception as e: logging.error(f"Error saving product data to database: {e}") return False # Do not mark as scraped ``` This function handles the job of saving the scraped product data into our SQLite database. It starts by checking if the products table already exists. If it doesn’t, the function creates the table using the same structure as the fields we’ve collected—like product name, price, stock status, and so on. For more detailed fields, like highlights and product info, which hold multiple pieces of data, we store them as JSON strings. This keeps their structure intact inside the database, making it easier to work with later. Once the table is ready, the function tries to save the product data. If everything is stored correctly, it returns True. If something goes wrong, it returns False. This return value is important. It lets the scraper know whether it’s safe to mark the product link as “processed.” We only mark links as done when we’re sure their data has been saved properly. That way, we avoid losing any products due to errors or interruptions, and we keep the scraper accurate and reliable. ### Status Update Function ``` def mark_link_as_scraped(link_id): """ Updates the `scraped` column in the `links` table to mark the link as processed. After a product page has been successfully processed and its data saved, this function marks the corresponding link as scraped (scraped=1) to prevent it from being processed again. Args: link_id (int): The ID of the link in the 'links' table. Returns: None Raises: Various database exceptions which are caught and logged. Note: This function relies on the database structure created by the link scraper. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("UPDATE links SET scraped = 1 WHERE id = ?", (link_id,)) conn.commit() conn.close() logging.info(f"Marked link ID {link_id} as scraped.") except Exception as e: logging.error(f"Error updating scraped status for link ID {link_id}: {e}") ``` This function is used to mark a product link as processed once its data has been successfully scraped and saved. It does this by updating the scraped field in the links table, changing its value from 0 to 1. By keeping track of which links are already processed, the scraper can skip them in future runs. This way, we avoid repeating work and keep everything running efficiently—even if the scraper is stopped and restarted. To make the update safe and precise, the function uses a parameterized SQL query, which updates only the row that matches the given link’s unique ID. This ensures that only the correct entry is changed, without affecting any others. It’s a clean and reliable way to manage our progress and make the scraper more sustainable over time. ### Main Scraping Workflow ``` def scrape_products(): """ Main function to scrape product details from unscraped links. This function orchestrates the entire scraping process: 1. Fetches unscraped links from the database 2. For each link, fetches the product page content 3. Parses product details from the page content 4. Saves the data to the database 5. Marks the link as scraped if the save was successful 6. Implements random delays between requests to avoid detection Returns: None Note: This function implements error handling and logging at each step. It uses random delays between 7-10 seconds to avoid overloading the server and to reduce the risk of being detected as a bot. """ unscraped_links = get_unscraped_links() if not unscraped_links: logging.info("No unscraped links found. Exiting scraper.") return for link_id, url, category in unscraped_links: logging.info(f"Processing {url} (Category: {category})") html_content = fetch_page_content(url) if html_content: product_data = parse_product_details(html_content) if product_data: scraped_date = datetime.now().strftime("%Y-%m-%d") # Only mark as scraped if saving was successful if save_product_data(product_data, category, scraped_date, url): mark_link_as_scraped(link_id) logging.info(f"Successfully saved and marked {url} as scraped.") else: logging.warning(f"Skipping marking {url} as scraped due to save failure.") time.sleep(random.uniform(7,10)) # Random delay between processing requests logging.info("Scraping completed for all available links.") if __name__ == "__main__": scrape_products() ``` This function is the main controller for the entire product scraping process. It runs through each step in a clear, orderly way and keeps everything running smoothly from start to finish. First, it grabs a list of product links from the database where scraping hasn’t been attempted yet (those marked with scraped = 0). Then, it processes each link one at a time, following this simple and reliable workflow: 1. Load the product page using our page loader. 2. Parse the product information using the small helper functions we created earlier. 3. Save the collected data into the database. 4. Update the product link to show that it’s been successfully scraped. This process is designed to be resilient and careful. If there are no links left to scrape, the function stops early—saving time and resources. And if something goes wrong while processing a link (like the page doesn’t load or parsing fails), the link isn’t marked as done. That way, we can try it again later without losing any data due to temporary issues. To stay polite and [avoid being flagged as a bot](https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/), the function also adds a small random delay between requests. This helps reduce the load on the website’s server and makes the scraper less likely to be blocked or rate-limited. Finally, this function is only run when it’s specifically called, so it can be used as a standalone tool or plugged into a larger scraping system. It’s flexible, efficient, and built to handle real-world challenges without skipping a beat. ### Conclusion Web scraping helps turn messy website content into clean, usable data. In this project, we scraped product details from Zepto’s Fruits and Vegetables section using Playwright, BeautifulSoup, and SQLite. We learned how to handle JavaScript-loaded content with Playwright, save data neatly using SQLite, and scrape responsibly by adding random delays and error handling. By splitting the process into two steps—collecting links first, then scraping product details—we made the scraper more reliable and easier to manage. This approach works well not just for product data, but also for price tracking, market research, and more. Just remember: as websites evolve, scrapers should too—and always be respectful and ethical in how you use them. Author I’m Shahana, a Data Engineer at Datahut, where I specialize in building reliable, scalable data pipelines that convert complex web content into clean, structured datasets—particularly for industries like e-commerce, grocery delivery, and retail analytics. At Datahut, we help clients automate data extraction from modern websites, including those that use JavaScript and infinite scrolling. In this blog, I walked through a real-world scraping project focused on Zepto’s Fruits and Vegetables section. We used Playwright, BeautifulSoup, and SQLite to build a solution that handles dynamic content, stores data efficiently, and follows responsible scraping practices—all while keeping the code clean and beginner-friendly. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. FAQ SECTION FAQ 1: Is it legal to scrape data from Zepto? - Zepto's publicly visible product data (names, prices, availability) is generally accessible, but always review the website's Terms of Service before scraping - Avoid scraping personal user data, login-protected pages, or any content explicitly restricted in Zepto's robots.txt file - Use scraping responsibly — add delays between requests, avoid overloading the server, and use the data only for lawful purposes like research or price monitoring - For a deeper understanding, refer to Datahut's guide on [whether scraping e-commerce websites is legal](https://www.blog.datahut.co/post/is-web-scraping-legal/) FAQ 2: Why do we need Playwright instead of a simple requests library for Zepto? - Zepto is a JavaScript-heavy website — its product listings don't exist in the raw HTML source; they load dynamically in the browser - A basic requests call only fetches the static HTML, which means you'd get an empty or incomplete page with no products - Playwright launches a real Chromium browser, waits for JavaScript to execute, and captures the fully rendered page — just like a human visiting the site - It also handles infinite scrolling automatically, which is essential since Zepto loads more products as you scroll down FAQ 3: What happens if the scraper stops midway through collecting data? - Because we use SQLite with a scraped flag (0 = pending, 1 = done), the scraper knows exactly where it left off - On restart, it simply queries for all links where scraped = 0 and continues from there — no data is lost or duplicated - Product data is only marked as scraped after it has been successfully saved, so a failed save automatically retries on the next run - This two-step design (links first, then product details) is specifically built for resilience against interruptions FAQ 4: Can this scraper be adapted for other categories or grocery apps? - Yes — for other Zepto categories, simply add the new category URL and a label to the urls dictionary in the main() function - For other grocery apps like Blinkit or BigBasket, the core structure (Playwright + BeautifulSoup + SQLite) stays the same; only the CSS selectors and URLs need to be updated to match the new site's layout - The modular design — where each function handles one task — makes it straightforward to swap out or update individual parts without rewriting the whole script - Datahut has published a similar project for [Blinkit's fruits and vegetables section](https://www.blog.datahut.co/post/how-to-scrape-blinkit-s-fruits-and-vegetables-data/) that follows the same approach FAQ 5: How do we handle changes in Zepto's website structure? - Zepto may update its HTML layout, CSS class names, or data attributes over time — when this happens, CSS selectors used in the parsing functions will stop matching and return "N/A" - The fix is to open the updated product page in Chrome, right-click the element you need, select "Inspect," and copy the new CSS selector path - Each piece of data (name, price, stock, highlights) is handled by its own dedicated function, so updating one selector never breaks the others - Setting up logging (which this scraper already does) helps you spot when fields start returning "N/A" consistently — that's usually the first sign a selector needs updating - Treating CSS selectors as the one maintenance cost of any scraper is a realistic expectation — the rest of the pipeline stays stable ### How to Scrape Product Data from Family Food Centre? (Step-by-Step Python Guide) URL: https://www.blog.datahut.co/post/how-to-scrape-product-data-from-family-food-centre/ Last updated: 2026-09-07T09:43:00.000Z [Have you ever wondered if there's a way to automatically collect information from online shopping websites](https://www.blog.datahut.co/post/web-scraping-vs-web-crawling-2026/) \- without manually copying and pasting everything? That’s exactly what web scraping helps us do. It’s like teaching a computer to visit web pages and pick out the information you need, just like how you would—but much faster and without the effort. In this blog, I’ll walk you through a simple web scraping project I worked on. The goal? To collect data from the fruits and vegetables section of the Family Food Centre website. If you’re not familiar, Family Food Centre is a well-known online supermarket in Qatar, offering a wide variety of fresh produce. The scraping process was divided into two clear steps First, we needed to visit the fruits and vegetables category page and collect all the product links listed there. Then, using those links, we visited each individual product page to gather detailed information—like the product name, price, and packaging details. We've used the same two-phase approach on other grocery platforms too - c[heck out how we did it for Blinkit.](https://www.blog.datahut.co/post/how-to-scrape-blinkit-s-fruits-and-vegetables-data/) ## Scrape Product Data from Family Food Centre: Tools Behind the Magic You might be wondering—how is web scraping even possible? How can a computer visit a website and pick out just the parts we care about? Well, that’s where some smart tools come in. Throughout this project, I used a few simple but powerful tools that made the entire scraping process much easier. Think of them as a team working together—each with its own special job. Let’s meet them. ### [Playwright ](https://playwright.dev/python/?ref=blog.datahut.co)– The One Who Browses for You Imagine you’re sitting in front of a website, clicking buttons, scrolling down, and waiting for new items to load. Now imagine handing that job over to a helper who can do all that for you—quickly and accurately. That’s what Playwright does. If you're hitting anti-bot walls, see [how curl\_cffi can help you scrape without getting blocked.](https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/) Playwright is a tool that acts like a real browser. It can open web pages, click on things, and scroll through content, just like a human would. This is especially useful for websites that load more products only when you scroll down (you’ve probably seen this on shopping sites). Playwright takes care of that smoothly, without you lifting a finger. ### [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co) – The One Who Finds What You Need Once the page is fully loaded, it’s time to pick out the useful details—like product names, prices, or packaging info. That’s where BeautifulSoup comes in. Think of a webpage as a messy room full of information. BeautifulSoup helps us search through that room and find exactly what we’re looking for. It takes the page’s HTML code (which is how websites are built behind the scenes) and lets us extract just the parts we care about. ### [SQLite](https://www.sqlite.org/index.html?ref=blog.datahut.co) – The One Who Keeps Everything Safe After collecting all this information, we need a place to store it. Not in a notebook—but in a digital one called SQLite. SQLite is a simple, lightweight database that runs on your own computer. You don’t need to install anything fancy or set up a server. It helps you save your data in a neat and organized way so you can come back to it later, run analysis, or even share it. ## Getting Started: The Foundation Code Now that everything is set up, it’s time to collect the actual product links. In this part, we go through the fruits and vegetables section of the website and gather links to each product listed there. These links will later help us visit each product page and pull out the detailed information we need. We’ll also handle multiple pages, just like a user clicking through “Next” to see more items. The steps below will walk you through how the script grabs those links and gets them ready for the next stage. ### Setting Up Our Tools: The Import Section Before jumping into the code, there’s one important step we need to take—setting up our tools. Each of these libraries plays a specific role in the process, and together, they help everything run smoothly. ``` import asyncio import sqlite3 import datetime from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` Let’s start by understanding the tools we’re bringing into our project—and why they’re important. First, we have asyncio. Think of it as the manager that keeps things moving behind the scenes. When our script is waiting for a page to load or some data to process, asyncio makes sure that time isn’t wasted. Instead of just sitting and waiting, it allows the program to keep working on other tasks. This helps our scraper run faster and more efficiently, especially when dealing with lots of web pages. Another necessary tool as we told before is sqlite3, our structured data storage system. Manually storing hundreds or thousands of product links is not feasible—this is where SQLite plays its role by efficiently storing scraped data. To go along with datetime, which timestamps our gathered data for tracking purposes, these pieces are the basis of our data storage solution. Finally, our stars of our scraping process—Playwright and BeautifulSoup—collaborate to scan web pages and extract useful content, making the entire scraping task smooth and easy. ### Global Settings: Our Project's Command Center ``` # Global configuration DB_NAME = "product_links.db" URL = "https://family.qa/default/produce/fruits-and-vegetables.html" ``` Every project needs a clear starting point—and some basic rules to follow. In our web scraping script, that structure comes from global variables. Think of them as fixed reference points that guide how and where the script works. In this case, we define two key variables: - DB\_NAME tells the script where to store the scraped data (our database file), and - URL tells it where to begin scraping (the starting web page). Even though these values are simple and don’t change during the run of the script, they play a big role. They’re like signposts that point the program in the right direction. Setting these values at the top of our script keeps things neat and easy to manage. For example, if we ever want to scrape a different section of the website, we just update the URL. Or if we want to save the data to a different file, we simply change the DB\_NAME. There’s no need to dig through the whole code to make those changes. By organizing our script this way, we make it easier to maintain, flexible to adapt, and ready to grow if we want to scale things up later. ### Database Creation: Building Our Digital Warehouse ``` def create_database(): """ Creates and initializes an SQLite database for storing product links. Creates a 'products' table with the following schema: - id: INTEGER PRIMARY KEY AUTOINCREMENT - link: TEXT (stores the product URL) - date: TEXT (stores the scraping date) - scraped: INTEGER (flag to track processed links, default 0) Raises: sqlite3.Error: If there's an error creating the database or table """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS products ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT, date TEXT, scraped INTEGER DEFAULT 0 ) """) conn.commit() except sqlite3.Error as e: print(f"Database error: {e}") finally: conn.close() ``` Before we start collecting data, we need a proper place to store it—something organized and reliable. That’s where our create\_database() function comes in. Its job is to set up a small, local database using SQLite, where all our product links will be stored safely. The process begins by connecting to the database using sqlite3.connect(DB\_NAME). If the database file doesn't exist yet, no worries—SQLite will automatically create one for us. This makes it very beginner-friendly and easy to work with. Once connected, we create something called a cursor. You can think of the cursor as a messenger—it helps us send instructions (SQL commands) to the database. Next, we ask the database to create a table called products—but only if it doesn’t already exist. This is done using the CREATE TABLE IF NOT EXISTS command. The table will have four columns: - id: a unique number for each entry (automatically increases for every new product), - link: where the product URL is stored, - date: which saves the date when the data was collected, - scraped: a flag to tell us whether the product has already been scraped. It starts with a value of 0, meaning "not yet scraped." After setting everything up, we save the changes and close the connection to keep things clean and free up resources. If anything goes wrong during this process, the function will catch the error and print it out. That way, we can quickly figure out what happened and fix it. ### The Link Saving System: Our Data Archival Process ``` def save_links_to_db(links): """ Saves product links to the SQLite database with the current date. Args: links (list): List of product URLs to be saved Notes: - Each link is saved with the current date and scraped=0 flag - Duplicate links are allowed to track historical data - Current date is stored in ISO format (YYYY-MM-DD) Raises: sqlite3.Error: If there's an error during database operations """ if not links: print("No links to save.") return conn = None try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() current_date = datetime.date.today().isoformat() inserted_count = 0 for link in links: cursor.execute("INSERT INTO products (link, date, scraped) VALUES (?, ?, 0)", (link, current_date)) inserted_count += 1 # Count inserted links conn.commit() print(f"Inserted {inserted_count} links into the database.") except sqlite3.Error as e: print(f"Database insert error: {e}") finally: if conn: conn.close() ``` When we collect the product links from the website, the next step is to store them neatly in our database. That’s exactly what the save\_links\_to\_db(links) function does—it takes a list of product URLs and saves each one into our SQLite database, tagging them with today’s date for reference. The function starts by checking whether the list of links is empty. If there are no links to save, it simply prints a message and stops right there—no need to move forward if there’s nothing to do. But if we do have links, the function connects to the database and sets up a cursor to send instructions. Then it uses datetime.date.today().isoformat() to get today’s date in a clean, standard format (like "2025-06-12"), which helps us keep track of when each link was added. Each link is then added to the products table with that date. We also include a scraped value, which is set to 0 for now—this tells us that we haven’t yet scraped the full product details from this link. Think of it as a little note saying, “Hey, this one’s still waiting to be processed.” The function keeps count of how many links it successfully saves and prints that out once it’s done, giving us helpful feedback. If something goes wrong—like a connection issue or an unexpected error—it catches the problem and prints it clearly so we can fix it. And just like good housekeeping, it makes sure to close the database connection at the end, no matter what. ### Link Extraction: Finding Treasures in HTML ``` def extract_product_links(html): """ Extracts product links from the HTML content using BeautifulSoup. Args: html (str): Raw HTML content of the page Returns: list: List of extracted product URLs Notes: - Targets links with class 'product-item-link' - Filters out None/empty links - Uses BeautifulSoup's CSS selector for efficient parsing """ soup = BeautifulSoup(html, "html.parser") product_links = [] for a_tag in soup.select("a.product-item-link"): link = a_tag.get("href") if link: product_links.append(link) return product_links ``` To work with data from a website, we can’t just rely on the raw HTML—it’s messy, unorganized, and full of information we don’t need. That’s where structured parsing comes in. It helps us focus on just the pieces we care about. In this part of the process, we use a function called extract\_product\_links(html) to pull out only the product links from the HTML code we’ve collected. Here’s how it works: First, we take the big block of HTML content and hand it over to BeautifulSoup, a library that helps us understand and navigate the structure of a web page. It’s a bit like turning a scrambled document into a searchable map where we can zoom in on the parts we want. Next, the function searches through this map to find all the tags (these are the building blocks of links on a webpage). But we don’t want just any links—we’re looking specifically for ones that have a class named "product-item-link", which, in this website’s layout, are the ones that point to individual product pages. Once it finds those tags, it uses .get("href") to grab the actual URLs hidden inside them. These links are added to a list, but only if they’re not empty (sometimes a tag might be there but missing the actual link, so we skip those). In the end, the function returns a clean list of product links—nothing more, nothing less. This step is crucial because it filters out all the extra clutter and leaves us with just the data we need to move forward: the direct links to the product pages we want to explore next. ### Page Navigation: Our Automated Browser Control ``` async def fetch_all_pages(): """ Asynchronously fetches all paginated pages and extracts product links. Returns: list: Consolidated list of product URLs from all pages Notes: - Uses Playwright for browser automation - Implements pagination handling - Waits for network idle state to ensure page loads - Uses a 60-second timeout for initial page load - Handles browser cleanup in case of errors Raises: Exception: Any error during page navigation or scraping """ async with async_playwright() as p: browser = await p.chromium.launch(headless=True) page = await browser.new_page() try: await page.goto(URL, timeout=60000) # 60s timeout to prevent timeouts await page.wait_for_load_state("networkidle") all_links = [] page_number = 1 while True: content = await page.content() links = extract_product_links(content) all_links.extend(links) print(f"Page {page_number}: Scraped {len(links)} links.") # Debugging output next_button = await page.query_selector("#layered-ajax-list-products > div:nth-child(3) > div.pages > ul > li.item.pages-item-next > a") if next_button: await next_button.click() await page.wait_for_load_state("networkidle") page_number += 1 else: break except Exception as e: print(f"Scraping error: {e}") finally: await browser.close() print(f"Total Scraped: {len(all_links)} links.") # Debugging output return all_links ``` To collect product links from multiple pages of a website, we use the fetch\_all\_pages() function. This function works like a tireless assistant that mimics how a human browses—clicking through each page and saving useful links along the way. It uses Playwright, which opens a browser in the background (called "headless" mode, since it doesn't actually show the browser window). Once the browser opens, it visits the starting URL and waits for the entire page to finish loading. This is important because many websites load content dynamically as you scroll or wait. After the page loads, the function grabs the HTML content and sends it to another helper function, extract\_product\_links(), which finds all the product URLs on that page. These links are saved into a growing list. Then, the function checks if there’s a “Next” button on the page. If it finds one, it clicks on it—just like a real user would—and waits for the next page to load fully. This cycle repeats: scrape, click “Next,” wait, and repeat. If there is no “Next” button, the function understands it has reached the last page and stops. Throughout the process, it prints helpful updates, like how many links were found on each page. Finally, once all pages are scraped, the browser is closed to save memory, and the full list of product links is returned. This way, we make sure no product is left behind—even if it's hidden on page 10! ### The Main Orchestra Conductor ``` async def main(): """ Main execution function that orchestrates the scraping process. Flow: 1. Creates/initializes the database 2. Fetches all product links 3. Saves the links to database 4. Reports success/failure Notes: - Handles the case when no links are scraped - Provides feedback about the operation """ create_database() links = await fetch_all_pages() if links: save_links_to_db(links) print(f"Saved {len(links)} product links to the database with the current date.") else: print("No links were scraped.") if __name__ == "__main__": asyncio.run(main()) ``` The main() function is like the project manager of our entire web scraping workflow—it controls the order of execution and ensures everything runs smoothly. It starts by making sure our data storage is ready using create\_database(). This step sets up the SQLite database where all product links will be saved, ensuring we have an organized place to store the data. Next, it kicks off the core scraping process by calling the fetch\_all\_pages() function. This function runs asynchronously, which means it can handle multiple tasks efficiently without waiting for each step to finish completely before moving on. After scraping, the function checks whether any product links were actually found. If links are available, it saves them into the database using save\_links\_to\_db(links) and prints a confirmation message, letting us know how many links were stored. If no links are found, it prints a message to let us know that the process returned no data—useful for debugging or making improvements later. At the bottom of the script, there's a standard Python pattern: if name == "\_\_main\_\_":. This ensures the main() function only runs if the script is executed directly (not when imported into another module). Inside this block, [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(main()) is called to handle all asynchronous operations efficiently. This ensures that our scraping happens smoothly, without blocking or freezing up the rest of the script. ## From Links to Details: Diving Deep into Product Data Now that we have successfully gathered all the product links, we can move on to the crucial next phase—extracting detailed product information. This step involves visiting each collected link and carefully retrieving the relevant data directly from the individual product pages. Details such as the product name, price, and packaging are systematically captured in this stage. In the following sections, we’ll explore how this process is structured in code, ensuring efficient navigation and reliable data extraction across the entire product catalog. ### Setting Up Our Data Collection Tools Just like before, we begin by importing the necessary libraries that facilitate seamless data extraction: ``` import asyncio import sqlite3 from datetime import datetime from playwright.async_api import async_playwright from bs4 import BeautifulSoup # Configuration constant for database DB_NAME = "product_links.db" ``` We’ll continue using the same tools as before, but now we’re shifting our focus. Previously, we used them to find product links. This time, we’ll use them to open each link and collect detailed information from the product pages. One thing that stays the same is how we store the data. We’re still using the same database we set up earlier. This helps keep everything organized in one place. By sticking to a consistent format for storing our data, we make it easier to keep track of what we’ve collected and avoid duplicates or confusion. Think of the product links as doors. Each door leads to a page full of information, and now we’re ready to step through each one, gather what we need, and neatly file it away in our database. This structured approach ensures that everything we collect is easy to find and manage later on. ### Building Our Product Information Warehouse ``` def create_data_table(): """ Creates a table in SQLite database to store detailed product information. Schema: - id: INTEGER PRIMARY KEY AUTOINCREMENT - link: TEXT (product URL) - item_name: TEXT (name of the product) - packing: TEXT (packaging information) - sku: TEXT (product SKU) - price: TEXT (product price) - category: TEXT (product category) - date: TEXT (date of scraping) Raises: sqlite3.Error: If database operations fail Notes: - Uses error handling to manage database connection - Creates table only if it doesn't exist """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT, item_name TEXT, packing TEXT, sku TEXT, price TEXT, category TEXT, date TEXT ) """) conn.commit() except sqlite3.Error as e: print(f"[ERROR] Database table creation failed: {e}") finally: if conn: conn.close() ``` Now that we're ready to store detailed product information, we need a proper place to keep it all. That’s where this function comes in—it sets up a dedicated table inside our SQLite database to hold the data we’re about to collect. First, the function opens a connection to our database. Think of this like opening the door to a storage room. Once inside, it gets ready to create a table by using something called a cursor. This cursor is like a tool that lets us send instructions to the database. The main instruction we give is to create a table named product\_data. We use a special command that checks whether the table already exists—if it does, we don’t create it again. If not, the table gets created with all the columns we need to hold our product details. Each row in this table will store one product’s information. The columns include things like the product link, its name, how it's packaged, its SKU (which is just a unique code used to track items), price, category, and the date we scraped the information. There’s also a column for an ID number that automatically counts up for each new row—this helps us keep everything in order. To make sure everything goes smoothly, the function also watches for any errors. If something goes wrong—like a problem with the database connection—it prints an error message to help us figure out what happened. And no matter what, it always closes the connection at the end, just like locking the door behind you when you leave the storage room. This keeps everything clean and avoids unnecessary memory use. ### Finding Products We Haven't Explored Yet ``` def get_unscraped_links(): """ Retrieves all unprocessed product links from the database. Returns: list: Tuples containing (link, date) for unscraped products Notes: - Queries links where scraped=0 flag - Returns empty list if query fails - Each tuple contains the product URL and its original scrape date """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("SELECT link, date FROM products WHERE scraped = 0") links = cursor.fetchall() return links except sqlite3.Error as e: print(f"[ERROR] Failed to fetch unscraped links: {e}") return [] finally: if conn: conn.close() ``` This function plays an important role in helping us keep track of our progress while scraping. It’s like checking a list to see which tasks are still left to do. Specifically, it looks through our database and pulls out only the product links we haven’t worked on yet. Let’s walk through what it does step-by-step. First, the function connects to the SQLite database—think of this like opening a file where we’ve saved all the product links. Then, it creates a cursor, which acts like a pen that can write or read inside this file. Using that cursor, the function runs a query that says: “Give me all the product links from the products table where the scraped value is 0.” In our setup, a scraped value of 0 means that the product hasn’t been processed yet. This simple check helps us avoid re-scraping the same product pages and keeps everything running efficiently. The results from the query come back as a list of tuples—each tuple includes the link to the product and the date it was added. This makes it easy for us to keep both pieces of information together and refer back to them later if needed. To make the function more reliable, we’ve added some error handling. If something goes wrong—like a glitch in the database or a typo in the query—the function will catch the error, print a helpful message, and safely return an empty list. This way, the rest of the scraping process won’t crash, and we’ll know something needs our attention. Finally, once everything is done, the database connection is closed to tidy up and free any resources. Just like shutting a drawer after you’ve finished looking through your papers, this step helps keep the system clean and ready for the next task. ### Keeping Track of Our Progress ``` def update_scraped_status(link): """ Marks a product link as scraped in the database. Args: link (str): URL of the product that has been processed Notes: - Updates scraped flag to 1 for the given link - Uses parameterized query for SQL injection prevention - Includes error handling for database operations """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("UPDATE products SET scraped = 1 WHERE link = ?", (link,)) conn.commit() except sqlite3.Error as e: print(f"[ERROR] Failed to update scraped status for {link}: {e}") finally: if conn: conn.close() ``` This function helps us keep our scraping workflow neat and organized by marking each product link as “done” once it’s been processed. Think of it like checking off items on a to-do list so we don’t accidentally work on the same thing twice. First, it connects to the SQLite database—this is like opening the notebook where we’ve been storing all our product links. Then, it creates a cursor, which acts like a tool for writing into that notebook. Next comes the update part. The function runs a special command that says: “Find this specific product link and change its status so that it’s marked as scraped.” Technically, it sets the scraped field to 1, which means “this one’s done.” To do this safely, it uses something called a parameterized query—which just means that instead of sticking the link directly into the SQL command (which can be risky), it uses a placeholder and fills it in separately. This method helps prevent errors or even security problems like SQL injection. Once the update is made, the function saves the change permanently by committing the transaction. That’s like hitting “Save” after editing a document, so the changes don’t get lost. To make sure everything runs smoothly, the function includes error handling. If anything goes wrong—maybe the link wasn’t found or the database connection failed—it prints an error message that tells us which link had trouble. That way, we know exactly where to look if we need to fix something. Finally, the function closes the database connection. This is a good habit to keep things tidy and avoid leaving connections open that might slow things down. ### Saving Our Product Treasures ``` def save_product_data(link, item_name, packing, sku, price, category, date): """ Stores extracted product information in the database. Args: link (str): Product URL item_name (str): Name of the product packing (str): Packaging information sku (str): Product SKU price (str): Product price category (str): Product category date (str): Scraping date Notes: - Uses parameterized queries for safe data insertion - Logs successful saves and errors - Ensures database connection is properly closed """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" INSERT INTO product_data (link, item_name, packing, sku, price, category, date) VALUES (?, ?, ?, ?, ?, ?, ?) """, (link, item_name, packing, sku, price, category, date)) conn.commit() print(f"[INFO] Saved product: {item_name}") except sqlite3.Error as e: print(f"[ERROR] Failed to save product data for {link}: {e}") finally: if conn: conn.close() ``` This function plays a key role in saving the detailed product information we’ve collected into our database. Once we’ve scraped data like the product’s name, price, and category, we need a way to store it all in one place so we can refer to it later—and that’s exactly what this function does. Let’s walk through it step by step. First, the function takes in a bunch of details about a product: the link to the product page, its name, how it’s packaged, its SKU (that’s a unique product code), the price, which category it belongs to, and the date we collected the data. All of this information gets passed to the function as inputs. Next, it connects to our SQLite database—think of this as opening up a storage box where we keep all our product records. After connecting, it creates something called a cursor, which is what we use to “write” into the database. Now comes the important part: saving the data. The function uses an INSERT INTO command, which is like saying, “Hey database, here’s a new row of product info—please add it to the table.” To keep things secure and clean, it uses parameterized queries. This simply means the actual values (like the name and price) are added separately, which helps avoid mistakes or security issues like SQL injection. If everything goes smoothly, the function prints a message to let us know the data was added successfully. But if something goes wrong—like if the database is locked or there’s a typo in the data—it logs an error message so we know what happened and can fix it. Finally, just like closing a file when you’re done reading or writing, the function closes the database connection. This is a good habit because it frees up memory and keeps the system running efficiently. ### The Text Extraction Helper ``` def extract_text(soup, selector): """ Safely extracts text content from HTML using CSS selectors. Args: soup (BeautifulSoup): Parsed HTML content selector (str): CSS selector for target element Returns: str: Extracted text or 'N/A' if element not found Notes: - Returns 'N/A' instead of None for missing elements - Strips whitespace from extracted text """ element = soup.select_one(selector) return element.text.strip() if element else "N/A" ``` This small but handy function helps us pull out specific bits of text from a webpage’s HTML. When we’re scraping data from a site, the page is often filled with lots of HTML code, and we just want a specific piece—like a product name or a price. That’s where this function comes in. Imagine you already have the HTML of a webpage loaded using BeautifulSoup—a popular Python tool that makes it easier to work with HTML. You also know the CSS selector for the exact item you want to grab. For example, you might say, “I want the text inside this particular
or .” The function takes these two things: the soup object (which is just the HTML you’ve already parsed), and the selector (which tells it what to look for). It uses soup.select\_one(selector) to find the first match. If it finds the element, it grabs the text inside, removes any unnecessary spaces at the beginning or end, and returns it. But what if the element doesn’t exist on the page? Instead of giving back a confusing None, the function simply returns "N/A". This is helpful because it keeps your data clean and consistent. Later on, if you’re working with a spreadsheet or a database, it’s much easier to handle missing values when they’re marked clearly as "N/A" rather than as blanks or errors. ### The Product Data Detective ``` def parse_product_data(html): """ Extracts product details from HTML content. Args: html (str): Raw HTML content of product page Returns: tuple: Contains (item_name, packing, sku, price, category) Notes: - Uses BeautifulSoup for HTML parsing - Employs CSS selectors for precise element targeting - Returns 'N/A' for missing data fields """ soup = BeautifulSoup(html, "html.parser") # Extract each product detail using CSS selectors item_name = extract_text(soup, "div.page-title-wrapper.product > h2 > span") packing = extract_text(soup, "div.product-info-price > div > div.stock.available > span") sku = extract_text(soup, "div.product-info-price > div > div.product.attribute.sku > div") price = extract_text(soup, "span.price") category = extract_text(soup, "div.product-page-brand-common-view > ul > li > a:nth-child(2)") return item_name, packing, sku, price, category ``` This method plays a central role in transforming messy raw HTML into clean, structured product data. Think of a web page as a complex puzzle full of information—we only want a few specific pieces, like the product name, packaging, SKU, price, and category. This method helps us pull out just those important parts in a neat, organized way. Here's how it works: First, it uses BeautifulSoup, a Python library that makes it easier to work with HTML, almost like giving us X-ray vision to see through the web page’s code. Once we have the HTML "decoded," the method calls the extract\_text() function multiple times. Each call is aimed at a different piece of product info—for example, one call might pull out the product’s name, while another grabs the price or SKU. Each of these values is picked out using CSS selectors, which work like directions telling the code exactly where to look in the HTML. It’s similar to saying, “Look inside this box, under that label, and grab whatever text you find.” Once all the fields are collected, the method bundles them into a tuple, which is like a little package of information that keeps all the product details together and in the correct order. If any field happens to be missing on the page, it automatically fills in "N/A" to keep things consistent—so we never end up with holes or mismatched records. This structured format is really helpful when you later want to save the data into a database or display it neatly elsewhere. Everything is in place, and you know exactly what to expect—making your data clean, predictable, and easy to work with. ### The Individual Page Scraper ``` async def scrape_product_page(page, link, date): """ Scrapes individual product page and saves extracted data. Args: page (Page): Playwright page object link (str): URL of the product to scrape date (str): Date when the link was originally found Notes: - Uses 60-second timeout for page loading - Validates extracted data before saving - Includes comprehensive error handling - Updates scraped status after successful extraction """ try: await page.goto(link, timeout=60000) # 60 seconds timeout await page.wait_for_load_state("networkidle") # Extract and parse product data content = await page.content() item_name, packing, sku, price, category = parse_product_data(content) # Validate essential data before saving if item_name == "N/A" and price == "N/A": print(f"[WARNING] No valid data found for {link}. Skipping...") return # Save data and update status save_product_data(link, item_name, packing, sku, price, category, date) update_scraped_status(link) print(f"[INFO] Scraped: {item_name} | Price: {price}") except Exception as e: print(f"[ERROR] Failed to scrape {link}: {e}") ``` This asynchronous method is like a dedicated worker that focuses on collecting detailed information from one specific product page at a time. It’s designed to handle modern, dynamic websites where content might take a moment to fully load. First, it receives three important pieces of information: a Playwright page object (which helps it interact with the web page like a real browser), the product's URL, and the date when this scraping attempt is happening. The method begins by visiting the product page using Playwright’s goto() function. It gives the page up to 60 seconds to load—this is especially useful when dealing with pages that take longer due to animations, pop-ups, or slow servers. It also waits until the network is quiet, meaning everything on the page (like images and scripts) has had time to finish loading. This ensures that all the product details are present before we start collecting them. Once the page is fully loaded, it grabs the raw HTML content and sends it over to another helper method called parse\_product\_data(). This helper organizes the messy HTML into neat product details like name, price, packaging, and so on. Before saving anything, the method checks to make sure the most important details—like the product name and price—are actually there. If either is missing, it skips saving that product and logs a warning. This prevents the database from being cluttered with incomplete or useless records. If all the required information is present, it proceeds to store the data in the database and updates the product's status so we know it’s already been processed. Along the way, if anything goes wrong—like the page doesn’t load or the data can’t be saved—it handles the issue gracefully and logs clear error messages. That way, you can figure out what went wrong without the whole scraping process breaking down. ### The Grand Orchestra of Data Collection Now we bring all the pieces together into a smooth, coordinated scraping workflow. Each function plays its part—fetching unprocessed links, visiting product pages, extracting clean data, and saving it to the database. Like a well-oiled machine, the system handles errors gracefully and keeps the process moving. This is where all our careful setup pays off, creating a reliable, automated pipeline for product data collection. #### The Product Collection Conductor ``` async def scrape_all_products(): """ Orchestrates the scraping of all unprocessed product pages. Notes: - Retrieves unscraped links from database - Launches headless browser instance - Processes each product page sequentially - Ensures proper cleanup of browser resources - Includes error handling and logging """ links = get_unscraped_links() if not links: print("[INFO] No new products to scrape.") return async with async_playwright() as p: browser = await p.chromium.launch(headless=True) page = await browser.new_page() for link, date in links: print(f"[INFO] Scraping {link}...") await scrape_product_page(page, link, date) await browser.close() ``` This method takes charge of the entire scraping process, acting like the conductor of our data-gathering operation. Its main job is to go through every product link that hasn’t been scraped yet and process them one by one. It all starts by calling a function named get\_unscraped\_links(). This function checks the database and returns a list of product URLs that haven’t been handled yet. If it turns out there are no more fresh links to scrape, the method politely prints a message and exits early. This prevents unnecessary work and saves time and resources. When there are links to work with, the real scraping begins. The method launches a browser—specifically, a lightweight, invisible (or “headless”) version of Chromium using Playwright. It then loops through each unscraped link, visiting each page and calling another function, scrape\_product\_page(), to extract the product’s information. This part of the code is asynchronous, which means it can handle tasks more efficiently by not waiting around while pages load. Instead, it keeps things moving, making the whole process faster and more responsive. Once all the links have been processed and every bit of information has been gathered and saved, the browser is closed. This final step is important because it releases system resources and ensures everything is tidy when the job is done. #### The Grand Finale: Our Main Function ``` async def main(): """ Main execution function that coordinates the scraping process. Flow: 1. Initializes database table 2. Scrapes all unprocessed products 3. Logs completion status Notes: - Asynchronous execution using asyncio - Provides feedback about operation progress """ create_data_table() await scrape_all_products() print("[INFO] Scraping completed.") if __name__ == "__main__": asyncio.run(main()) ``` The main() function serves as the central starting point of the entire scraping operation—it’s where the process truly begins. It starts with a foundational step: calling create\_data\_table(). This ensures that the database is set up correctly, with the appropriate table(s) ready to receive incoming product data. Without this, the scraping process would lack a place to store its results, making this an essential initial task. After preparing the database, main() moves on to its core responsibility—kicking off the actual data collection. It does this by invoking scrape\_all\_products(), the orchestrated method that handles scraping each individual product link. Since scraping involves asynchronous operations (to make the process faster and more responsive), Python’s [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(main()) is used when calling the main function. This allows asynchronous tasks to execute smoothly in a controlled event loop. Finally, once all product pages have been visited and the data is safely stored, the function prints a success message. This serves as a confirmation that everything went according to plan—the data pipeline ran, the information was captured, and the system operated as expected. ## Conclusion By following this carefully organized process, we've built a dependable system that can automatically collect product data from Family food center. Each step—from setting up a clean and reliable database to using tools like Playwright and BeautifulSoup—works together to keep everything running smoothly and accurately. Along the way, we’ve also added checks and messages to help us monitor progress and avoid collecting the same data twice. One of the key strengths of this setup is that it uses asynchronous programming. This means we can load and process several web pages at once, which saves a lot of time and makes the whole system faster and more scalable. Whether we’re working with just a few products or thousands, this structure holds up well and doesn’t need constant attention. In the end, what we’ve created is more than just a scraper—it’s a solid foundation for collecting useful data in a smart, repeatable way. This opens the door to doing more with the information we gather, like tracking prices, studying market trends, or even making informed business decisions. It’s a practical and powerful tool for anyone looking to turn web data into real-world insight. AUTHOR I’m Shahana, a Data Engineer at Datahut, where I specialize in building smart, scalable data pipelines that transform messy web data into structured, usable formats—especially in domains like retail, e-commerce, and competitive intelligence. [At Datahut](https://datahut.co/?ref=blog.datahut.co#contact), we help businesses across industries gather valuable insights by automating data collection from websites, even those that rely on JavaScript and complex navigation. In this blog, I’ve walked you through a real-world project where we created a robust web scraping workflow to collect product information efficiently using Playwright, BeautifulSoup, and SQLite. Our goal was to design a system that handles dynamic pages, pagination, and data storage—while staying lightweight, reliable, and beginner-friendly. If your team is exploring ways to extract structured product or pricing data at scale—or if you're just curious how web scraping can support smarter decisions—feel free to connect with us using the chat widget on the right. We’re always excited to share ideas and build custom solutions around your data needs. FAQ SECTION 1\. What tools do you need to scrape product data from Family Food Centre? - Playwright — to automate browser actions like scrolling and clicking "Next" on paginated pages - BeautifulSoup — to parse the loaded HTML and extract product details like name, price, and packaging - SQLite — to store scraped links and product data locally in an organized database - asyncio — to run the scraping process asynchronously, making it faster and more efficient - Python's datetime module — to timestamp each record for tracking when data was collected 2\. Why is the scraping process split into two phases? - Phase 1 focuses only on collecting product URLs from the category listing pages - This avoids overloading the scraper by separating link discovery from data extraction - Phase 2 visits each saved URL individually to pull detailed information like price, SKU, and packing - Splitting phases makes it easier to resume if the scraper stops midway — unscraped links remain in the database with a scraped = 0 flag - It also makes debugging simpler since each phase can be tested and fixed independently 3\. How does the scraper handle multiple pages on the Family Food Centre website? - Playwright loads the first page and waits for it to reach a networkidle state before extracting links - The scraper then looks for a "Next" button using a CSS selector specific to the site's pagination structure - If the button is found, it is clicked automatically and the next page is allowed to fully load - This loop continues until no "Next" button is found, signaling the last page has been reached - A page counter prints progress updates so you can monitor how many pages and links have been scraped 4\. How is scraped data stored and tracked to avoid duplicates? - All product links are saved in a products table in an SQLite database with a scraped column defaulting to 0 - Detailed product info (name, price, SKU, packing, category) is stored separately in a product\_data table - Once a product page is successfully scraped, its scraped flag is updated to 1 using the update\_scraped\_status() function - On every new run, get\_unscraped\_links() queries only rows where scraped = 0, so already-processed links are never revisited - Each record is also timestamped with the scrape date for historical tracking and analysis 5\. What happens if a product page fails to load or has missing data? - Every scraping function is wrapped in a try/except block to catch errors gracefully without crashing the entire script - If a page times out (beyond the 60-second limit), the error is logged and the scraper moves on to the next link - Before saving, the script checks whether both item\_name and price return "N/A" — if so, that product is skipped with a warning - Missing individual fields (like SKU or packing) are stored as "N/A" to keep the database consistent and query-friendly - All errors are printed to the console with the specific link that caused the issue, making it easy to identify and retry problem URLs ### CeraVe vs. Cetaphil on Amazon: Who Really Wins and Why URL: https://www.blog.datahut.co/post/cerave-vs-cetaphil-on-amazon/ Last updated: 2026-09-07T09:43:02.000Z ## Why CeraVe and Cetaphil Dominate Amazon Skincare The skincare aisle on [Amazon](https://www.blog.datahut.co/post/amazon-product-data-scraping/) is not a level playing field. Two brands - [CeraVe](https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/) and [Cetaphil](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/) \- have carved out a permanent duopoly in the [dermatologist-recommended ](https://www.aad.org/?ref=blog.datahut.co)segment, and their dominance is built on data-driven marketplace strategies that go far deeper than good formulations. This analysis breaks down how both brands position, price, discount, and emotionally connect with customers - using 226 active products and over 2.3 million real customer reviews as the raw material. This analysis draws on our deep dives into each brand individually: [Cetaphil on Amazon 2026](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/) (143 products, 707K reviews) and [CeraVe on Amazon 2026](https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/) (83 products, 1.6M reviews). ## Market Overview: Two Giants, [Two Marketplace Strategies](https://www.blog.datahut.co/post/competitive-market-intelligence/) Both brands serve the same core demographic - people with sensitive skin, parents, and those seeking clinically reliable skincare. Yet their [Amazon architectures](https://www.blog.datahut.co/post/amazon-product-data-scraping/) are built on fundamentally different philosophies. ![CeraVe vs. Cetaphil](https://www.blog.datahut.co/content/images/2026/07/img-323.png.webp) ![cerave vs cetaphil](https://www.blog.datahut.co/content/images/2026/07/img-324.png.webp) ## SKU Strategy: Depth vs. Breadth On Amazon, SKU count correlates directly with search visibility. More products means more impressions, more search query coverage, more opportunities to capture a sale. But CeraVe's performance proves that raw SKU count isn't the only path to dominance. Lotion and Cream categories dominate both catalogs: 53% of CeraVe's catalog and 49.7% of Cetaphil's are in these two moisturizer formats. Both brands agree that moisturization is the primary purchase driver — they just pursue dominance differently. - Liquid formats (CeraVe: 11 SKUs, Cetaphil: 17 SKUs) form a meaningful secondary tier — pump-format cleansers and targeted treatments bridging skincare and body care. - Gel formats (CeraVe: 10 SKUs, Cetaphil: 11 SKUs) serve oily and combination skin customers who demand lightweight, non-greasy solutions. - Specialty formats diverge sharply: CeraVe stays experimental (Serum, Oil, Patch, Paste) while Cetaphil extends into Wipes, Foam, Bar, and Drop formats for full lifecycle coverage including baby and travel. ## Pricing Tier Analysis: How CeraVe and Cetaphil Position Products How each brand distributes products across price tiers reveals their thinking about customer segments, from sensitive skin newcomers buying their first cleanser, to loyal customers investing in specialised treatments. ![Pricing Tier Analysis: How CeraVe and Cetaphil Position Products](https://www.blog.datahut.co/content/images/2026/07/img-325.png.webp) ![cerave vs cetaphil amazon analysis](https://www.blog.datahut.co/content/images/2026/07/img-326.png.webp) The premium tier ($30+) is a shared vulnerability. CeraVe products above $30 average a 3.54 rating; Cetaphil's average 3.47\. Consumers view these [clinical brands as accessible ](https://www.nih.gov/?ref=blog.datahut.co)daily necessities, not luxury items — and the review data says so explicitly. Cetaphil's exposure is larger, with 27 premium SKUs averaging $53.34. ## Discount Strategy: Who Discounts What and Why Discount behavior is one of the most honest strategic signals in e-commerce. It reveals which products a brand is actively pushing, where customer acquisition costs are high, and where established trust makes promotional spend unnecessary. ![cerave and cetaphil discounts](https://www.blog.datahut.co/content/images/2026/07/img-327.png.webp) ### The Lotion Signal: Confidence in Full View The most telling number in the entire discount analysis: CeraVe's Lotion carries only a 9.87% average discount; Cetaphil's Lotion only 9.54%. Their bestselling formats don't need promotional support. Brand reputation and repeat purchase habits drive volume without discounting. ## Review Concentration & Customer Loyalty ![Review Concentration & Customer Loyalty of cerave and cetaphil](https://www.blog.datahut.co/content/images/2026/07/img-328.png.webp) Understanding how each brand's total reviews are distributed across its catalog reveals whether success depends on a few hero products or a broader foundation of consistent performers. This distinction matters enormously for long-term marketplace resilience. ## Customer Psychology: Why People Buy CeraVe vs Cetaphil Star ratings tell us how satisfied customers are. Reviewing attribute data tells us why. The emotional triggers that convert browsers into buyers differ fundamentally between these two brands - and that gap is where smart strategy lives. ![Customer Psychology: Why People Buy CeraVe vs Cetaphil ](https://www.blog.datahut.co/content/images/2026/07/img-329.png.webp) The customer journeys mapped by this data are distinct and telling: - CeraVe's journey: Moisturizing attracts trial → Effectiveness drives loyalty → Value enables repeat purchase → Compatibility builds long-term trust. - Cetaphil's journey: Protection attracts trial → Gentleness keeps customers → Softness delivers sensory confirmation → Value for money seals the repeat purchase. For the full data behind each brand's strategy, explore the individual analyses: [CeraVe on Amazon](https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/) and [Cetaphil on Amazon](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/). ## Final Verdict: Two Different Moats [CeraVe](https://www.cerave.com/?ref=blog.datahut.co) wins through clinical depth and concentrated power. By investing heavily in a smaller portfolio of Creams and Lotions, it has built an impenetrable wall of 1.6 million reviews. Its 4.45 average rating across 83 SKUs reflects a brand that consistently delivers on the promise of visible hydration. [Cetaphil](https://www.cetaphil.in/?ref=blog.datahut.co) wins through algorithmic breadth and safety positioning. Its 143 SKUs capture diverse long-tail searches that CeraVe's concentrated catalog can't reach. Its customers are deliberate, protective buyers managing sensitive skin - and fragrance-free formulation is a conversion trigger, not just a product feature. Both brands share one clear vulnerability: the premium price tier. Consumers see these as accessible daily necessities. Any product pushing above $30 faces an uphill battle against that perception - and the review ratings prove it. Connect with[ ](https://www.datahut.co/?ref=blog.datahut.co)[Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the information you need, hassle-free. ![Managed web scraping service](https://www.blog.datahut.co/content/images/2026/07/img-330.png.webp) ## FAQ SECTION ## 1\. Which is better on Amazon: CeraVe or Cetaphil? - CeraVe performs better in engagement, with over 1.6M reviews and higher average ratings - Cetaphil offers a wider product range, covering more niche and sensitive-skin needs - CeraVe wins on effectiveness and customer trust signals - Cetaphil wins on gentleness and safety-focused positioning - The better choice depends on skin type and specific needs ## 2\. Why is CeraVe more popular on Amazon than Cetaphil? - Strong review concentration on top products - Higher average ratings across fewer SKUs - Focus on moisturization and visible results - Better product-level dominance instead of catalog expansion - Builds trust through consistent performance and social proof ## 3\. Why do customers choose Cetaphil over CeraVe? - Known for gentle, non-irritating formulations - Preferred for sensitive and reactive skin types - Strong association with fragrance-free skincare - Wider product formats (wipes, bars, baby care) - Focus on skin protection and comfort ## 4\. Do discounts affect sales of CeraVe and Cetaphil on Amazon? - Bestselling products from both brands have low discount dependency (\~9–10%) - Indicates strong brand trust and repeat purchase behavior - Discounts are mainly used for:New product launchesLess popular SKUsCompetitive categories - Core products sell well without heavy promotions ## 5\. What is the biggest weakness of both CeraVe and Cetaphil? - Both struggle in the premium price segment ($30+) - Customers view them as daily essentials, not luxury skincare - Higher-priced products receive:Lower ratingsReduced perceived value - Indicates a clear pricing ceiling in consumer mindset ### CeraVe on Amazon 2026: Best Sellers, Pricing & Review Analysis URL: https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/ Last updated: 2026-09-07T09:43:04.000Z CeraVe's presence on Amazon's [skincare category](https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/) is one of the most strategically disciplined on the platform. With 83 actively [priced products](https://www.blog.datahut.co/post/price-comparison-on-amazon-how-web-scraping-helps-companies-win-the-e-commerce-game/), over 1.6 million verified customer reviews, and an average rating of 4.45 out of 5, the brand has built a catalog that prioritizes depth over breadth and the data proves it's working. At an average sale price of $21.93, CeraVe on Amazon sits in the accessible mid-range, but unlike [competitors who flood listings](https://www.blog.datahut.co/post/competitive-market-intelligence/) with variations, it competes with precision. This analysis decodes the mechanics behind [CeraVe's Amazon strategy](https://www.statista.com/topics/11222/amazon-beauty/?ref=blog.datahut.co): from SKU distribution across product forms, to a disciplined discounting philosophy, to the exact voice-of-customer signals that drive repeat purchase. ## CeraVe on Amazon 2026: Key Data at a Glance Based on analysis of 83 actively priced products and 1,618,499 verified customer reviews, here are the headline findings from the CeraVe Amazon analysis: ![CeraVe on Amazon 2026](https://www.blog.datahut.co/content/images/2026/07/img-268.png.webp) CeraVe's sibling brand in the dermatologist-recommended segment, Cetaphil, follows a different path - broader catalog, softer positioning. Read our [Cetaphil Amazon 2026 analysis](https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/) to see how the two strategies compare. ## How Does CeraVe Perform on Amazon? Ratings, Pricing & Review Data CeraVe's 83 active products span 10+ distinct product forms, generating 1,618,499 total [customer reviews](https://www.blog.datahut.co/post/scraping-amazon-reviews-python-scrapy/), establishing the brand as one of the most reviewed [dermatologist-recommended](https://www.cerave.com/about-cerave/developed-with-dermatologists?ref=blog.datahut.co) skincare lines on [Amazon](https://www.blog.datahut.co/post/how-to-do-amazon-market-research/). This review density signals strong repeat-purchase behavior and sustained customer trust. An average rating of 4.45 out of 5 represents consistent quality across the catalog. The $21.93 average sale price positions CeraVe squarely in the accessible mid-range, appealing to its core audience of sensitive skin consumers, eczema sufferers, and dermatologist-referred shoppers. ### Top Product Forms by Review Volume Lotion leads with 663,242 total reviews across 20 SKUs. These lightweight, non-greasy formulas are the gateway product for customers discovering the brand - high trial volume, high repeat purchase, and the deepest category moat. Cream holds the second position with 546,069 total reviews across 24 SKUs, reflecting both diversity of skin conditions served and CeraVe's mastery of barrier-repair formulation. Cream's 24 SKUs make it the largest single product form by count, yet it maintains exceptional quality consistency. Liquid contributes 11 SKUs and includes cleansers and toners, positioning CeraVe across the full daily skincare routine rather than just moisturization. Q: What is CeraVe's average rating and price on Amazon? CeraVe products on Amazon maintain an average rating of 4.45 out of 5, with an average sale price of $21.93, placing it in the accessible mid-range skincare segment where value perception and clinical efficacy converge most powerfully. ## How CeraVe Competes for Shelf Space: SKU Distribution Strategy On Amazon, a higher SKU count improves search visibility and sales potential. CeraVe prioritizes its strongest formats instead of addressing every skincare niche, reflecting a deliberate and defensible portfolio strategy. ### SKU Distribution Breakdown - Cream (24 SKUs) and Lotion (20 SKUs) account for 51.8% of CeraVe's portfolio. This demonstrates that depth, quality, and volume are more effective than an overly broad catalog. - Liquid (11 SKUs) forms a strong second tier, appealing to customers who prefer pump or liquid-format cleansers and treatments. This supports a full skincare routine offering. - Gel (10 SKUs) addresses the lightweight hydration segment, targeting customers with oily or combination skin who seek effective, non-greasy solutions. - Foam, Bar, Ointment, and Balm (2–3 SKUs each) address specialized needs such as exfoliation, hand care, and targeted hydration, while maintaining category focus. - Serum, Oil, Patch, Paste, and Wipes (1–2 SKUs each) reflect CeraVe's willingness to test adjacent markets while maintaining brand consistency and clinical efficacy. This approach is intentional. CeraVe deepens its presence in its strongest categories while maintaining a presence across the skincare ecosystem, encouraging customer loyalty once trust is established. ## CeraVe on Amazon: Which Products See the Deepest Discounts? Discount activity on Amazon signals which products a brand is prioritizing, where customer acquisition costs are high, and where established trust reduces the need for promotions. [CeraVe's discount patterns are deliberate and targeted.](https://www.blog.datahut.co/post/how-scraping-amazon-data-can-help-you-price-your-products-right/) ![Cerave discounts on Amazon](https://www.blog.datahut.co/content/images/2026/07/img-269.png.webp) ### Key Takeaway CeraVe employs two parallel discount strategies simultaneously: significant promotional investment in developing or premium-adjacent categories (Oil, Patch, Bar) and maintaining perceived value in its core category (Lotion). At an overall 13.27% catalog average, well below the 20–30% discount levels common on Amazon, CeraVe's restraint reflects genuine [brand equity](https://www.investopedia.com/terms/b/brandequity.asp?ref=blog.datahut.co), not pricing indiscipline. Q: How does CeraVe handle pricing and discounts on Amazon? CeraVe averages a 13.27% discount across its catalog, with the highest promotions (30–33%) reserved for Oil and Patch to encourage trial, while flagship Lotion products are protected at under 10% discount, a clear signal that demand in core categories is driven by trust, not price incentives. ## Price Tier Analysis: Where CeraVe Wins and Where It Struggles The distribution of CeraVe's 83 active products across price tiers reveals its approach to customer segmentation from first-time buyers seeking affordable daily moisturizers to returning customers managing chronic skin conditions. ### Budget Tier (≤$15)- 40 SKUs, Highest Rating in Catalog The budget segment holds 40 SKUs and averages a 4.57 rating, the highest of any price band. CeraVe's entry-level products are not stripped-down compromises but full-featured solutions at accessible prices. This is a significant competitive advantage: new customers are won at the lowest price point and retained at the highest satisfaction levels. ### Mid-Range Tier ($15–$30) - 34 SKUs, Sweet Spot Cetaphil performs best in this tier by SKU count, with 34 products averaging a 4.54 rating. This is CeraVe's core strategic positioning and strongest value proposition, where clinical efficacy at an attainable price resonates most powerfully with shoppers. Protecting and expanding this tier is essential to maintaining category leadership. ### Premium Tier ($30+) - 9 SKUs, Biggest Vulnerability Products in this tier average only a 3.54 rating, a significant 1-point gap below the budget tier. Customers at this price point benchmark CeraVe against specialized clinical and luxury skincare brands, and expectations are not consistently met. Strategic Implication: The premium rating gap is a positioning issue, not just a product issue. CeraVe's brand identity is built on clinical accessibility, not luxury. Addressing this gap through product reformulation, clearer premium differentiation, or improved customer expectation management represents the most immediate growth opportunity in the catalog. Q: What is the biggest vulnerability in CeraVe's Amazon catalog? Products priced above $30 average only a 3.54 star rating, a full point below the budget tier (4.57) and mid-range tier (4.54). This is CeraVe's most measurable weakness and clearest short-term growth opportunity for product and marketing teams. Both CeraVe and Cetaphil hit the same ceiling above $30\. We break down exactly who handles it better and why in our [CeraVe vs. Cetaphil head-to-head analysis](https://www.blog.datahut.co/post/cerave-vs-cetaphil-on-amazon/). ## CeraVe Best-Sellers vs. Full Catalog: Where Do 1.6M Reviews Actually Come From? Analyzing CeraVe's review concentration across 83 active SKUs reveals whether the brand depends on a few blockbuster products or maintains genuine portfolio depth. - No single product dominates the catalog - the top SKU holds a meaningful but not outsized share of total reviews, reflecting a balanced portfolio where success is distributed. - The top 10 products command 59.58% of all review volume - concentrated enough to anchor the business, distributed enough to maintain resilience against disruption. - The remaining 70+ products collectively share over 40% of all review volume, signaling that CeraVe's extended catalog serves real, recurring needs - not just tail SKUs holding space. - For competitors, the top 10 CeraVe products present a significant social proof barrier that requires strong differentiation or sustained investment to overcome. - This distributed structure also creates natural launch opportunities -new SKUs can carve out meaningful share without cannibalizing existing best-sellers. ## Why Do Customers Buy CeraVe? What Amazon Reviews Really Say Star ratings tell us how satisfied customers are. Review attribute data tells us why. By analyzing the specific features and benefits customers mention most across CeraVe's catalog, we unlock what actually drives purchasing decisions and what sustains loyalty. ![Cerave amazon review](https://www.blog.datahut.co/content/images/2026/07/img-270.png.webp) ### The Customer Decision Flywheel Together, these attributes map the complete customer journey: Moisturizing attracts trial, Effectiveness drives loyalty, Value enables repeat purchase, and Compatibility builds long-term trust. This is the flywheel CeraVe executes consistently and it maps directly to the brand's dermatologist-backed positioning. Q: What features do customers value most in CeraVe products? Moisturizing is the [#1](https://www.blog.datahut.co/blog/hashtags/1) purchase driver with 12,088 review mentions. Effectiveness and Value for Money follow closely. Customers choose CeraVe specifically for its proven hydration, clinical efficacy, and accessible pricing, not for luxury appeal or trend relevance. ## CeraVe Amazon Ratings by Product Form: Full Breakdown CeraVe maintains consistently high satisfaction scores across all product forms - a reflection of rigorous quality standards and formulation discipline. The gap between the highest and lowest rated formats is notably narrow. ![Cerave amazon ratings](https://www.blog.datahut.co/content/images/2026/07/img-271.png.webp) Maintaining all product forms above 4.45 at catalog scale is a significant quality achievement. The narrow rating band (4.70 to 4.55) confirms that CeraVe's brand promise holds across all formats, not just in the core moisturizer lines that generate the most visibility. Q: Which CeraVe product form has the highest per-product customer engagement on Amazon? Serum leads with the highest average reviews per product, followed by Bar and Balm, formats with focused, loyal customer bases generating disproportionate engagement relative to their SKU counts. ## CeraVe Amazon Strategy: What the Data Really Tells Us ### What CeraVe Is Doing Right - Focusing on creams and lotions provides CeraVe with strong search visibility, creating advantages that are difficult for new competitors to match. - CeraVe maintains disciplined discounting by limiting lotion discounts to under 10% and investing in promotions for newer or higher-priced products, demonstrating a sophisticated[ pricing strategy](https://www.blog.datahut.co/post/price-comparison-on-amazon-how-web-scraping-helps-companies-win-the-e-commerce-game/). - Customer reviews highlight moisturizing, effectiveness, and value for money, which align with CeraVe's clinical positioning and confirm the brand delivers on its promises. - CeraVe consistently maintains product ratings above 4.45 across all formats, demonstrating a high level of formulation discipline that is difficult to sustain at scale. ### Where CeraVe Must Improve - The premium rating gap, with an average of 3.54 for products priced above $30, indicates a clear disconnect between customer expectations and product performance at higher price points. - The oil segment remains underserved. Despite significant promotional investment, limited SKUs prevent CeraVe from fully capturing the demand being generated. - Serum and specialty formats demonstrate strong engagement per product. To support premium pricing above $30, clearer management of customer expectations is needed. ![Managed web scraping services](https://www.blog.datahut.co/content/images/2026/07/img-272.png.webp) ## Frequently Asked Questions (FAQs) 1\. How many products does CeraVe have on Amazon? CeraVe has 83 active products across 10+ product formats on Amazon. - The catalog includes lotions, creams, gels, serums, and cleansers. - It covers daily care, treatment-focused, and specialty SKUs. - This represents a focused catalog that prioritizes depth over breadth. - The range supports multiple skin types, conditions, and routine steps. 2\. What is CeraVe's best-performing price tier on Amazon? The Budget tier (under $15) has the highest average rating at 4.57, while the Mid-Range tier ($15–$30) is the strongest by SKU volume and strategic importance. - The Budget tier includes 40 SKUs with an average rating of 4.57, reflecting the highest satisfaction in the catalog. - The Mid-Range tier consists of 34 SKUs with a 4.54 average rating, representing the core value proposition. - Both tiers outperform the Premium tier (over $30) by more than one full rating point. - The $15–$30 range offers the best alignment of pricing, performance, and customer expectations. 3\. Does CeraVe rely on hero products or a broad catalog? CeraVe relies on the strength of a broad catalog, anchored by two dominant product forms rather than a single hero SKU. - The top 10 products account for 59.58% of all reviews, indicating a balanced portfolio rather than dependence on a single blockbuster. - The remaining 70 or more products share over 40% of the total review volume. - This demonstrates genuine portfolio depth and low dependency on any single SKU. - This approach creates a resilient and defensible catalog structure that is challenging for competitors. 4\. How disciplined is CeraVe's discounting strategy on Amazon? CeraVe's discount strategy is moderate and highly intentional, with an average discount of 13.27% across the catalog. - The deepest discounts, ranging from 30% to 33%, are reserved for Oil and Patch products to encourage trial. - The flagship Lotion is discounted by only 9.87%, signaling strong customer demand. - This rate is well below the Amazon category average of 20% to 30%, indicating genuine brand equity. - Discounting is used as a growth investment rather than as a tool for generating demand. 5\. What are the top customer review attributes for CeraVe? Customers prioritize three core attributes: Moisturizing, Effectiveness, and Value for Money. - Moisturizing is mentioned 12,088 times and is the top purchase driver across the catalog. - Effectiveness is cited in 9,700 mentions, reflecting customers' expectations for visible, proven results. - Value for Money appears in 5,980 mentions, highlighting the perception of premium quality at an accessible price point. - These three attributes align directly with CeraVe's brand positioning and marketing claims. 6\. What is the biggest weakness in CeraVe's Amazon positioning? CeraVe's premium tier (over $30) averages only a 3.54-star rating, representing the most significant and actionable gap in the catalog. - Only nine SKUs exceed $30, but they significantly impact overall portfolio perception. - A one-point rating gap compared to the budget tier highlights misaligned customer expectations. - CeraVe's accessible, clinical identity makes it difficult to justify luxury price points. - This issue can be addressed through product reformulation, clearer premium positioning, or improved expectation management. 7\. What data methodology was used in this CeraVe Amazon analysis? This analysis is based on structured Amazon data collected and processed by Datahut. - Covers 83 active SKUs and 1,618,499 verified customer reviews - Includes pricing, discount rates, ratings, and product type classifications - Uses web scraping pipelines for data extraction and normalization - Applies AI tagging for review attribute analysis and customer signal mapping ### Carrefour UAE Web Scraping Case Study: Tracking Vegetable Prices, Discounts & Availability Over Time URL: https://www.blog.datahut.co/post/carrefour-uae-web-scraping-case-study/ Last updated: 2026-09-07T09:43:06.000Z Ever wondered what grocery data scraping can reveal about real-time pricing and product availability? In this case study, we scrape the [Carrefour UAE](https://www.carrefouruae.com/mafuae/en?ref=blog.datahut.co) [vegetables category](https://www.carrefouruae.com/mafuae/en/c/F11660500?ref=blog.datahut.co) over five days to analyze: - Price fluctuations - Discount patterns - Product availability trends This guide walks through a complete Carrefour UAE Web Scraping pipeline, from collecting product URLs to extracting structured data and preparing it for analysis using Python. ## Step-by-Step Web Scraping Workflow for Carrefour UAE ### Step 1: Scraping Product URLs from Carrefour UAE Vegetables Category Before any meaningful analysis can happen, the most important first step is collecting the right product page links, and this blog begins by focusing on how URLs were gathered from the vegetables section of the [Carrefour UAE](https://www.blog.datahut.co/post/web-scraping-carrefour-data-using-python-and-selenium/) website in a clean and structured way. Think of this step as creating a shopping list before entering a store—without knowing which aisles to visit, nothing else can move forward smoothly. The vegetables category page acts as the entry point, where dozens of product cards appear as the page loads and expands, each hiding a link that leads to a detailed product page. Instead of copying these links manually, an automated script carefully opens the category page, waits for the content to appear, and scans each product block to extract its URL while ignoring duplicates. As more items load through scrolling or the “Load More” action, the script continues gathering links until the full list is complete, saving them in a structured database and a backup file for safety. This approach turns a repetitive manual task into a reliable process, ensuring that every vegetable product page is captured once and stored neatly, laying a strong foundation for deeper data extraction and time-based analysis in later stages. ### Step 2: Extracting Product Data from Carrefour UAE Product Pages Once the list of vegetable product URLs was safely stored, the next step was to visit each link and gently collect the details hidden inside those pages, turning simple addresses into meaningful data. This stage can be compared to walking through each aisle after noting down its location, taking time to observe what is actually on the shelf. Using the URLs saved in the database, the script opened every product page one by one, allowed the page to load naturally, and handled small interruptions like cookie popups so the content could be read clearly. Each page was then converted into a structured format, making it easy to extract important information such as the vegetable name, price, discount, and availability without confusion. As the data was collected, it was stored back into the database and marked as processed, ensuring the same page would not be visited again in future runs, while a copy was also saved in a JSON file for easy review or sharing. Care was taken to handle missing details or slow-loading pages gracefully, [logging issues ](https://www.loggly.com/ultimate-guide/python-logging-basics/?ref=blog.datahut.co)instead of stopping the process, which allowed the scraper to continue smoothly across many product pages. By the end of this step, the project moved from a simple list of links to a well-organized dataset that clearly described each vegetable product on the Carrefour website, setting the stage for deeper insights and time-based analysis. ### Step 3: Data Cleaning and Preparation Using OpenRefine After collecting raw vegetable product data from the Carrefour website, the next important step was cleaning it so the information could be trusted and easily analyzed, and this is where OpenRefine became especially useful. The process began by opening [OpenRefine](https://southampton-rsg.github.io/openrefine-data-cleaning/aio.html?ref=blog.datahut.co) and uploading the scraped file, which immediately displayed the data in a clear, spreadsheet-like view that made small issues easy to spot. Price values often contained currency symbols, extra dots, or commas that could cause problems during analysis, so these were carefully removed or standardized to ensure every price followed the same clean format. Some columns appeared out of order or contained mixed values, so they were rearranged and refined to keep related information grouped together in a logical way. Duplicate product URLs, which can quietly appear during repeated scraping runs, were identified and removed to avoid counting the same vegetable more than once. Throughout this process, OpenRefine acted like a smart cleaning desk, allowing each adjustment to be previewed before applying it, which helped maintain accuracy. By the end of this stage, the once-raw dataset was transformed into a neat, consistent, and analysis-ready table, making it far easier to explore trends and patterns in the [Carrefour UAE](https://www.blog.datahut.co/post/decoding-carrefour-s-television-marketplace-trends-and-insights/) vegetables data with confidence. ## All-in-One Tool-sets for Smarter Data Extraction A smooth [scraping](https://www.okta.com/identity-101/data-scraping/?ref=blog.datahut.co) workflow depends heavily on choosing the right tools, and this project brings together a set of Python libraries that make the entire journey—from opening a webpage to saving clean product details—feel steady and manageable even for someone new to data extraction. Browser automation plays a central role here, with [Playwright](https://www.lambdatest.com/playwright?ref=blog.datahut.co) and Selenium stepping in to load each vegetable product page just like a real user would, allowing the script to scroll, wait, and interact naturally with different elements. To help these automated actions blend in more smoothly, Playwright Stealth adjusts small browser behaviors so the site treats the scraper like a regular visitor rather than a bot. Once a page loads, BeautifulSoup steps forward, turning the raw HTML into a readable structure that makes it easy to extract names, prices, discounts, and other details without wrestling with messy code. For organizing everything behind the scenes, sqlite3 stores product links in a lightweight database, keeping track of what has already been processed, while JSON provides a simple backup format that can be opened or shared with ease. Supporting libraries like logging and datetime help document each step and track when data is collected, and modules such as [random](https://www.w3schools.com/python/ref%5Fmodule%5Frandom.asp?ref=blog.datahut.co) and [time](https://www.geeksforgeeks.org/python/python-time-module/?ref=blog.datahut.co) add natural delays that help the scraper behave more realistically. Even elements like Firefox browser options and Selenium's waiting functions contribute to stability by giving each product page enough time to load fully before any extraction begins. All of these tools work together in a calm, coordinated way, creating a pipeline that handles dynamic pages, organizes the collected information, and keeps the entire process clear and easy to understand. ## Step 1: Scraping Product URLs from the Vegetables Section ### Importing Libraries ``` IMPORTS import sqlite3 import json import logging from datetime import datetime from playwright.sync_api import sync_playwright from playwright_stealth import stealth_sync from bs4 import BeautifulSoup import time ``` The code uses a few essential [Python tools](https://www.blog.datahut.co/post/web-scraping-tools/)—such as Playwright for loading the webpage, BeautifulSoup for reading the HTML, and SQLite for storing the scraped results—to collect basic product details from the site. ### Understanding the Logging Setup ``` LOGGING SETUP logging.basicConfig( filename='/home/anusha/Desktop/DATAHUT/carrefouruae/Log/carrefour.log', filemode='a', format='%(asctime)s - %(levelname)s - %(message)s', level=logging.DEBUG ) """This block configures Python's built-in logging module to track the scraper’s runtime activity, errors, and debug information""" ``` The [logging](https://docs.python.org/3/howto/logging.html?ref=blog.datahut.co) setup acts like a small diary for the scraper, quietly recording what happens during the run, and the main parameters simply decide where the log file should be saved, how new entries are added, what details each message includes, and the level of information that needs to be captured, creating a clear and friendly trail that helps trace errors or unusual behavior while working with data from the Carrefour UAE vegetables section ### SQLite Table for Storing Scraped Vegetable Product URLs ``` DATABASE SETUP conn = sqlite3.connect('/home/anusha/Desktop/DATAHUT/carrefouruae/Data/carrefour.db') c = conn.cursor() c.execute('''CREATE TABLE IF NOT EXISTS product_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, scraped_date TEXT, processed INTEGER DEFAULT 0 )''') conn.commit() """This block connects to a SQLite database and ensures that a table for storing scraped product URLs exists""" ``` The database setup acts as a small storage room where each scraped vegetable product link from the Carrefour UAE section can be safely kept for later use, and the code simply connects to a SQLite file, opens a cursor to communicate with it, and creates a table only if it does not already exist; the table includes an automatically increasing ID, a place to store each product URL, a date field to remember when the link was collected, and a small marker that shows whether the data from that URL has been processed, allowing the scraping workflow to run in an organised and friendly manner while working with the vegetables page at the [base\_url ](https://developer.mozilla.org/en-US/docs/Learn%5Fweb%5Fdevelopment/Howto/Web%5Fmechanics/What%5Fis%5Fa%5FURL?ref=blog.datahut.co)and target\_url provided earlier. ### JSON Backup System for Product URLs ``` JSON SAVE SETUP json_file = '/home/anusha/Desktop/DATAHUT/carrefouruae/Data/carrefour_urls.json' try: with open(json_file, 'r') as f: json_data = json.load(f) except FileNotFoundError: json_data = [] """This block manages backup storage of scraped product URLs in a JSON file. It ensures that previously scraped data is preserved between script runs""" ``` The [JSON](https://www.w3schools.com/whatis/whatis%5Fjson.asp?ref=blog.datahut.co) save setup works like a small notebook that keeps a backup copy of every vegetable product URL collected from the Carrefour UAE section, and the code first points to a file where the data should be stored, then attempts to open it so any previously saved information can be loaded back into memory; if the file is not found, an empty list is created instead, allowing the script to start fresh without errors. ### Basic HTTP Headers for Scraping ``` HEADERS headers = { "Accept-Encoding": "gzip, deflate, br", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:136.0) Gecko/20100101 Firefox/136.0", "Connection": "keep-alive", } """This dictionary stores the HTTP headers that will be sent along with requests made by the browser (Playwright)""" ``` The [headers](https://www.zenrows.com/blog/web-scraping-headers?ref=blog.datahut.co) dictionary acts like a small introduction the scraper gives when requesting a page from the Carrefour UAE vegetables section, and each entry plays a simple role in helping the website understand how the request should be handled; the Accept field tells the server what kinds of content can be received, the User-Agent identifies the browser style being used, the Connection field keeps the link open for smoother communication, and the encoding line explains how compressed data can be handled, creating a cleaner and more reliable interaction while fetching details from the base\_url and target\_url defined earlier. ### Core Scraper Function for Collecting Product Links ``` SCRAPER FUNCTION def scrape_carrefour(): """Scrape product URLs from Carrefour UAE (Vegetables category)""" base_url = "https://www.carrefouruae.com" target_url = "https://www.carrefouruae.com/mafuae/en/c/F11660500" with sync_playwright() as p: browser = p.firefox.launch(headless=False) context = browser.new_context() stealth_sync(context) page = context.new_page() try: page.goto(target_url, timeout=60000) logging.info("Navigated to Carrefour Vegetables page.") ``` The first part of the scraper sets the stage by preparing everything needed to read the vegetables section from the site, starting with two basic URL references that guide the browser to the right category page. The function opens a Playwright browser, adds stealth settings to reduce detection, and creates a new page where the content will load. Once the page is opened, the script gently navigates toward the vegetables section and logs the action as a way to record what is happening behind the scenes. This simple setup works like opening the front door of a store before exploring the shelves, giving a clear starting point before any real extraction begins. ``` # Accept cookies if popup appears try: page.wait_for_selector('#onetrust-accept-btn-handler', timeout=10000) page.click('#onetrust-accept-btn-handler') logging.info("Accepted cookies.") except Exception: logging.info("No cookie popup detected.") product_urls = [] last_height = 0 scraped_urls_in_run = set() # track URLs scraped in *this run* only while True: soup = BeautifulSoup(page.content(), 'html.parser') product_tags = soup.find_all('div', class_='max-w-[134px]') for tag in product_tags: a_tag = tag.find('a', href=True) if a_tag: product_url = base_url + a_tag['href'].split('?')[0] if product_url in scraped_urls_in_run: continue # skip already scraped in this run scraped_urls_in_run.add(product_url) scraped_date = datetime.now().strftime('%Y-%m-%d %H:%M:%S') ``` In the second part, the scraper handles small website interactions and starts gathering the product links that appear on the screen. It begins by checking whether a cookie popup shows up and, if it does, closes it so the page can load fully without interruptions. After that, the script creates a few helper lists to store the links found during the run and avoid duplicates. As the page scrolls, the HTML content is read with BeautifulSoup, and each product box is scanned to locate the anchor tag that carries the product’s path. The URL is cleaned, its timestamp is recorded, and the link is added to a temporary set that keeps track of what has already been captured, ensuring a smooth and organized scraping process. ``` # Insert into DB c.execute('INSERT INTO product_urls (url, scraped_date, processed) VALUES (?, ?, ?)', (product_url, scraped_date, 0)) conn.commit() # Append to JSON json_data.append({ "url": product_url, "scraped_date": scraped_date, "processed": 0 }) product_urls.append(product_url) logging.info(f"Scraped {len(product_urls)} URLs in this iteration.") ``` The third part focuses on storing every product link safely so it can be used later without needing to scrape again. Whenever a new URL is found, the [script inserts ](https://www.sqlite.org/lang%5Finsert.html?ref=blog.datahut.co)it into a SQLite table along with the exact time it was collected and a small marker showing that the link has not yet been processed. At the same time, a copy of the same information is added to the JSON backup list, creating a secondary record that protects the data even if the workflow is restarted. This dual-saving approach works like writing a note in two places—one in a structured database and another in an easy-to-read file—making the entire process more reliable. ``` # Attempt to click "Load More" try: load_more_selector = 'button:has-text("Load More")' page.wait_for_selector(load_more_selector, timeout=10000) page.click(load_more_selector) time.sleep(3) logging.info("Clicked Load More.") except Exception: logging.info("No more Load More button, ending scroll.") break logging.info(f"Total URLs scraped: {len(product_urls)}") except Exception as e: logging.error(f"Error while scraping: {e}") finally: browser.close() ``` This part controls the scrolling experience by checking whether the page offers a [“Load More” button](https://medium.com/@datajournal/pagination-in-web-scraping-627c528de91c?ref=blog.datahut.co) that reveals additional vegetable products. The script waits for the button, clicks it when available, pauses briefly to let new items appear, and logs the action for clarity. If the button does not show up, it simply means the page has reached the end of the list, so the loop breaks and the scraper finishes its work. This closing step feels like turning each page of a catalog until the last one appears, ensuring that no product link is missed while keeping the process clear and predictable. ### Saving the Final Data into a Clean JSON File ``` # Save JSON with open(json_file, 'w') as f: json.dump(json_data, f, indent=4) logging.info("Scraping completed and JSON saved.") ``` At the end of the scraping run, the script simply opens the JSON file in write mode, stores the updated list of collected product details in a neatly formatted structure, and logs a short message to confirm that the data has been safely saved for future use. ### Triggering the Vegetable Scraper Only When the Script Runs Directly ``` MAIN if __name__ == "__main__": scrape_carrefour() """ Main Execution Block: - When the script is run directly, execute the scraping process""" ``` The main block simply checks whether the file is being run on its own and, if so, starts the vegetable-scraping function, ensuring the process activates only during direct execution and not when the script is imported into another program. ## Step2: Collecting Detailed Information from Individual Product Pages ### Importing Libraries ``` IMPORTS import sqlite3 import json import os import random import time import logging from datetime import datetime from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.firefox.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException ``` This import block simply gathers the essential tools needed for the scraping workflow, bringing in modules for database storage, JSON backups, time tracking, logging, and HTML parsing, while [Selenium](https://www.geeksforgeeks.org/software-testing/selenium-basics-components-features-uses-and-limitations/?ref=blog.datahut.co) and its supporting classes handle browser automation by opening pages, waiting for elements to load, and managing timeouts so the script can smoothly collect vegetable product data without revealing any website links directly in the code. ### Setting Up Basic Configuration Paths for Storing Scraped Data ``` CONFIGURATIONS DB_PATH = "/home/anusha/Desktop/DATAHUT/carrefouruae/Data/carrefour.db" LOG_DIR = "/home/anusha/Desktop/DATAHUT/carrefouruae/Log/data_scraper" JSON_OUTPUT = "/home/anusha/Desktop/DATAHUT/carrefouruae/Data/product_data.json" """ - DB_PATH: Path to the SQLite database where product data will be stored. - JSON_OUTPUT: Path to the JSON file where data will be exported as backup. - LOG_DIR: Directory where log files will be saved. """ ``` The configuration section simply defines three important paths that guide where the scraped vegetable information will be stored, with one pointing to the [SQLite database](https://www.geeksforgeeks.org/sql/introduction-to-sqlite/?ref=blog.datahut.co) for structured records, another leading to the JSON file used as a backup copy, and the last specifying the folder where log files should be saved, creating a clear and organized foundation before the scraper begins collecting data from the Carrefour UAE vegetables section. ### Logging Setup ``` LOGGING os.makedirs(LOG_DIR, exist_ok=True) log_file = os.path.join(LOG_DIR, "product_scraper.log") logging.basicConfig( filename=log_file, filemode='a', level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s' ) ``` The logging setup creates a folder if it does not already exist, chooses a file where all activity will be recorded, and then defines how each log entry should look by specifying the file path, the mode for adding new messages, the level of detail to capture, and a clear format showing the time, message type, and description, making it easier to trace what happens during the vegetable-scraping process. ### Initializing a Simple Selenium Browser to Collect Vegetable Product Data ``` SELENIUM SETUP options = Options() options.headless = False driver = webdriver.Firefox(options=options) driver.set_page_load_timeout(60) """This block is responsible for starting and configuring the web browser that will be controlled by our script""" ``` The [Selenium setup](https://medium.com/@datajournal/web-scraping-with-selenium-955fbaae3421?ref=blog.datahut.co) creates a controlled browser window that the script can operate, with the options setting used to decide whether the browser stays visible, the driver launching Firefox to perform the automated actions, and the timeout ensuring the page has enough time to load, making the scraping process for the Carrefour UAE vegetables section smooth. ### SQLite Table to Store Product Details ``` DATABASE SETUP conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER, url TEXT, name TEXT, price REAL, discount TEXT, scraped_date TEXT ) """) conn.commit() """This block creates and prepares the SQLite database where all scraped Carrefour product details will be stored""" ``` The database setup opens a connection to the chosen storage file, creates a cursor to send instructions, and sets up a table only if it does not already exist, with columns for the product link, name, price, discount, and the time it was scraped, forming a clean and friendly structure for saving item details collected from the vegetables section. ### Safe Extraction Helper to Cleanly Collect Data ``` HELPERS def safe_extract(soup, selector, attr=None, multiple=False, default=None): """Safely extract text or attribute values from an HTML element using BeautifulSoup """ try: if multiple: return [e.get(attr) if attr else e.get_text(strip=True) for e in soup.select(selector)] else: element = soup.select_one(selector) return element.get(attr) if attr else element.get_text(strip=True) except Exception: return default ``` The helper function acts like a gentle, reliable assistant that handles the small but important task of pulling information from each HTML element without breaking the scraping process, and it works by checking whether a single item or multiple items should be gathered, selecting the correct element from the page, and then returning either its text or a specific attribute based on what is needed; if anything unexpected happens—such as the element not existing or the structure being different than usual—the function quietly returns a default value instead of causing an error. ### Handling the Cookie Consent Popup for a Smooth Scraping Flow ``` HELPER TO HANDLE COOKIE POP-UP def handle_cookie_popup(): """Handles the cookie consent pop-up on the Carrefour website""" try: WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler"))).click() logging.info("Accepted cookies.") except Exception: pass ``` The cookie-handling helper takes care of the small but necessary step of clearing the consent popup that appears before the vegetables page can load properly, and it does this by patiently waiting for the accept button to become clickable, pressing it when available, and quietly moving on if the popup is not present; this simple action keeps the scraper from getting stuck at the very beginning and ensures that the rest of the data-collection process can continue smoothly. ### Processing Unfinished Vegetable Links and Extracting Complete Product Details ``` MAIN SCRAPER cursor.execute("SELECT id, url, scraped_date FROM product_urls WHERE processed=0") rows = cursor.fetchall() logging.info(f"Total URLs to scrape: {len(rows)}") """It takes product URLs that have not yet been processed from the `product_urls` table, visits each page, extracts product details, and saves them into the `product_data` table""" all_data = [] for row in rows: product_id, url, scraped_date = row try: """ Step 1: Open the product page and prepare soup object - Load the URL in Selenium - Handle cookie popup if it appears - Wait for a short random delay to mimic human browsing - Convert the loaded page into a BeautifulSoup object """ driver.get(url) handle_cookie_popup() time.sleep(random.uniform(3, 6)) soup = BeautifulSoup(driver.page_source, 'html.parser') ``` The main scraper begins by pulling all vegetable product links from the database that are still marked as unprocessed, creating a clear list of pages that still need attention, and for each of these links, the script carefully opens the product page in the browser, waits for it to load naturally, handles any cookie popups that may appear, and then allows a small random delay to make the browsing pattern look more human; once the page settles, the [HTML](https://www.w3schools.com/html/html%5Fintro.asp?ref=blog.datahut.co) is passed into [BeautifulSoup](https://pypi.org/project/beautifulsoup4/?ref=blog.datahut.co), which turns it into a structured format the script can read, and from there, the scraper gradually extracts details such as the product name, price, image, and other information, storing everything in a separate data table so nothing gets missed or overwritten, making the overall scraping journey from raw links to meaningful product data smooth and organized. ### Extracting the Main Product Details Like Name, Price, and Discount ``` EXTRACT PRIMARY DATA """Step 2: Extract product name, price and discount""" name = safe_extract(soup, "h1 span") price_main = safe_extract(soup, "div.flex.items-baseline div.text-xl") price_dec = safe_extract(soup, "div.flex.items-baseline div.text-sm.ml-2xs") try: price_str = f"{price_main}{price_dec}".replace('..', '.').replace(' ', '').replace('AED', '').replace(',', '.') price = float(price_str) except: price = 0.0 discount = safe_extract(soup, "div.text-md.leading-5.font-bold.ml-2xs.text-c4red-500") ``` The section responsible for extracting the primary details focuses on gathering the most important information a shopper would notice first, and it starts by pulling the product name from the page so the item can be clearly identified; next, it reads the price, which is often split into two small parts on the page, and these pieces are gently combined, cleaned, and converted into a proper number so the value can be stored accurately, with a fallback of zero in case the format is unusual; finally, the scraper checks whether a discount label is displayed and captures it when present, allowing the data to reflect both regular and offer prices, and this careful step-by-step approach helps create a complete and reliable record of each vegetable product. ### Storing Clean Product Details Safely in the Database ``` INSERT INTO DB """ Step 3: Save extracted data into the database """ cursor.execute(""" INSERT INTO product_data (id, url, name, price, discount, scraped_date) VALUES (?, ?, ?, ?, ?, ?) """, ( product_id, url, name, price, discount, scraped_date )) conn.commit() cursor.execute("UPDATE product_urls SET processed=1 WHERE id=?", (product_id,)) conn.commit() ``` Once the scraper finishes collecting the main details of each vegetable product, this section ensures that everything is stored neatly in the database so nothing gets lost, and it does this by inserting the product’s unique ID, page link, name, cleaned price, discount value, and the original scraped date into a structured table designed to hold the final information; after saving the entry, the script immediately updates the status of that specific link in the product URL table to show that it has already been processed, which prevents the same page from being scraped again in future runs, and this simple two-step cycle—first storing the data, then marking it as complete—helps the entire workflow stay organized, and consistent. ### JSON Export with a Simple In-Memory Backup ``` APPEND TO all_data FOR JSON OUTPUT """Step 4 : Save a copy into memory (all_data list) - This allows exporting later into JSON, CSV, or other formats""" all_data.append({ "url": url, "name": name, "price": price, "discount": discount, "scraped_date": scraped_date }) logging.info(f"Scraped product: {product_id}") except Exception as e: logging.error(f"Error scraping URL {url}: {e}") continue ``` After each vegetable product is successfully scraped and stored in the database, this part of the script creates a simple in-memory copy by adding the cleaned details to a list called all\_data, which acts like a temporary collection tray that gathers every item during the run; keeping this extra copy makes it much easier to export all products at once into formats like JSON or CSV later on, without needing to re-query the database or repeat the scraping process. Each entry in the list includes the product link, name, final price, discount value, and the timestamp marking when it was collected, giving the entire dataset a well-structured and easy-to-use format. Along the way, the script logs each product ID to keep a clear record of progress, and if an error occurs for a particular link, the scraper simply notes the issue and moves on, ensuring the full extraction from the Carrefour UAE vegetables section continues smoothly without stopping midway. ### Save To JSON ``` SAVE TO JSON with open(JSON_OUTPUT, "w") as f: """Save all collected product data into a JSON file""" json.dump(all_data, f, indent=4) logging.info(f"Scraping completed. Data saved to {JSON_OUTPUT}") ``` The final step of the scraper focuses on turning all the collected vegetable product details into a neatly formatted JSON file, making the data easy to reuse for analysis, reporting, or further automation. By opening the chosen output file in write mode and passing the complete all\_data list into json.dump, the script creates a structured snapshot of every product scraped in this session, with clear indentation that keeps the file readable. Once the export is complete, a log message confirms the successful save, helping maintain a smooth flow from scraping to storage and giving the entire workflow a polished finish that connects naturally with everything done earlier in the Carrefour UAE vegetables section scraping process. ### Closing the Scraper Safely by Releasing All Used Resources ``` CLEANUP driver.quit() conn.close() """This block ensures that all resources used during the scraping process are properly released after the script finishes execution""" ``` The cleanup step brings the scraping process to a gentle and responsible finish by closing the browser window with driver.quit() and shutting down the database connection with conn.close(), making sure that every tool opened during the run is properly released so the system stays smooth, stable, and ready for the next time the vegetable data from the Carrefour UAE section needs to be collected or reviewed. ## Conclusion Bringing this study to a close, the five-day snapshot of the Carrefour UAE vegetables section shows how even everyday grocery items carry quiet patterns when observed over time. Tracking the same set of vegetables day after day highlighted that prices and discounts are not as static as they appear during a single visit, with small shifts reflecting stock movement, promotional cycles, and short-term demand. Some vegetables remained steady throughout the period, offering a sense of pricing consistency, while others showed brief changes that hinted at ongoing offers or supply adjustments. Beyond the numbers, the process itself demonstrated how time-based scraping turns simple product data into a story, revealing how online shelves evolve rather than staying frozen in one moment. This [time-series approach ](https://www.geeksforgeeks.org/machine-learning/time-series-analysis-and-forecasting/?ref=blog.datahut.co)transforms routine scraping into a meaningful learning exercise, showing that even a short window of data collection can uncover valuable insights and build a strong foundation for deeper analysis of real-world e-commerce behavior. ## Libraries and Versions Name: playwright Version: 1.48.0 Name: playwright-stealth Version: 1.0.6 Name: selenium Version: 4.18.1 Name: beautifulsoup4 Version: 4.13.3 ## AUTHOR I’m Anusha P O, Data Science Intern at [Datahut](https://www.datahut.co/?ref=blog.datahut.co), with a strong focus on building automated data collection workflows that convert raw web content into clean, structured datasets ready for analysis. In this blog, the focus shifts to practical [data extraction](https://www.geeksforgeeks.org/data-analysis/what-is-data-extraction/?ref=blog.datahut.co) from large e-commerce platforms, using the Carrefour UAE website as a real-world example. Modern online retailers display thousands of products across dynamic pages, and understanding how to responsibly collect and organize this information is an essential skill for anyone working with data today. At Datahut, the goal is to design scalable and reliable web data solutions that support informed business decisions. If there is interest in using public web data for market analysis, pricing intelligence, or category insights, feel free to connect through the chat widget and explore how raw online data can be turned into actionable intelligence. FAQ SECTION 1\. What data can you scrape from Carrefour UAE? - Product name and category - Price and discount details - Stock availability - Product URLs and metadata 2\. Which tools are best for scraping Carrefour UAE data? - Playwright for dynamic page loading - Selenium for browser automation - BeautifulSoup for HTML parsing - SQLite and JSON for data storage 3\. Is it legal to scrape Carrefour UAE website data? - Scraping publicly available data is generally allowed - Must follow website terms of service - Avoid overloading servers with frequent requests - Use data responsibly for analysis 4\. Why track vegetable prices over time? - Identify pricing trends and fluctuations - Monitor discount cycles - Analyze demand and supply patterns - Improve competitive pricing strategies 5\. How do you clean scraped ecommerce data? - Remove duplicates and missing values - Standardize price formats - Convert data types (string to numeric) - Structure data for analysis tools ### Cetaphil on Amazon in 2026: Pricing Strategy, Best Sellers & Market Position URL: https://www.blog.datahut.co/post/cetaphil-amazon-analysis-2026/ Last updated: 2026-09-07T09:43:08.000Z Cetaphil on[ Amazon has built one of the most review-dense skincare catalog](https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/)s on the platform - 143 active products, over 707,466 verified customer reviews, and an average rating of 4.15 out of 5\. But not all of those products are winning equally. We analysed every Cetaphil best seller and SKU across the full catalog to understand exactly what's driving customer trust, where the brand's pricing strategy is working, and where a measurable gap is silently holding it back. The brand, owned by [Galderma](https://www.galderma.com/?ref=blog.datahut.co), competes in Amazon's highly [competitive skincare category](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/) on the strength of [clinical credibility rather than trends](https://www.aad.org/public/everyday-care/skin-care-basics/dry/dermatologists-tips-relieve-dry-skin?ref=blog.datahut.co). This report uses structured [Amazon data](https://sell.amazon.com/blog/amazon-seo?ref=blog.datahut.co) to show how Cetaphil earns and in one key segment, loses - customer confidence. ## Cetaphil Sale 2026: Amazon Key Data at a Glance Based on analysis of 143 actively [priced products](https://www.blog.datahut.co/post/price-optimization-solutions/) and 707,466 verified customer review, these are the key findings from Cetaphil Amazon Analysis: Wondering how Cetaphil stacks up against its closest rival? See our full [CeraVe on Amazon 2026 analysis](https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/) — 83 products, 1.6M reviews, and a very different catalog strategy. ## How Does Cetaphil Perform on Amazon? Ratings, Pricing & Review Data ![How Does Cetaphil Perform on Amazon? Ratings, Pricing & Review Data](https://www.blog.datahut.co/content/images/2026/07/img-276.png.webp) Cetaphil's presence on [Amazon spans 143 active products](https://www.blog.datahut.co/post/how-to-do-amazon-market-research/) across 10+ product forms, making Cetaphil Amazon one of the most review-dense skincare catalogs on the platform. This high review density indicates strong repeat-purchase behavior and customer engagement. An average rating of 4.15 out of 5 demonstrates consistent quality across the catalog. The $24.07 average sale price positions Cetaphil in the accessible mid-range, appealing to its core audience of sensitive skin consumers, parents, and those referred by dermatologists. Q. Are there Cetaphil sales in 2026? Yes - based on Amazon data, Cetaphil products are frequently discounted, especially during seasonal sales events. Most discounts range between X–Y%, with top SKUs receiving the deepest cut. ## Cetaphil Best Sellers on Amazon: [Which Products Get the Most Reviews?](https://www.blog.datahut.co/post/how-to-scrape-review-data-from-amazon-and-do-sentiment-analysis/) Lotion: 197,375 total reviews across 39 SKUs. This category focuses on daily moisturizing for dry to sensitive skin. With 39 Lotion SKUs, Cetaphil appears in nearly every moisturizer-related search on Amazon. Cream: 187,472 total reviews across 31 SKUs, averaging over 6,000 reviews per SKU, which is the highest engagement rate among core formats. Cream products primarily serve customers managing eczema-prone or compromised skin who use them as daily therapeutic solutions. Gel: 80,229 reviews across 11 SKUs, representing Cetaphil's offering for lightweight hydration and oil-free cleansing. This format attracts customers who prefer lighter textures before exploring the broader moisturizer range. Q: What is Cetaphil's average rating and price on Amazon? Cetaphil products on Amazon maintain an average rating of 4.15 out of 5, with an average sale price of $24.07, placing it firmly in the accessible mid-range skincare segment ## How Cetaphil Competes for Shelf Space: SKU Distribution Strategy On Amazon, a higher SKU count increases search visibility, impressions, and sales opportunities. Cetaphil has focused on its strongest formats instead of covering every skincare niche. ![How Cetaphil Competes for Shelf Space: SKU Distribution Strategy](https://www.blog.datahut.co/content/images/2026/07/img-277.png.webp) ### SKU Distribution Breakdown - Lotion (39 SKUs) and Cream (31 SKUs) together comprise 48.9% of the active catalog, reflecting a strategic focus on dominating the moisturization segment rather than spreading resources across many formats. - Liquid (17 SKUs) forms a meaningful third tier, capturing customers who prefer pump-format or liquid-based cleansers and treatments. - Gel (11 SKUs) targets the lightweight hydration segment, appealing to a different customer profile than Cetaphil's traditional dry-skin base. - Foam, Bar, Drop, and Wipes (4–5 SKUs each) address specialized needs such as baby care, travel-friendly cleansing, precision serums, and sensitive eye care. While smaller in volume, these formats are important for comprehensive customer lifecycle coverage. - Serum (3 SKUs) and Oil (2 SKUs) indicate Cetaphil's initial expansion into the premium, high-margin skincare segment, extending beyond its traditional mass-market positioning. This depth-over-breadth approach creates visibility advantages that are difficult for new entrants to replicate. ## Cetaphil on Amazon: Which Products See the Deepest Discounts? Discount activity on Amazon signals which products a brand is prioritizing, where customer acquisition costs are high, and where established trust reduces the need for promotions. [Cetaphil's discount patterns are deliberate and targeted.](https://www.bigcommerce.com/articles/ecommerce/pricing-strategy/?ref=blog.datahut.co) ![Cetaphil Amazon Discount Strategy:](https://www.blog.datahut.co/content/images/2026/07/img-278.png.webp) Key Takeaway: Cetaphil employs two parallel discount strategies: significant promotional investment in developing categories (Ointment, Gel) and maintaining perceived value in its core category (Lotion). Data indicates both strategies are effective. Q. How does Cetaphil handle pricing and discounts on Amazon? Cetaphil averages an 11.85% discount across its catalog, with the highest promotions (33%) on Ointments to encourage trial, and lower discounts (9.54%) on Lotions to protect brand value where loyalty is strongest. ## Price Tier Analysis: Where Cetaphil Wins and Where It Struggles The distribution of Cetaphil's 143 active products across price tiers reflects its approach to customer segmentation, from new customers seeking affordable cleansers to loyal customers purchasing specialized barrier repair treatments. ![Cetaphil Price Tier Analysis](https://www.blog.datahut.co/content/images/2026/07/img-279.png.webp) ### Budget Tier (≤$15) - 48 SKUs This tier has an average price of $10.95 and a 4.21 average rating. Body Wash and basic cleansers anchor this segment, serving as entry points for new or budget-conscious customers. ### Mid-Range Tier ($15–$30) - 68 SKUs Sweet Spot Cetaphil performs best in this tier, with 68 SKUs, the highest average rating of 4.38, and a strong concentration of core Lotion and Cream products. Clinical efficacy at an accessible price resonates most with shoppers in this segment. Protecting and expanding this tier is essential. ### Premium Tier ($30+) - 27 SKUs Biggest Vulnerability Products in this tier average $53.34 and have the lowest catalog rating at 3.47\. Customers at this price point have higher expectations, often comparing Cetaphil to specialized clinical and luxury skincare brands, and these expectations are not consistently met. Strategic Implication: The premium rating gap reflects a positioning challenge, not just a product issue. Cetaphil's identity centers on accessibility and dermatological trust rather than luxury. Addressing this gap through product reformulation, clearer premium positioning, or improved customer expectation management represents the most immediate growth opportunity. Q: What is the biggest vulnerability in Cetaphil's Amazon catalog? Products priced above $30 average only a 3.47 star rating despite an average price of $53.34 - a 0.91-point gap below their mid-range products (4.38). This is Cetaphil's most significant measurable weakness and the clearest short-term growth opportunity. ## Cetaphil Best-Sellers vs Full Catalog: Where Do 707,466 Reviews Actually Come From? ![Cetaphil reviews on Amazon](https://www.blog.datahut.co/content/images/2026/07/img-280.png.webp) Analyzing Cetaphil best sellers on Amazon through 707,466 total reviews shows whether the brand's success relies on a few top products or a broader, more resilient foundation. - The single top product holds only 5.54% of all reviews (39,207 out of 707,466) - no one SKU dominates the catalog. - The top 5 products together account for 26.54% of reviews; the top 10 command 42.80%. - The remaining 133+ products share 57.2% of all review volume - a distributed structure that signals genuine catalog breadth. - For competitors, the top 10 Cetaphil products present a significant social proof barrier that requires strong differentiation or sustained promotional investment to overcome. - This distributed review structure also creates opportunities for new product launches, allowing new Cetaphil products to gain traction without directly competing against established hero SKUs. ## Why Do Customers Buy Cetaphil? What Amazon Reviews Really Say Cetaphil Amazon reviews tell us more than satisfaction levels - they reveal exactly why customers buy and stay loyal. Review attribute data tells us why customers buy and why they stay. By analyzing review tags from Amazon's AI-powered summary system across Cetaphil's catalog, we surface the exact attributes driving purchase decisions. ![Cetaphil Amazon reviews](https://www.blog.datahut.co/content/images/2026/07/img-281.png.webp) ![Cetaphil emotional drivers](https://www.blog.datahut.co/content/images/2026/07/img-282.png.webp) Scent Finding: With 5,285 mentions, Scent ranks 4th , but NOT because customers love a fragrance. They are specifically calling out the absence of harsh fragrance as a purchase driver. Fragrance-free is a competitive differentiator, not just a formulation spec. Together, these attributes map the complete customer decision journey: Protection and moisturization bring customers in. Gentleness (signaled through scent-free formulation) keeps them from leaving. Softness and cleanliness deliver sensory confirmation. Value for money seals the repeat purchase. Q: What features do customers value most in Cetaphil products? Skin Protection is the [#1](https://www.blog.datahut.co/blog/hashtags/1) purchase driver with over 7,000 review mentions. Moisturizing efficacy and fragrance-free formulation follow closely. Customers buy Cetaphil specifically for its safe, gentle, dermatologist-trusted formulas , not for luxury or trend appeal. ## Cetaphil Amazon Ratings by Product Form: Ointment 4.80, Wipes 4.73 - Full Breakdown Cetaphil maintains high satisfaction scores across all product forms, with a rating range of only 0.32 points between the lowest and highest-rated formats. This reflects one of the highest levels of quality consistency in the skincare category. ![Cetaphil Amazon Ratings ](https://www.blog.datahut.co/content/images/2026/07/img-283.png.webp) Maintaining all product forms above a 4.48 rating at scale is a significant quality achievement. This confirms that Cetaphil's formulation standards are consistent across the entire portfolio, not only in core moisturizer lines. Q: Which Cetaphil product form has the highest customer engagement on Amazon? Cetaphil Wipes have the highest individual product engagement, averaging 15,159 reviews per listing, despite only 4 SKUs. Masks (11,921 average reviews) and Scrubs (9,101 average reviews) follow. ## Cetaphil on Amazon: Strategy Analysis Report The data demonstrates that Cetaphil has earned its position through consistent performance. Its 707,466 reviews and 4.15 average rating reflect sustained delivery on its core promise, rather than reliance on viral trends or aggressive promotions. ## What Cetaphil Is Doing Right - Depth over breadth: Lotion and Cream dominance creates significant search visibility advantages that are difficult for new entrants to match. investing in growth categories while protecting core loyalty and executing both successfully. - Customer language alignment: The leading review attributes (Skin Protection, Moisturizing, Fragrance-Free) align directly with the brand's positioning, indicating that Cetaphil delivers on its advertised promises. - Quality consistency: Maintaining all product forms at a rating above 4.48 at scale demonstrates strong formulation discipline. ## Where Cetaphil Must Improve - The premium rating gap (3.47 average for products over $30) is the most actionable finding. It highlights a measurable disconnect between customer expectations and product delivery at the premium price point. - Oil (4.75 rating, 2 SKUs) indicates an underserved, high-satisfaction segment with likely unmet demand. - Serum and Oil categories are in early stages. Premium positioning in these segments requires clearer customer expectation management before prices exceed $30. Cetaphil isn't alone in this struggle - CeraVe faces the exact same premium pricing wall. We compared both brands head-to-head in [CeraVe vs. Cetaphil on Amazon: Who Really Wins and Why](https://www.blog.datahut.co/post/cerave-vs-cetaphil-on-amazon/). ## Stop Guessing, Start Winning: Amazon Competitor Intelligence with Datahut This Cetaphil Amazon analysis is based on structured, [reliable marketplace data](https://online.hbs.edu/blog/post/data-driven-decision-making?ref=blog.datahut.co) \- the type of intelligence that distinguishes category leaders. If your brand or clients require: - Near real-time competitor pricing and discount tracking - Brand share-of-reviews and visibility benchmarking against category leaders - Hidden revenue leak detection, assortment blind spots, or emerging customer trend monitoring - Dashboards actionable by marketing, sales, and merchandising teams - Clean, reliable data pipelines for AI models and demand forecasting FAQ SECTION ### 1\. How many products does Cetaphil have on Amazon? Cetaphil Amazon has 143 active products across 10+ product formats. • Includes lotions, creams, cleansers, serums, and oils • Covers both daily care and treatment-focused SKUs • Reflects a broad, diversified skincare catalog • Supports multiple skin types and use cases ### 2\. What is Cetaphil’s best-performing price tier on Amazon? The $15–$30 price tier is Cetaphil’s strongest segment by performance. • Contains 68 SKUs (largest concentration) • Highest average rating: 4.38 / 5 • Balances affordability with perceived quality • Represents the brand’s core value positioning ### 3\. Does Cetaphil rely on hero products or a broad catalog? Cetaphil relies on broad catalog strength rather than a few hero products. • Top product contributes only 5.54% of total reviews • 57.2% of reviews spread across 133+ products • Indicates low dependency on single SKUs • Creates a stable and resilient product portfolio ### 4\. Why is fragrance-free important for Cetaphil customers? Fragrance-free is a key purchase driver for Cetaphil customers with sensitive skin. • “Scent” appears in 5,285+ reviews • Customers prefer absence of fragrance, not added scent • Reduces irritation for sensitive skin users • Acts as a primary filtering factor during purchase ### 5\. What data methodology was used in this Cetaphil Amazon analysis? This analysis is based on structured Amazon data collected and processed by Datahut. • Covers 143 active SKUs and 707,466 reviews • Includes pricing, discounts, ratings, and product types • Uses web scraping pipelines for data extraction • Applies AI tagging for review attribute analysis ### Web Scraping vs. Web Crawling: Which One Do You Need? [2026 Guide] URL: https://www.blog.datahut.co/post/web-scraping-vs-web-crawling-2026/ Last updated: 2026-09-07T09:43:10.000Z You've seen both terms in job posts, tutorials, and tools, and everywhere. You've probably nodded along. The confusion around web scraping vs. web crawling stems from one simple fact - they sound similar but are fundamentally different processes. And mixing them up is costing you time. ## Why the Confusion Exists around Web scraping and Web crawling? If you ask ten developers to explain the difference between web scraping and web crawling, you’ll probably get ten different answers. People often use these terms interchangeably in job posts, documentation, and tutorials. But they actually describe different tasks, each with its own tools, challenges, and uses. Let’s clear up the confusion. ## What Is Web Crawling? What Is [Web Scraping](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/)? Let's dive in deeper to know more about Web Scraping vs. Web Crawling: [Web crawling](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch/) is the systematic browsing of [the web](https://www.w3.org/standards/?ref=blog.datahut.co) to find and index URLs across different sites. The main goal is to navigate and map out what’s out there. This is what Googlebot and other search engine crawlers do: they travel the internet, following links across millions of sites to build the search indexes that power results on Google, Bing, and other search engines. Web scraping is about extracting specific data from chosen web pages, like prices, reviews, emails, or stats. Want to see this in action? [Here's how we scrape product data from e-commerce sites at scale.](https://www.blog.datahut.co/post/web-scraping-to-extract-product-data-from-e-commerce-sites/) The goal is to collect data, not to explore. Scrapers help content aggregation platforms by gathering listings, articles, and product data from many sources, bringing all this information together in one place. Imagine a huge library. A crawler is like a librarian who walks every aisle, noting which books exist and where they are. A scraper is like a researcher who goes to certain shelves and copies down specific passages. One finds and maps everything; the other digs out the details. ## Web Scraping vs. Web Crawling: Side-by-Side Comparison Understanding web scraping vs. web crawling is crucial because choosing the wrong tool wastes time and resources. ## How Web Crawling and Scraping Work Together: The 2-Step Data Pipeline In practice, crawling and scraping often work together as part of a two-step process. You might not know every URL you need at first, especially on a big e-commerce site with thousands of product pages in many categories. - Stage 1: Crawl. Systematically follow links across the site to find all product page URLs. This often includes reading the site's XML sitemap file to speed up discovery instead of following every link by hand. - Stage 2: Scrape. Visit each URL you found and pull out the specific data you need, like name, price, or rating, from each page, one category at a time. Here's the thing: web scraping vs. web crawling isn't really an either-or choice. Most production systems use both. Scaling this two-step pipeline to millions of pages? [Read our guide on scraping Amazon and large e-commerce sites at scale.](https://www.blog.datahut.co/post/how-to-scrape-amazon-and-other-large-e-commerce-websites-at-a-large-scale/) ![How Web Crawling and Scraping Work Together: The 2-Step Data Pipeline](https://www.blog.datahut.co/content/images/2026/07/img-89.jpg.webp) Large crawls almost always surface the same URL multiple times through different navigation paths. Deduplication, filtering out URLs you've already visited or processed, ensures each page is crawled and scraped only once, keeping your pipeline efficient. Think of it this way: you crawl to find the haystack, and you scrape to pull out the needles. ## [Best Web Scraping and Crawling Tools](https://www.datahut.co/?ref=blog.datahut.co) in 2026: Scrapy, Playwright, and Beyond Your choice of library determines whether your web scraping or web crawling pipeline views a static snapshot of a page. This decision ultimately defines whether you are able to interact with an application. ### Requests-Based Frameworks ([Scrapy](https://scrapy.org/?ref=blog.datahut.co), BeautifulSoup) These tools are the speed demons of web data extraction. They make direct HTTP requests to a server and retrieve the raw HTML. Most production-grade frameworks in this category include built in proxy support to enhance reliability. This feature allows you to route requests through different IP addresses from the very start of your scraping process. Best for: Massive crawls of static web pages, sitemap parsing across large sites, or targets that provide data via an API. New to Python scraping? [Our Python web scraping tutorial walks you through Scrapy and BeautifulSoup step by step](https://www.blog.datahut.co/post/python-web-scraping-tutorial/). Workflow: You define seed URLs, and a spider follows hyperlinks across domains to populate a search engine index or a private database. Data is typically exported to CSV, JSON, Excel, or a database, depending on the downstream use case. Limitation: Requests-based frameworks and tools can’t run JavaScript. If a site uses React or Vue to show its content, these tools will only see an empty page. ### Browser Automation ([Playwright](https://playwright.dev/?ref=blog.datahut.co), Puppeteer, Selenium) A headless browser is a full version of Chrome or Firefox that you control with code. It shows web pages just like a regular browser, including complex JavaScript and dynamic content. Best for: Sites with heavy JavaScript, infinite scrolling, or content that requires user interaction to reveal, like clicking a button to load product prices. Trade-off: These tools use a lot of resources. Running 100 [Playwright](https://www.blog.datahut.co/post/how-to-scrape-tablet-data-from-amazon-using-playwright-step-by-step-tutorial/) instances needs much more CPU and RAM than running 100 Scrapy requests. ## Schema Drift: The Silent Scraper Killer Nobody Warns You About Crawlers are usually reliable. As long as there are links (a tags) on a page, they work. Scrapers, however, are more fragile. A site renames a CSS class from .product-price to .price-now. Your scraper silently returns nothing - no error, no alert - just missing data in your dashboard. Scrapers look for specific data using CSS selectors or [XPath](https://www.blog.datahut.co/post/xpath-for-web-scraping-step-by-step-tutorial/) expressions. These are exact instructions, like "find the element with this class name" or "go to this spot in the HTML." If a site changes a CSS class, moves a div, or shifts a button, your selectors stop working and the scraper breaks. This is called schema drift. When a site’s structure changes, your extraction logic can quietly stop working. It’s not a matter of if it will happen, but when. Scraping isn’t a “set it and forget it” tool. Scrapers need regular maintenance, and schema drift is a hidden cost that most tutorials don’t mention. Before starting a scraping project, remember you’ll likely spend time updating CSS selectors and XPath queries later on. Schema drift, IP blocks, JavaScript barriers, web scraping, and crawling can break often. Datahut’s managed scraping service handles these issues and delivers structured, analysis-ready data without the hassle of ongoing maintenance. Spent more time fixing broken selectors than actually using your data? That's schema drift, and it's normal. Datahut handles it for you, so your team gets clean, analysis-ready data without babysitting scrapers. [Contact Datahut today.](https://www.datahut.co/contact?ref=blog.datahut.co) ## Anti-Bot Detection in 2026: IP Blocking, CAPTCHAs, and How to Stay in the Game Modern websites don’t always allow automated access. Understanding these defenses helps you become a more effective and responsible practitioner. IP Blocking: Sites detect repeated requests from the same IP address and ban it. Proxy rotation, automatically cycling requests through a pool of different IP addresses, is the standard countermeasure. Look for scraping frameworks with built-in proxy support to manage this at scale. CAPTCHA: The classic bot-detection gate. Some advanced scrapers use CAPTCHA-solving services, though this raises significant ethical questions. User-Agent Switching: Scripts that mimic a regular Chrome browser by sending a real browser's User-Agent string in request headers. Rate Limiting: Sending too many requests too quickly can raise suspicion. Good scrapers add delays between requests to avoid drawing attention and to prevent overloading the server. ## Web Scraping or Web Crawling? The 60-Second Decision Checklist Think of crawling as using a wide-angle lens for breadth and discovery, and scraping as using a microscope for depth and precision. Which approach do you need right now? Use crawling if you need to find out what exists on a site when you don’t already have the URLs, or if you want to index content across several domains. Use scraping if you already know where the data is and just need to collect it, like product prices, reviews, or category listings exported to Excel or a database. Use Both if: You need to first discover pages across a large site via sitemap parsing or link-following, then extract specific data from each one. Crawl first, scrape second. ## Is Web Scraping Legal in 2026? What You Need to Know Before You Start Most tutorials skip this part, but don’t ignore it. How you scrape is just as important as what you scrape. - Always follow applicable legal protocols. Check out this - Limit your request rate and use proxy rotation carefully. Don’t overload a server with hundreds of requests per second. - Be aware of [GDPR](https://gdpr-info.eu/?ref=blog.datahut.co) and o[ther data privacy laws when collecting personal data](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/), especially if you’re gathering content from several sites. - If you’re unsure, check for an official API first. It’s usually the cleaner, faster, and more ethical option. ## The One-Line Takeaway Crawling answers the question, "what exists?" It indexes the web across different sites, powering things like search results and site audits. Scraping answers, "what does it say?" It pulls out specific data using CSS selectors, XPath, and proxy rotation to get what you need at scale. Use a crawler when you need to explore, a scraper when you know your target, and both when the project is big enough. ### Frequently Asked Questions: Web Scraping vs. Web Crawling 1. What is the difference between web scraping and web crawling? Web scraping extracts specific data from pages. Web crawling discovers and indexes URLs across the web. - Crawling finds pages, and scraping collects data from them. - Search engine bots (like Googlebot) crawl; data pipelines scrape. - Most large projects use both: crawl to discover, scrape to extract - Crawling is broad and fast; scraping is targeted and structured. 1. Can AI do web scraping automatically? Yes, AI-powered scrapers can extract data without hard-coded rules. - LLM-based tools understand page structure in natural language. - They adapt when a site's layout changes, no manual fixes needed. - Human oversight is still needed for legal compliance and data quality. 1. Is web scraping legal in 2026? Web scraping is legal for publicly available data, but has important limits. - Scraping personal data or bypassing logins can create legal risk. - The HiQ vs. LinkedIn ruling (2022) protects the scraping of public data in the US. - GDPR compliance is mandatory when scraping data from EU users. 4\. Do I need both web scraping and web crawling for my project? It depends on how many pages you need to discover versus extract from. - Known pages only → scraping alone is enough. - Thousands of URLs to discover → you need crawling + scraping together. - E-commerce monitoring, price tracking, and SEO audits typically need both. - Tools like Scrapy and Playwright support both in a single pipeline. 1. What tools are best for web scraping and crawling in 2026? The best tool depends on your use case, technical skill, and scale. 1. Scrapy Python, open-source, best for large-scale crawls 2. Playwright/Puppeteer handles JavaScript-heavy, dynamic pages. 3. BeautifulSoup is lightweight and beginner-friendly for small scrapes. 4. How does AI improve web data extraction compared to traditional scraping? AI extractors understand page content semantically — not just structurally. - Traditional scrapers break when CSS selectors or HTML layouts change. - AI models identify fields like 'price' or 'product name' by meaning, not markup. - Setup is faster — no need to write XPath or CSS rules per site. - More resilient across site redesigns and A/B layout tests. 1. What is a headless browser, and when do I need one for scraping? A headless browser loads full web pages (including JavaScript) without a visible interface. - Required for scraping JavaScript-rendered pages (React, Vue, Angular SPAs) - Unnecessary for static HTML pages — adds overhead without benefit. - Playwright and Puppeteer are the most popular headless browser tools. - Slower than plain HTTP scraping — use only when JS rendering is needed ### How to Steal the Product Copy Formula of Winning Brands: A Simple 4-Step Data-Driven Framework URL: https://www.blog.datahut.co/post/reverse-engineer-product-copy-formula/ Last updated: 2026-07-28T10:18:59.000Z Do you know what actually differentiates great product copy from the good? Great product copy is rarely the result of creative inspiration. It's the outcome of a data-driven process. Good copy is often just creative writing. Bad copy is what happens when someone simply asks ChatGPT to generate it. ## Product copy formula: Stop treating product copy like an art project In this blog, we break down the exact 4-step process you can use to write great product content by reverse-engineering the content generation process of category leaders. Let's say you want to launch a new product in a category where companies like Olay and Cetaphil are dominating. What is the best way to write product content? The brands that dominate their category didn't get there by chance. Their listings are the result of: - Continuous [A/B testing](https://cxl.com/blog/ab-testing-guide/?ref=blog.datahut.co) - [Consumer psychology audits](https://www.nngroup.com/articles/how-users-read-on-the-web/?ref=blog.datahut.co) - Marketplace algorithm optimisation - Thousands of iterations across titles, bullet points, and imagery When a brand like Olay sells millions of products across marketplaces, every part of their product pages has been refined over time: titles, bullet points, images, ingredient claims, and product positioning were all battle-tested. When you study an Olay listing, you aren't just looking at a description — you are looking at a compressed history of optimisation. Every ingredient claim and benefit earned its spot through testing. By reverse-engineering these category leaders, you can bypass years of trial and error and start with a proven formula. ## The 4-Step Framework to write great product copy that sells The most important thing is to understand the patterns that category champions use. This gives you a clear picture of what works and what doesn't. The best way to find those patterns is to analyze the data of the leaders. That's exactly what this framework helps you do. Let's unlock the best way to find your best product copy formula. ### Step 1 - Data Collection The first step is to collect product listings from the category you want to compete in. For this demonstration, we are using 44 Olay product listings from the [Body Wash category on Amazon](https://www.amazon.com/stores/Olay/page/AD2BF519-62C5-4C66-9EDE-763EC687B6C0?ref=blog.datahut.co). The data was extracted using the [Datahut web scraping platform](https://www.datahut.co/solutions/ecommerce-web-scraping?ref=blog.datahut.co). We are demonstrating how to write the product title, but you can follow the same process for bullet points, descriptions, and more. To make this easier to understand, imagine you are launching a new body wash product on [Amazon](https://ecombrainly.com/amazon-conversion-rate/?ref=blog.datahut.co) and want to learn how leading brands structure their product titles. Here is an example of a typical title from the dataset: Olay Body Wash for Women, Fresh Radiance, 24/7 Skin-Loving Freshness, Visibly Radiant, Plant Based Cleansers, Vitamin B3 & Antioxidant Blend, For All Skin Types, Strawberry & Mint Scent, 29 fl oz To an amateur, this looks like a long, cluttered sentence. To a data scientist, it is a highly structured sequence of high-value variables: \[Brand\] + \[Product Type\] + \[Target User\] + \[Benefit Claims\] + \[Key Ingredients\] + \[Skin Type\] + \[Fragrance\] + \[Size\] Once you collect titles across a category, you can start identifying patterns that successful brands consistently use. You can collect this data yourself using web scraping, or access ready-to-use datasets through [Datahut](https://www.datahut.co/contact?ref=blog.datahut.co). ### Step 2: From Messy Text to Analytics Ready Data When you scrape product listings, you end up with messy, unstructured text. To a human, the structure is obvious but to a computer, it's just a string of words. Annotation converts this into structured product intelligence. For example, this title: Looks like one long sentence. But an annotation tool turns it into structured information like this: This process is called data annotation. Annotated data looks like this Each title is broken down into attributes such as Brand, Category, Variant, Target User, Key Attribute, Benefit, Key Ingredient, Size/Weight, Scent, Skin Type, Technology, Bundle, and Free-of claims. By labeling titles across these dimensions, we convert unstructured product text into structured data. This makes it possible to analyse patterns in how leading brands construct their listings what benefits they emphasise, which ingredients they highlight, and how they position variants and claims. This structured annotation allows us to reverse-engineer the formula behind effective product titles and identify the signals that marketplaces and shoppers respond to. You can use any annotation tool available. We also built a lightweight browser-based annotation tool that works with any LLM API to speed up the process. You can download it and run it directly in your browser. You can get it here: [Get the annotation tool](https://tally.so/r/0QLrM9?ref=blog.datahut.co) You can also see the demo video of annotation here : [Demo Video](https://www.youtube.com/watch?v=8WbbICYAMew&ref=blog.datahut.co) ![Data annotation tool for better product copy](https://www.blog.datahut.co/content/images/2026/07/img-427.png.webp) After annotation the data looks like this ``` [ { "text": "Olay Super Serum Body Wash for Extra Dry Skin, 24hr Long Lasting Hydration, 5+ Ingredient Complex for Bright Even Firm Luminous Skin, 18.5 fl oz", "annotations": [ { "text": "Olay", "label": "Brand" }, { "text": "Body Wash", "label": "Category" }, { "text": "Extra Dry Skin", "label": "Skin Type" }, { "text": "Hydration", "label": "Benefit" }, { "text": "Bright", "label": "Benefit" }, { "text": "Even", "label": "Benefit" }, { "text": "Firm", "label": "Benefit" }, { "text": "Luminous Skin", "label": "Benefit" }, { "text": "Super Serum", "label": "Variant" }, { "text": "18.5 fl oz", "label": "Variant" }, { "text": "Serum", "label": "Key Ingredient" } ] }, ``` ### Step 3: Finding the Winning Formula After you finish annotating all the titles, export this entire data as a json file. You can now convert this to a csv file with the labels as headers. After annotating all 44 listings, the patterns become undeniable. Olay titles aren't just descriptions; they are modular stacks of information. ![The anatomy of Olay product titles and the patterns you can clearly see.](https://www.blog.datahut.co/content/images/2026/07/img-428.png.webp) The annotated Olay product titles show a very consistent structure. - Brand, Category, Benefit, and Size/Weight appear in all 44 listings, making them the core components of every title. - Skin Type, Key Ingredient, and Target User appear in most listings (42), indicating that Olay frequently highlights who the product is for and what key ingredients drive the benefit. - Scent appears in about three-quarters of the titles (32) as an additional descriptive element, while attributes like Overall, the data suggests that Olay titles prioritize clear product identification, benefit communication, and ingredient signalling. Additional attributes used selectively to enhance differentiation. From the position analysis of the annotated fields, the typical title structure looks like this: Brand + Category + Target User + Variant + Benefit + Key Ingredient + Skin Type + Scent + Technology + Free Of + Size/Weight + Bundle ### Step 4: Ingredient Signalling (The SEO Secret) Analysis of the annotated data reveals which "Power Words" Olay uses to trigger customer trust and search relevance. In this category, [Vitamin B3 / Niacinamide](https://health.clevelandclinic.org/niacinamide?ref=blog.datahut.co) was mentioned 41 times across 44 titles. This tells a new competitor two things: 1. Market Expectation: If you don't mention Niacinamide or a similar "Serum Complex," you are invisible to the category's top shoppers. 2. SEO Priority: These ingredients are likely the highest-volume search terms for the category. #### The Power of High-Frequency Signaling By counting ingredient mentions in annotated titles, we can see the signals brands believe resonate most with shoppers. Beyond the dominance of Vitamin B3, the data shows a clear hierarchy of secondary ingredients: - Serum Complex (18 mentions): This suggests Olay is moving the category toward a "skincare-as-bodycare" trend. - Plant Based Cleansers (16 mentions): This indicates a strategic pivot toward "clean beauty" claims to capture health-conscious segments. - Hyaluronic Acid (15 mentions): Highlighting a well-known "hero ingredient" provides an instant mental shortcut for "hydration" to the consumer. - Antioxidant Blend (11 mentions) and Vitamin C (8 mentions): These are used regularly to reinforce specific skin-brightening or anti-aging benefits. ![Product Copy Formula of Winning Brands](https://www.blog.datahut.co/content/images/2026/07/img-429.png.webp) ### Why Ingredient Data Matters This type of analysis allows you to understand ingredient trends within the category before you even write your first draft. Instead of choosing ingredients based on what sounds "nice," you choose them based on proven search volume and consumer trust signals established by the market leaders. When you align your product's key ingredients with these "Top Ingredients in Titles," you aren't just describing a product—you are aligning your listing with the established mental models of your target audience. ## Is it just for the titles ? Not at all. The same reverse-engineering process applies to every element of a product listing — and each one is worth studying carefully. #### Bullet Points Bullet points are where brands do their heaviest persuasion work. By annotating bullet points across category leaders, you can identify which benefit claims appear most frequently, how brands sequence their messaging (does hydration come before brightening, or after?), and what proof points — dermatologist-tested, clinically proven, allergy-tested — the top sellers consistently include. These aren't random choices. They reflect what shoppers actually respond to. #### Product Descriptions Descriptions give brands more room to tell a story, but the best ones still follow a recognizable structure. Analyzing descriptions across a category reveals how leaders handle objections, which emotional triggers they lean on, and how they transition from ingredient claims to lifestyle outcomes. This is where you find the deeper narrative formula that drives conversions. #### Image Alt Text and Backend Keywords The same ingredient and benefit signals that dominate titles often show up in backend search terms and image metadata. Studying what category leaders emphasize visually — and how they describe those images — gives you a fuller picture of their SEO strategy beyond just the visible copy. #### A+ Content and Brand Stories For brands that invest in enhanced content, the patterns are even more revealing. Layout choices, headline structures, comparison charts, and lifestyle imagery all follow repeatable formulas that can be decoded and adapted. Titles are just the easiest place to start because the patterns are most visible there. Once you've built the habit of reverse engineering category leaders at the title level, applying the same logic to the rest of the listing becomes natural. The compounding effect on your content quality will be significant. ## One important caveat Reverse-engineering category leaders gives you a strong starting point, not a final answer. What worked for Olay an established brand with years of trust built in and may not land identically for a new entrant with no brand recognition yet. Treat this analysis as your first draft hypothesis. Use it to build listings that are informed by evidence, then run your own A/B tests to validate what actually converts for your specific product and audience. The framework removes the blank page problem; your own testing is what refines the result. ## Closing Thoughts When you look at a product listing on Amazon, it might appear to be just a piece of marketing copy. But behind every high performing listing is a long history of experimentation, optimisation, and data-driven decisions. The brands dominating a category rarely guess their way to the perfect title. Over time, they test different ingredients, benefits, claims, and positioning until they discover what resonates with both marketplace algorithms and customers. By collecting category data, annotating it, and analysing the patterns, you can uncover the same signals that experienced brands and consulting teams rely on. Instead of starting with a blank page, you start with evidence of what already works in the market. In other words, writing product content is not about creativity alone. It’s about understanding the structure behind successful listings and using data to guide your decisions. And once you approach product content this way, every category becomes something you can study, decode, and systematically improve. ## How We Can Help You Write Better Product Copy Writing great product copy starts with great data — and that's exactly where Datahut comes in. #### 1\. Get the Data You Need, Fast Manually collecting product listings from Amazon, Walmart, or any other marketplace is slow, inconsistent, and prone to gaps. Datahut's web scraping platform extracts clean, structured product data at scale — titles, bullet points, descriptions, ingredient claims, pricing, and more — so you can start your analysis immediately instead of spending weeks gathering raw inputs. #### 2\. Learn From What's Already Working We've worked across dozens of product categories and marketplaces, and we've seen firsthand what separates listings that convert from those that don't. Our team can consult with you on the best practices we've observed across categories — from title structure and ingredient signalling to benefit hierarchy and SEO keyword positioning. You get the benefit of patterns we've identified across thousands of listings, without having to build that knowledge from scratch. If you're launching a new product or optimising an existing catalog, we can help you go from a blank page to a data-backed content strategy — faster than you'd expect. ## Frequently Asked Questions ### Frequently Asked Questions #### 1\. What is the best way to write product titles for Amazon? The most reliable approach is to reverse-engineer titles from the top-selling products already in your category rather than writing from scratch. Collect 20–50 titles from category leaders, break each one down into its components (brand, product type, target user, key benefit, ingredients, skin type, scent, size), and look for the patterns that appear most consistently. Whatever shows up in 80%+ of titles is effectively a category requirement . Your listing needs it to be considered relevant by both the algorithm and the shopper. What appears in 50–70% of titles is where differentiation happens. Start with the required elements, then use the variable ones to position your specific product. Amazon's character limit (typically 200 characters) means you'll need to prioritize, so build your title around the highest-frequency signals first. #### 2\. How do category leaders like Olay write their product listings? The honest answer is that no one at Olay sits down and writes a listing the way a copywriter would. At scale, listings are the output of a testing infrastructure . Product teams run A/B experiments on titles, change one variable at a time (swap an ingredient claim, test a different benefit phrase, add or remove a scent descriptor), and measure the impact on click-through rate and conversion. What you see on the page today is the version that survived that process. For a new entrant without that testing history, reverse-engineering gives you a shortcut: you're essentially inheriting the conclusions of experiments you didn't have to run yourself. The important thing to understand is that you're reading the end state, not the starting point Olay's current listings look nothing like their first drafts. #### 3\. What is data annotation in the context of product content? Annotation is the process of turning a product title which is just a string of text into a structured record where every component has a label. So instead of "Olay Body Wash for Women, Vitamin B3, 29 fl oz" as a single sentence, you end up with a row of data: Brand = Olay, Category = Body Wash, Target User = Women, Key Ingredient = Vitamin B3, Size = 29 fl oz. Once you have 40 or 50 titles annotated this way, you can run simple analysis — what percentage of titles include a skin type claim? Which ingredients appear most often, and in what position? Annotation is what converts a pile of scraped text into something you can actually learn from. It doesn't require specialist tools; a spreadsheet works fine for small datasets, though LLM-assisted annotation tools speed things up significantly at scale. #### 4\. How do I find the right ingredients to include in my product title? If i give you an example from the data set we just analysed, start by collecting titles from the top 20–30 products in your category. List every ingredient or ingredient-adjacent claim that appears (Niacinamide, Hyaluronic Acid, Peptide, Plant-Based Cleansers, etc.). Count how many titles each one appears in. Ingredients that show up in more than 70% of titles are effectively category table stakes — shoppers in that category are already searching for them, and excluding them signals that your product doesn't belong. For ingredients that appear in 30–50% of titles, you have a choice: including them helps you capture a specific search segment, but only if your product actually contains them. Never include an ingredient in a title purely for SEO if it's not meaningfully present in the formula — this generates returns, negative reviews, and potential compliance issues on Amazon. Use the frequency data to prioritize, but let your actual product formulation make the final call. #### 5\. How many competitor listings do I need to analyse for this to be useful? For most product categories on Amazon, 30–50 listings from the top sellers gives you enough signal to identify reliable patterns. Below 20, you risk mistaking one brand's idiosyncratic style for a category norm. Above 100, you hit diminishing returns unless the category is unusually fragmented with many distinct sub-segments. The key is to focus on the right listings: pull from the top 20–30 organic search results for your primary category keyword, not just one brand. If a single brand like Olay dominates the category, include their full catalog but balance it with 10–15 listings from other top sellers so you're reading category patterns rather than one company's house style. Rerun the analysis every 2–3 months ingredient trends and title conventions shift, and what worked two years ago may no longer reflect how the algorithm or the shopper has evolved. ### Why Your Competitors Know More About the Market Than You Do: Competitive Market Intelligence URL: https://www.blog.datahut.co/post/competitive-market-intelligence/ Last updated: 2026-09-07T09:43:13.000Z That gap- three weeks of lost pricing advantage, missed sales opportunities, and reactive efforts- is not a speed issue. It is an intelligence issue. And the uncomfortable truth? While you were reacting, [your competitor was already analyzing the results of that price](https://www.blog.datahut.co/post/dynamic-pricing/) change, studying your response, and planning the next move. This is today’s competitive landscape. Some companies have clear visibility, while most lack critical insight. ![Why Your Competitors Know More About the Market Than You Do](https://www.blog.datahut.co/content/images/2026/07/img-390.png.webp) Sources: IDC / Seagate Data Age Report, Forbes Technology Council (Dark Data analysis), McKinsey Digital Research, Enterprise Analytics & AI Adoption Surveys ## The Intelligence Gap Is Wider Than You Think ### Competitive Market Intelligence Most business leaders believe they remain competitive through superior products, faster teams, or more effective marketing. While these factors are important, they depend on a more fundamental element: understanding current market dynamics. Companies that consistently lead their categories are not necessarily working harder; they are leveraging more accurate and timely information. They identify pricing shifts early, recognize demand changes before they become trends, and address [competitor weaknesses before customers switch.](https://www.blog.datahut.co/post/boost-your-online-retail-marketing-roi-through-competitive-intelligence/) The data gap between these companies has widened significantly in recent years, as tools for gathering external market intelligence have advanced, and not all organizations are utilizing them. ![Competitive data edge: internal vs external](https://www.blog.datahut.co/content/images/2026/07/img-391.png.webp) ## What Your Competitors Are Actually Tracking (That You’re Not) [Modern competitive intelligence](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-dos-and-donts-of-dynamic-pricing-in-retail?ref=blog.datahut.co) does not involve covert activity. All relevant data is publicly available, including product listings, prices, job boards, customer reviews, and social signals. The key difference is whether your organization has a system to collect, process, and act on this information. Here are the four blind spots that consistently separate companies that lead from companies that follow: ## Blind Spot [#1](https://www.blog.datahut.co/blog/hashtags/1): Real-Time Pricing Shifts Amazon adjusts prices approximately 2.5 million times per day.[ Major retailers reprice their catalogs weekly](https://www.blog.datahut.co/post/category-price-index/). Even B2B software companies quietly change their pricing pages without announcement. If you check competitor prices quarterly using spreadsheets, you are not tracking pricing; you are recording historical data. By the time you respond, the market has already changed. Companies with[ pricing intelligence](https://www.bcg.com/publications/2024/overcoming-retail-complexity-with-ai-powered-pricing?ref=blog.datahut.co) systems receive automated alerts when a competitor reduces prices by more than 5%, launches a promotional bundle, or introduces a new tier, before customers question their own pricing. ## Blind Spot [#2](https://www.blog.datahut.co/blog/hashtags/2): Product Assortment Changes Competitors quietly discontinue products that were cannibalizing their margins. They soft-launch new SKUs to test demand. They repackage existing products with new bundling strategies. Every one of these moves is a signal. A discontinuity tells you there was a problem. A new bundle tells you where they think demand is going. A quiet launch tells you what category they are moving into next. Without systematic monitoring of competitor product catalogs, these signals often arrive too late, typically when customers mention them during conversations. ## Blind Spot [#3](https://www.blog.datahut.co/blog/hashtags/3): Customer Sentiment Across the Category Your customers review your products, and your competitors’ customers review theirs. This data- millions of public reviews on Amazon, Google, Trustpilot, Reddit, and app stores- provides insight into what is working, what is not, and what customers desire. Leading brands analyze this data to identify unresolved product complaints from competitors. These gaps become opportunities for new features. By addressing them, you can gain market share without increasing marketing spend. ## Blind Spot [#4](https://www.blog.datahut.co/blog/hashtags/4): Strategic Signals Hidden in Plain Sight Job postings are especially underutilized as intelligence sources. If a competitor is hiring multiple data scientists and a Head of AI, they are likely preparing for a product change. Aggressive hiring in a new region often signals market entry. This is not speculation; it is pattern recognition based on public data. ## The Kodak Problem: Why Smart Companies Still Get Blindsided Even experienced executives may find this concerning: most companies that were blindsided were not complacent. They were monitoring the market, but not focusing on the right indicators. ### Kodak Kodak tracked every move Fujifilm made in the film market for decades. It was diligent and well-resourced competitive intelligence work. Meanwhile, digital photography was emerging from the electronics industry - a category Kodak did not think of as competition. By the time the threat was obvious, the market had already shifted. ### Blockbuster Blockbuster closely monitored Hollywood Video and Movie Gallery. Its competitive intelligence on physical rental competitors was thorough and current. Meanwhile, Netflix developed a different business model with no late fees, no storefronts, and unlimited selection. Blockbuster focused on the wrong competitors. ### BlackBerry BlackBerry monitored all other smartphone manufacturers and understood the hardware competition thoroughly. Apple and Google were developing platforms rather than just better phones. By the time BlackBerry recognized the true competition, the market had shifted. ## 5 Questions Your Competitors Can Already Answer And You Probably Can’t Consider this benchmark: data-driven companies in your category can answer these questions within minutes. Assess how long it would take your team to respond to each. ![data-driven companies queries](https://www.blog.datahut.co/content/images/2026/07/img-392.png.webp) If your honest answer to most of these is ‘days’ or ‘we do not track that,’ you have identified an intelligence gap. The good news is that this data is publicly available. The key is whether you have a system to collect and use it. ## Real-World Examples: What ‘Knowing More’ Actually Looks Like ## The E-Commerce Brand That Won Black Friday A mid-sized consumer electronics brand began tracking competitor pricing across five product categories - not just their top three rivals, but the entire category of 40+ sellers. Six weeks before Black Friday, their monitoring system flagged that a major competitor was quietly testing price reductions on headphones. The brand adjusted its pricing and margin strategy three weeks before the sale period. They entered Black Friday prepared, rather than reacting in real time. Their headphone category revenue grew 31% year over year, compared to their competitor’s 4%. ## The Skincare Brand That Found a Product Brief in 50,000 Reviews A personal care brand sought to expand into a new product segment but was unsure where to begin. The team analyzed over 50,000 public customer reviews across competitor products in the target category, focusing not only on low ratings but also on the specific language used in 3-star reviews. They identified a consistent pattern: customers appreciated the performance of leading products but mentioned issues with scent in 23% of reviews. No competitor had addressed this concern. The brand launched an unscented version, which became its fastest-growing product within six months. ## The B2B SaaS Company That Saw a Competitor’s Pivot Coming A B2B analytics company observed its main competitor posting seven consecutive job listings for enterprise integration engineers and a VP of Partnerships over 90 days. The team identified this as a likely signal of a platform expansion into enterprise workflows, posing a direct threat to their strongest customer segment. They accelerated their integration roadmap by two quarters. By the time the competitor announced its new enterprise tier, the company had already strengthened relationships with 80% of its at-risk accounts. Net revenue retention that year was 118%. ## The Data You’re Missing and Where It Lives: How to Start Building Your Market Intelligence System You do not need to implement everything at once. The highest-ROI starting point is typically [pricing intelligence](https://www.bain.com/how-we-help/retailers-are-you-getting-the-full-value-of-your-dynamic-pricing-strategy?ref=blog.datahut.co), as it is directly linked to revenue, highly actionable, and easily automated. It provides a practical roadmap to go from reactive to proactive in 90 days: ## Month 1: Define What You Need to Know 1. Identify your top 10 competitors, including fast-growing smaller players as well as the most obvious ones. 2. Define the five questions from Section 4 that are most relevant to your business. 3. List the data types from Section 6 that would address those questions. 4. Audit what your team currently tracks and how often it is monitored. ## Month 2: Set Up Monitoring 5\. Begin with price tracking on competitor product pages. This can be accomplished using basic web data tools. 6\. Set up review monitoring across your category, including both your products and those of competitors. 7\. Create job posting alerts for your top five competitors on LinkedIn. 8\. Subscribe to competitor press release feeds and establish news alerts. ## Month 3: Build the Response Loop 9\. Schedule a weekly 30-minute competitive review meeting with key stakeholders. 10\. Establish response thresholds, such as reviewing within 48 hours if a competitor drops prices by more than 5%. 11\. Document your first three insights and the decisions they influenced. 12\. Scale effective practices and discontinue those that did not generate actionable insights. ## Get Your Free Competitive Intelligence Checklist The checklist covers all 10 data points broken down by business type - e-commerce, B2B SaaS, retail, and D2C. Download it, share it with your team, and use it as the starting point for your first competitive intelligence audit. Or if you want us to do it for you, [Datahut](https://www.datahut.co/contact?ref=blog.datahut.co) collects, cleans, and delivers external market data to businesses across every major category. Talk to our team and get a free sample dataset from your industry. ![10 data points broken down by business type](https://www.blog.datahut.co/content/images/2026/07/img-393.png.webp) [https://tally.so/r/81kWLO](https://tally.so/r/81kWLO?ref=blog.datahut.co) ## FAQ SECTION How do companies gather competitive intelligence? Most leading companies combine three approaches: internal data analysis (their own sales, CRM, customer feedback), primary research (customer interviews, surveys), and external data collection (automated monitoring of competitor websites, review platforms, job boards, and news). The third category has seen the most investment in recent years because it scales without proportional headcount increases. Is competitive intelligence legal? Yes — when it involves publicly available data. Monitoring competitor pricing pages, public product listings, open job postings, customer reviews, and press releases is entirely legal and widely practiced. The legal line is around deceptive methods (posing as a customer to extract information) or accessing private, password-protected systems. Ethical competitive intelligence stays within the bounds of what companies have made publicly available. What is the difference between competitive intelligence and market research? Market research typically focuses on understanding customers — their preferences, behaviors, and needs - often through surveys or focus groups. Competitive intelligence focuses on understanding competitors and market dynamics, their pricing, products, strategies, and moves. Effective market intelligence combines both and increasingly draws on external data to answer questions faster and on a larger scale than traditional research methods allow. How often should you track competitor [pricing](https://www.datahut.co/pricing?ref=blog.datahut.co)? It depends on your category's pricing velocity. Fast-moving e-commerce categories (electronics, fashion, consumer goods) can shift daily — weekly monitoring is the minimum and real-time alerts are ideal. Slower-moving B2B software categories typically reprice quarterly. A good rule: track at the same frequency that a mispricing could materially hurt your business. What data does Datahut collect for competitive intelligence? Datahut collects and delivers clean, structured external data including competitor pricing, product catalog changes, customer reviews, market signal data, job postings, and more — from any public website, at any scale. Data is delivered in ready-to-use formats (CSV, JSON, API) on custom schedules. No engineering team required on your end. About Datahut [Datahut](https://www.datahut.co/?ref=blog.datahut.co) is a web scraping and data services company that helps [businesses](https://www.datahut.co/solutions/ecommerce?ref=blog.datahut.co) collect, clean, and activate external market data. We work with e-commerce brands, retailers, financial services firms, and B2B companies to build competitive intelligence systems that don't require a full engineering team. Learn more at datahut.co · Read more on [blog.datahut.co](http://blog.datahut.co/?ref=blog.datahut.co) ### How to Scrape Tablet Data from Amazon Using Playwright (Step-by-Step Tutorial) URL: https://www.blog.datahut.co/post/how-to-scrape-tablet-data-from-amazon-using-playwright-step-by-step-tutorial/ Last updated: 2026-09-07T09:43:15.000Z [Amazon India’s](https://www.amazon.in/?ref=blog.datahut.co) tablets section becomes especially interesting during the [Great Indian Festival](https://www.42signals.com/blog/how-amazons-great-indian-festival-sale-became-indias-shopping-phenomenon/?ref=blog.datahut.co), when prices fluctuate rapidly, rankings change by the hour, and new offers appear across thousands of product listings. This dataset matters because it captures how real tablet products are presented, priced, and promoted during one of the busiest sale periods of the year, offering a clear window into pricing trends, brand competition, and product visibility on a large [e-commerce](https://www.shopify.com/blog/what-is-ecommerce?ref=blog.datahut.co) platform. In this blog, [Playwright](https://playwright.dev/python/docs/api/class-playwright?ref=blog.datahut.co) is used as the core tool because it behaves like a real browser, allowing dynamic pages to load fully before data is collected, even when content changes frequently during a sale. Over a carefully managed three-day scraping process, product URLs and detailed tablet information such as names, brands, and sale prices are extracted in a stable and controlled manner. ## Scrape Tablet Data from Amazon Using Playwright Let us analyze how to Scrape Tablet Data from Amazon using Playwright step by step: ## Stage 1: URL Scraping from the Amazon Tablets Section Before detailed product information can be collected, the first and most time-consuming step is gathering the correct product page URLs, and this stage focuses on how tablet links were carefully [scraped from Amazon’s](https://www.blog.datahut.co/post/is-it-legal-to-scrape-amazon-unethical-uses-of-amazon-web-scraping/) tablets section during the high-traffic Great Indian Festival period. The tablets listing page acts as the main entry point, where hundreds of products appear across multiple pages and continuously change due to offers, rankings, and availability, much like shelves being rearranged in a busy store every few hours. The [scraping process](https://www.blog.datahut.co/post/tutorial-how-to-scrape-amazon-data-using-python-scrapy/) was spread across three days, allowing the system to handle rate limits, page refreshes, and temporary blocks more safely while ensuring no valid tablet listings were missed. The script opens the category page, waits for products to load fully, and then scans each visible product card to extract only the product URLs, carefully skipping accessories or unrelated items. As [pagination](https://www.smashingmagazine.com/2016/03/pagination-infinite-scrolling-load-more-buttons/?ref=blog.datahut.co) moves the listing forward, newly discovered links are stored in a database with progress tracking, so already-collected URLs are not repeated in later runs in a day. In next day urls are collected in a separate table. ## Stage 2: Scraping Data by Visiting Each Amazon Tablet Product Page After the tablet product URLs were safely collected, the next step focused on visiting each link individually and carefully extracting the details displayed on those pages, turning simple URLs into meaningful product data. This stage can be compared to opening every product box after listing them, taking time to read the label, brand, and price instead of judging from the shelf alone. During the Great Indian Festival, when prices and availability changed frequently, this process was intentionally spread across three days to reduce pressure on the website and ensure stable, accurate data collection. Each stored URL was opened one by one, the page was given enough time to load fully, and the visible content was converted into a readable structure so important details such as the tablet name, brand, sale price, and scrape time could be captured without confusion. Once the [information was extracted](https://www.blog.datahut.co/post/how-to-scrape-amazon-and-other-large-e-commerce-websites-at-a-large-scale/), it was saved into a structured database and marked as completed, which prevented the same page from being revisited in later runs and allowed the [scraping process](https://www.blog.datahut.co/post/is-it-legal-to-scrape-amazon-unethical-uses-of-amazon-web-scraping/) to pause and resume safely if needed. ## Stage 3: Cleaning and Preparing the Dataset After scraping raw product data from the Amazon Tablet product pages, the next step was cleaning and preparing the dataset so it could be analyzed reliably. Using [OpenRefine](https://openrefine.org/?ref=blog.datahut.co), the [scraped](https://www.blog.datahut.co/post/python-web-scraping-tutorial/) file was loaded into a spreadsheet-style view, making inconsistencies easy to detect. Prices were in Indian Rupees (₹) and removing currency symbols, commas, and formatting issues make it . During this stage, non-relevant products such as tablet medicines, which appeared alongside tablets on Amazon, were identified and removed to keep the dataset category-specific. Duplicate product URLs generated during scraping were also eliminated. By the end of this process, the dataset was clean, consistent, and fully ready for accurate analysis of [Amazon](https://www.blog.datahut.co/post/the-secret-weapon-of-successful-amazon-sellers/) tablet products. ## Essential Python Libraries Powering the Amazon Tablet Scraper At first glance, a web scraper may appear to be just a short script that pulls data from a webpage, but behind that simplicity lies a group of carefully chosen Python libraries working together to keep the process stable and beginner friendly. In this project, asyncio manages the flow of tasks so the scraper can wait patiently while Amazon tablet pages load without freezing the entire program, which is especially useful when handling many product URLs. Playwright acts as the real browser, opening Amazon’s tablets section, loading [JavaScript-driven](https://www.firecrawl.dev/glossary/web-scraping-apis/what-is-javascript-rendering-web-scraping?ref=blog.datahut.co) content, and capturing the full page just as a human visitor would during the Great Indian Festival rush. Once the page content is available, BeautifulSoup steps in to read the raw HTML and gently extract useful details like product links, names, and prices without confusion. To keep everything organized, sqlite3 provides a lightweight local database where URLs and scraped tablet data are stored safely, while json allows the same information to be saved in a portable format for sharing or analysis. Throughout the process, logging quietly records each step, making it easier to trace progress or understand issues later, and standard tools like [os](https://www.geeksforgeeks.org/python/os-module-python-examples/?ref=blog.datahut.co) and [datetime](https://deepnote.com/blog/ultimate-guide-to-the-datetime-library-in-python?ref=blog.datahut.co) help manage files and timestamps, ensuring the scraper runs in a clean, predictable way. Together, these libraries form a reliable backbone that turns a complex, dynamic Amazon page into structured tablet data in a way that feels approachable and easy to follow for beginners. ## Step 1: Scraping Product URLs from the Tablets Section on Amazon India ### Importing Libraries ``` Import necessary libraries import asyncio import logging import os import sqlite3 import json from datetime import datetime from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` This set of commonly used Python libraries make [web scraping](https://www.geeksforgeeks.org/blogs/what-is-web-scraping-and-how-to-use-it/?ref=blog.datahut.co) simple and reliable—[Playwright](https://playwright.dev/?ref=blog.datahut.co) is used to open and load Amazon’s tablets listing page just like a real browser, BeautifulSoup helps read and understand the page’s HTML content, and SQLite stores the collected tablet details neatly so they can be reused later, while logging and JSON help track progress and save data in an organized way. ### Foundational Configuration Settings ``` Configurations START_URL = "https://www.amazon.in/s?i=computers&rh=n%3A1375458031&s=popularity-rank&fs=true&ref=lp_1375458031_sar" DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon_tablets/Data/amazon_tablets.db" JSON_PATH = "/home/anusha/Desktop/DATAHUT/Amazon_tablets/Data/amazon_tablets.json" ``` This configuration section sets the foundation for the entire scraping process by clearly defining where the data comes from and where it should be saved, making the workflow easy to understand. The START\_URL points to [Amazon India’s tablets listing page](https://www.amazon.in/s?i=computers&rh=n%3A1375458031&s=popularity-rank&fs=true&ref=lp%5F1375458031%5Fsar), which ensures the scraper consistently starts from a stable and relevant source page, while DB\_PATH specifies the exact location where scraped tablet details will be stored in a local SQLite database for structured access and easy querying later, and JSON\_PATH provides a parallel option to save the same data in JSON format, which is widely used for data sharing, APIs, and analysis; together, these paths act like a roadmap for the scraper, clearly separating data collection from data storage and helping newcomers understand how raw web data moves from a webpage into reusable files for further analysis or reporting. ### Logging Setup for Reliable Tablet Data Scraping ``` Logging LOG_DIR = "/home/anusha/Desktop/DATAHUT/Amazon_tablets/Log" os.makedirs(LOG_DIR, exist_ok=True) logging.basicConfig( filename=os.path.join(LOG_DIR, f"url_scraper_{datetime.now().strftime('%Y%m%d_%H%M%S')}.log"), level=logging.INFO, format="%(asctime)s [%(levelname)s] - %(message)s" ) ``` This [logging setup](https://docs.python.org/3/library/logging.html?ref=blog.datahut.co) acts like a quiet notebook running alongside the scraper, carefully noting what happens at each step so that progress and issues can be reviewed later without interrupting the main task. The code first defines a dedicated folder to store log files and ensures it exists, which helps keep scraping runs organized instead of mixing messages with other project files, and then configures Python’s logging system to automatically create a time-stamped log file that records important events in a clear, readable format; this makes it easier to understand when the scraper started, which pages were processed, and whether any errors occurred, and it also encourages good development habits by showing how real-world data projects rely on logs to debug problems and track long-running tasks, similar to how a delivery receipt helps confirm where a package has been and when it arrived. ### SQLite Database Setup for Storing URLs ``` DB Setup def init_db(): conn = sqlite3.connect(DB_PATH) cur = conn.cursor() cur.execute(""" CREATE TABLE IF NOT EXISTS tablets2 ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, scraped_date TEXT ) """) conn.commit() return conn ``` This [database setup](https://www.tutorialspoint.com/sqlite/index.htm?ref=blog.datahut.co) creates a simple and dependable place to store tablet product links collected from Amazon’s tablets listing page, making the [scraping process](https://www.blog.datahut.co/post/challenges-that-make-amazon-data-scraping-so-painful/) more organized and easier to manage over time. The init\_db function connects to a local SQLite database file and checks whether a table named tablets2 already exists, creating it only if needed so the scraper can be run multiple times without errors, while the table structure itself is kept intentionally simple by storing each product URL along with the date it was scraped, which helps to understand how raw website links are gradually turned into structured data that can be queried, updated, or reused later, much like maintaining a neatly labeled notebook instead of loose pages scattered across a desk. ### Filtering Non-Tablet Products While Scraping Tablet Listings ``` Filtering Logic EXCLUDE_KEYWORDS = [ "Timer", "HandBag", "Back Cover", "Case Cover", "Flip Case", "Volume Button Side Button Out Keys", "Coin Tissue", "Power Volume Button Flex Internal Keys", "Sticker Multicolour", "Photo Stand", "Pencil Case", "Vinyl Stickers", "Tablet Cutter", "NAGARJUNA", "Kashaya", "Kashayam" ] ``` This filtering logic helps keep the scraped data clean and focused by clearly defining which items should be ignored while collecting tablet product links from Amazon’s tablets section. The EXCLUDE\_KEYWORDS list contains common words and phrases that usually belong to accessories, covers, stickers, or unrelated products, and during scraping these terms act like a simple checkpoint—if a product title contains any of them, it is skipped—so only genuine tablet listings move forward in the process, which makes the final dataset more accurate and easier to analyze later. ### Filter unwanted products by HTML structure and keywords ``` Filter unwanted products def is_valid_product(html): """Filter unwanted products by HTML structure and keywords""" # Exclude Ayurvedic & non-tablet products with s-line-clamp-3 if 's-line-clamp-3' in html: return False # Exclude by keywords for word in EXCLUDE_KEYWORDS: if word.lower() in html.lower(): return False return True ``` This product validation function acts as a final quality check to ensure that only genuine tablet listings are saved during the scraping process. The is\_valid\_product function examines the raw HTML of each product card and immediately filters out unrelated items, such as ayurvedic products or accessories, by checking for specific page patterns and previously defined keywords, and if any unwanted signal is found the product is skipped without stopping the scraper; this simple logic helps to understand how small checks, applied at the right stage, can significantly improve data accuracy and reduce noise, much like quickly scanning a document for obvious mismatches before filing it away. ``` Scraper async def scrape_tablets(): conn = init_db() cur = conn.cursor() all_results = [] scraped_date = datetime.now().strftime("%Y-%m-%d") async with async_playwright() as p: browser = await p.firefox.launch(headless=False) page = await browser.new_page() await page.goto(START_URL, timeout=60000) while True: logging.info(f"Scraping page: {page.url}") await page.wait_for_selector("div[data-cy='title-recipe']", timeout=30000) html = await page.content() soup = BeautifulSoup(html, "html.parser") ``` This scraper function is the heart of the entire process, carefully guiding the program from opening Amazon’s tablets listing page to collecting clean and usable product links across multiple pages. In the first part, the scrape\_tablets function prepares the environment by connecting to the SQLite database, creating a cursor for saving data, initializing a list to store results for JSON output, and recording the current date so every scraped link is time-stamped, which helps track when the data was collected. Using Playwright, a real browser session is launched and directed to the Amazon tablets page defined in START\_URL, and once the page loads, the HTML content is captured and passed to [BeautifulSoup](https://pypi.org/project/beautifulsoup4/?ref=blog.datahut.co), which translates the complex webpage structure into something readable and searchable, similar to turning a crowded bookshelf into neatly arranged sections. ``` # Extract product links (normal + sponsored) links = soup.select("div[data-cy='title-recipe'] a.a-link-normal, a.a-link-normal.s-line-clamp-4") logging.info(f"Found {len(links)} links on page") for link in links: href = link.get("href") if not href: continue full_url = "https://www.amazon.in" + href if href.startswith("/") else href product_html = str(link.parent) if not is_valid_product(product_html): logging.info(f"Excluded product: {full_url}") continue ``` The second part focuses on finding the actual product links on the page, including both regular and sponsored listings, by selecting common HTML patterns Amazon uses for tablet titles. Each link is checked carefully to ensure it contains a valid URL, converted into a full Amazon link if needed, and then reviewed using the earlier filtering logic so that accessories or unrelated products do not slip through, reinforcing the idea that good scraping is not just about collecting more data, but about collecting the right data. ``` # Save to DB cur.execute("INSERT INTO tablets2 (url, scraped_date) VALUES (?, ?)", (full_url, scraped_date)) conn.commit() # Save to list for JSON all_results.append({ "url": full_url, "scraped_date": scraped_date }) logging.info(f"Scraped: {full_url}") ``` In the third part, every valid tablet link is saved immediately into the database along with the scrape date, ensuring nothing is lost even if the scraper stops unexpectedly, and at the same time the same information is added to a Python list so it can later be written into a JSON file, showing how the same data can be stored in multiple formats for different use cases, such as analysis or sharing. ``` # Check for "Next" button next_button = soup.select_one("a.s-pagination-next") if next_button and "href" in next_button.attrs: next_url = "https://www.amazon.in" + next_button["href"] await page.goto(next_url, timeout=60000) else: logging.info("No more pages. Exiting.") break await browser.close() ``` The final part handles pagination, which allows the scraper to move smoothly from one results page to the next by checking for the presence of Amazon’s “Next” button and loading the next page when available, and once no further pages are found, the loop ends gracefully and the browser is closed, completing the journey from the first tablet listing to the last without manual intervention, much like flipping through pages of a catalog until the end is reached. ### Running the Scraper Using Main Entry Point ``` Main if __name__ == "__main__": asyncio.run(scrape_tablets()) ``` This main block acts as the official starting switch for the scraper, telling Python exactly when the scraping process should begin. By checking if name == "\_\_main\_\_":, the code ensures that the scrape\_tablets function runs only when this file is executed directly and not when it is imported elsewhere, and [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(scrape\_tablets()) then safely starts the asynchronous scraping task. ## Step2: Collecting Detailed Information from Individual Product Pages ### Importing Libraries ``` Import necessary libraries import asyncio import sqlite3 import logging import os from datetime import datetime from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` ### Logging Configuration ``` Logging Setup LOG_DIR = "/home/anusha/Desktop/DATAHUT/Amazon_tablets/Log" os.makedirs(LOG_DIR, exist_ok=True) logging.basicConfig( filename=os.path.join(LOG_DIR, "data_scraper.log"), level=logging.DEBUG, format="%(asctime)s - %(levelname)s - %(message)s" ) """This section configures the logging system""" ``` This logging setup quietly records everything the scraper does in the background, making it easier to understand how the program behaves while collecting tablet data from Amazon’s listings page. ### SQLite Database Setup for Storing Product Details ``` Database Setup """This section sets up the SQLite database""" DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon_tablets/Data/amazon_tablets.db" def init_db(): conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, name TEXT, brand TEXT, sale_price TEXT, scraped_date TEXT ) """) ``` In the first part of the database setup, the code creates a structured and reliable space to store detailed tablet information scraped from Amazon’s tablets listing page, starting by defining a clear file path for the SQLite database and then opening a connection to it. A table named product\_data is created only if it does not already exist, with carefully chosen columns such as product URL, name, brand, sale price, and scrape date, which helps to understand how raw product listings are gradually transformed into organized records that are easy to query later, similar to arranging product details into clearly labeled columns in a spreadsheet instead of leaving them scattered across notes. ``` cursor.execute(""" CREATE TABLE IF NOT EXISTS tablets2 ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, scraped INTEGER DEFAULT 0 ) """) try: cursor.execute("ALTER TABLE tablets2 ADD COLUMN scraped INTEGER DEFAULT 0") except sqlite3.OperationalError: pass conn.commit() return conn ``` The second part focuses on tracking progress during scraping, which becomes especially important when dealing with many pages. Here, a separate table called tablets2 is created to store only product URLs along with a scraped flag that indicates whether a link has already been processed, allowing the scraper to pause and resume without repeating work. The additional check to add the scraped column safely, without breaking the program if it already exists, introduces a practical, real-world pattern used in long-running data tasks, and the final commit ensures all changes are saved before returning the database connection for use in the rest of the scraping workflow. ### Extracting Individual Tablet Details Using Playwright and BeautifulSoup ``` Scraping Logic """Scrape product details from a single Amazon product page""" async def scrape_product(page, url): try: await page.goto(url, timeout=60000) await page.wait_for_timeout(3000) soup = BeautifulSoup(await page.content(), "html.parser") ``` In the first part of this scraping logic, the function focuses on visiting a single Amazon tablet product page and preparing its content for extraction in a safe and readable way. The browser page is directed to the product URL, given a short pause to fully load dynamic content, and then the complete HTML is passed into BeautifulSoup, which converts the complex webpage into a structured format that can be easily searched, helping beginners see how each product page is handled one at a time rather than all at once. ``` def safe_select(selector, attr=None): el = soup.select_one(selector) if not el: return None return el.get(attr) if attr else el.get_text(strip=True) name = safe_select("span#productTitle") if not name: logging.info(f"Skipped empty product: {url}") return None brand = safe_select("tr.po-brand span.po-break-word") or safe_select("a#bylineInfo") sale_price = safe_select("span.a-price span.a-price-whole") scraped_date = datetime.now().strftime("%Y-%m-%d %H:%M:%S") return { "url": url, "name": name, "brand": brand, "sale_price": sale_price, "scraped_date": scraped_date } except Exception as e: logging.error(f"Error scraping {url}: {e}") return None ``` The second part carefully pulls out specific product details while avoiding common errors that can stop a scraper. A small helper function checks whether each piece of information exists before trying to read it, which prevents the program from breaking when a field is missing, and then key details such as the product name, brand, and sale price are collected along with the current date to record when the data was captured. If a page does not contain a valid product title, it is safely skipped, and any unexpected issue is logged for later review, showing how thoughtful checks and error handling help transform raw Amazon pages into clean, usable tablet data suitable for storage and analysis. ### Main Runner Logic for Scraping and Storing Product Data ``` Main Runner async def main(): """This function controls the overall scraping process""" conn = init_db() cursor = conn.cursor() cursor.execute("SELECT id, url FROM tablets2 WHERE scraped = 0") urls = cursor.fetchall() logging.info(f"Total URLs to scrape: {len(urls)}") async with async_playwright() as p: browser = await p.chromium.launch(headless=False) page = await browser.new_page() for idx, (url_id, url) in enumerate(urls, start=1): logging.info(f"Scraping {idx}/{len(urls)}: {url}") data = await scrape_product(page, url) ``` In the first part of the main runner, the main function acts as the control center that coordinates the entire scraping flow from start to finish. It begins by setting up the database connection and selecting only those tablet URLs that have not yet been scraped, using a simple status flag to avoid repeating work, and this list of pending URLs is then logged so it is clear how much data remains to be collected. A Playwright browser session is opened and kept alive while each product page is visited one by one, and for every URL the scraper calls the product-level scraping function, which helps to understand how large tasks are broken into smaller, manageable steps that work together smoothly. ``` if data: try: cursor.execute(""" INSERT INTO product_data (url, name, brand, sale_price, scraped_date) VALUES (?, ?, ?, ?, ?) """, ( data["url"], data["name"], data["brand"], data["sale_price"], data["scraped_date"] )) cursor.execute( "UPDATE tablets2 SET scraped = 1 WHERE id = ?", (url_id,) ) conn.commit() logging.info(f"Saved & marked scraped: {data['name'][:50]}...") except Exception as db_err: logging.error(f"DB error for {url}: {db_err}") await browser.close() conn.close() ``` The second part focuses on saving the collected product details and updating progress in a reliable way. When valid data is returned, the product information is inserted into the main product table, and the corresponding URL is immediately marked as scraped so it will not be processed again in future runs, which is especially useful for long scraping jobs that may need to be paused and resumed. Any [database-related issue ](https://www.loadview-testing.com/blog/5-most-common-database-performance-issues-fixes/?ref=blog.datahut.co)is safely logged without stopping the entire process, and once all URLs are processed the browser and database connections are closed cleanly, completing the journey from raw Amazon tablet links to structured, ready-to-use product data. ### Entry Point for Starting the Scraping Process ``` Entry Point if __name__ == "__main__": asyncio.run(main()) """This block makes sure the script runs only when executed directly""" ``` This entry point acts as the final trigger that starts the entire scraping workflow at the right moment. By checking if name == "\_\_main\_\_":, the script ensures that the main scraping function runs only when this file is executed directly and not when it is imported into another program, and [asyncio](https://docs.python.org/3/library/asyncio.html?ref=blog.datahut.co).run(main()) then safely launches the asynchronous process that ties together database setup, page scraping, and data storage. ## Conclusion This project demonstrates how a well-planned web scraping workflow can turn a fast-moving Amazon tablets sale page into a clean, reliable dataset, even during a high-traffic event like the Great Indian Festival spread across multiple days. Starting from carefully collecting product URLs, moving through controlled page-by-page data extraction, and ending with thoughtful cleaning and preparation, each stage shows that effective scraping is less about speed and more about patience, structure, and accuracy. By using simple tools such as Playwright, BeautifulSoup, SQLite, and OpenRefine, raw and constantly changing product listings are gradually shaped into organized information that can be analyzed with confidence. More importantly, this approach highlights good data practices—handling errors gracefully, avoiding duplicates, tracking progress, and respecting dynamic website behavior—which are essential lessons for beginners stepping into real-world data engineering. In the end, the project is not just about Amazon tablets, but about understanding how disciplined data collection lays the foundation for meaningful insights and informed decision-making. ## AUTHOR I’m Anusha P O, a Data Science Intern at [Datahut](https://www.datahut.co/?ref=blog.datahut.co), with a strong interest in building automated web data workflows that transform large volumes of online information into clean, analysis-ready datasets. This blog focuses on extracting data from the Amazon India tablets section, a category that showcases how modern [e-commerce](https://business.adobe.com/blog/basics/ecommerce-definition?ref=blog.datahut.co) platforms organize products across dynamic and frequently updated pages. By working with real tablet listings—covering details such as product names, brands, prices, and rankings—the blog walks through practical techniques for collecting and structuring data in a way that is reliable, beginner friendly, and scalable. At [Datahut](https://www.blog.datahut.co/), the broader objective is to help businesses unlock value from public web data for use cases like pricing analysis, product research, and competitive insights, and this walk through demonstrates how thoughtful scraping and data organization can turn everyday online listings into meaningful intelligence. ## Frequently Asked Questions (FAQs) ### 1\. What is Playwright and why is it used for web scraping Amazon? Playwright is a modern browser automation framework that allows developers to control browsers like Chromium, Firefox, and WebKit programmatically. It is widely used for web scraping because it can handle dynamic websites, JavaScript-rendered content, and user interactions, making it ideal for extracting product data from complex e-commerce platforms like Amazon. ### 2\. What tablet data can be extracted from Amazon using Playwright? Using Playwright, you can extract various tablet product details from Amazon such as product name, price, ratings, number of reviews, product specifications, brand name, and product URL. This data can be useful for price monitoring, competitor analysis, and market research. ### 3\. Is it legal to scrape product data from Amazon? Web scraping legality depends on how the data is collected and used. Publicly available product information can generally be collected for research and analytics, but it’s important to follow Amazon’s terms of service, respect robots.txt guidelines, and avoid aggressive scraping that may harm website performance. ### 4\. Why is Playwright better than traditional scraping libraries? Unlike traditional scraping libraries that only fetch HTML content, Playwright can fully render JavaScript-driven pages. It simulates real user behavior such as scrolling, clicking, and waiting for elements to load, which makes it more reliable when scraping modern e-commerce websites. ### 5\. What are common challenges when scraping Amazon product data? Common challenges include dynamic page structures, anti-bot detection mechanisms, CAPTCHA prompts, rate limiting, and frequent layout changes. Using techniques like request delays, rotating user agents, and structured selectors can help improve scraping reliability. ### How Skincare Brands Use Web Scraping to Track Competitors and Win on Amazon URL: https://www.blog.datahut.co/post/web-scraping-for-skincare-brands-that-want-to-win/ Last updated: 2026-09-07T09:43:17.000Z The global skincare market crossed $189 billion in 2025\. But revenue share isn't being won on formulation alone- it's being won by brands that can answer questions like: Why are consumers abandoning Product X? Which ingredient is about to become the next retinol? Is a [competitor quietly raising prices?](https://www.realdataapi.com/tata-cliq-personal-care-product-scraper-provides-reliable-insights.php?ref=blog.datahut.co) [Web scraping](https://www.blog.datahut.co/post/web-scraping-best-practices-tips/) has become the intelligence backbone for category-leading brands. From DTC disruptors to legacy conglomerates, structured data collection from competitors, retailers, and consumers provides the signals needed to move faster and smarter than the competition. This guide breaks down six proven use cases, with the data intelligence being non-negotiable, and real-world case study outcomes ## 6 Proven Web Scraping Strategies for Skincare Brands Each use case addresses a specific intelligence gap that costs brands revenue, market position, or consumer trust when left unchecked. ## Competitor Product & Pricing Analysis Systematically scrape [CeraVe](https://www.blog.datahut.co/post/cerave-on-amazon-2026-best-sellers-pricing-review-analysis/), The Ordinary, Neutrogena, and niche DTC brands for SKU names, ingredient lists, pack sizes, and price points. Build a living competitive map that updates weekly. Sources: [cerave.com](http://cerave.com/?ref=blog.datahut.co), [theordinary.com](http://theordinary.com/?ref=blog.datahut.co), [neutrogena.com](http://neutrogena.com/?ref=blog.datahut.co), brand PDPs Example: The Ordinary's Niacinamide 10% + Zinc 1% is priced at $7.90 for 30ml. A competitor scraping this weekly would notice if The Ordinary quietly drops it to $5.90 during a sale and can respond with a bundle offer or promotional price before losing customers.[ A living competitive map might look like:](https://www.blog.datahut.co/post/category-price-index/) Suddenly you can see you're 3× the price per ml with no obvious differentiation communicated on your PDP - a conversion killer. ## Ingredient Trend Monitoring Track which actives are gaining Google search volume, appearing in new product launches, and being discussed positively on Reddit, before they peak. Niacinamide, bakuchiol, and PDRN are recent examples of early signals brands missed. Sources: Google Trends, Reddit, new launches databases Example: In 2021, "bakuchiol" had 40K monthly Google searches. By 2023, it was 180K. Brands that started formulating in 2021 launched at peak demand. Brands that noticed in 2023 were 18 months too late. A weekly tracker watching these signals would flag: - Tranexamic Acid — Reddit mentions up 340% in 12 months - PDRN (Salmon DNA) — appearing in 67% more new Korean beauty launches vs. last year - Polyglutamic Acid — Google Trends index up from 12 → 58 in 8 months That's your R&D roadmap, built from public data. ## Customer Review Mining Extract thousands of reviews from Sephora, Ulta, Amazon, and Reddit. Feed structured review data to an LLM to surface recurring pain points, unmet needs, and positive drivers that quantitative research would never catch. Sources: Sephora, Ulta, Amazon, r/SkincareAddiction Example: A brand scrapes 11,000 one and two-star reviews for competitor SPF moisturizers across Amazon and Sephora. After feeding them to an LLM, the top complaints cluster into three themes: - "Leaves white cast" — mentioned in 38% of negative reviews - "Pilling under makeup" — 29% - "Greasy finish" — 24% Their own formula already solved the white cast issue. They rewrite their PDP headline to "Zero white cast, zero pilling" - directly addressing the [#1](https://www.blog.datahut.co/blog/hashtags/1) and [#2](https://www.blog.datahut.co/blog/hashtags/2) frustrations across the entire category. Conversion rate increases without changing the product at all. ## Consumer Sentiment Analysis Scrape forums, Quora, MakeupAlley, and Reddit threads to build a pain-point database. Identify which skin concerns (acne, hyperpigmentation, redness) lack satisfying solutions — your next product brief is hiding in these threads. Sources: Reddit, Quora, MakeupAlley, forums Example: A brand scrapes 6 months of posts from r/SkincareAddiction tagged "hyperpigmentation." The LLM analysis reveals a pattern: users love Vitamin C for its brightness but constantly complain that it oxidises quickly, smells bad, or irritates sensitive skin. Nobody has launched a stable, fragrance-free, sensitivity-tested Vitamin C that markets itself on exactly those three solved problems. That's a product brief. Straight from 50,000 unfiltered consumer opinions. No focus group needed. ## Retail & Marketplace Monitoring Track your own brand's listings on Amazon, Sephora, and Ulta for unauthorized sellers, price undercutting, Buy Box ownership shifts, and review velocity changes. This is brand protection as much as intelligence. Sources: Amazon, Sephora, Ulta, ASOS Example: A brand selling their $38 moisturiser on Amazon notices through weekly scraping that: - 6 third-party sellers are listing it at $22–$28 - They've lost the Buy Box 40% of the time this month - Their average star rating dropped from 4.6 → 4.2 in 30 days (likely counterfeit product driving bad reviews) Without scraping, this goes unnoticed for months. With it, they issue MAP violation notices within the week, protect the margin, and flag the counterfeit issue to Amazon before it damages the brand permanently. ## Regulatory & Safety Database Scraping Monitor EWG Skin Deep, INCI Decoder, and CosDNA for ingredient safety updates, banned substance lists by region (EU, US, Korea), and emerging consumer concerns around specific chemicals before they become PR crises. Sources: EWG Skin Deep, INCI Decoder, CosDNA Example: The EU banned 23 new cosmetic ingredients in 2023 alone. A brand selling in both the US and EU markets scrapes EWG Skin Deep and the EU Cosmetics Regulation database monthly. Their scraper flags that Butylphenyl Methylpropional (Lilial) — a fragrance ingredient in their best-selling cleanser — has just been added to the EU banned list. They have 8 months to reformulate before their EU retail contracts are at risk. Without the scraper, they'd find out when a retailer pulls the product from shelves. ## Why [Data Intelligence](https://gitnux.org/beauty-skincare-industry-statistics/?ref=blog.datahut.co) is Non-Negotiable in Beauty ![Why Data Intelligence is Non-Negotiable in Beauty](https://www.blog.datahut.co/content/images/2026/07/img-452.png.webp) ## Real-World Outcome from [Skincare Intelligence Programs](https://zipdo.co/beauty-skincare-industry-statistics/?ref=blog.datahut.co) ## Key Takeaway MAP violations are invisible without systematic monitoring. A daily scraping program turns a brand-protection problem into a solved, automated process - recovering margin and consumer trust simultaneously. The 90-day payback window on a scraping infrastructure investment is rarely matched by any other operational spend in the e-commerce stack. ## Conclusion In a $198B category growing at 6.5% annually, intuition is no longer a strategy - it’s a liability. The brands winning in 2025 aren’t guessing what consumers want, reacting late to ingredient trends, or discovering pricing leaks months after margin erosion. They’re building structured intelligence systems powered by continuous, compliant web scraping. From tracking price-per-ml gaps against competitors like CeraVe and The Ordinary, to mining review data on Amazon and conversations inside Reddit, the advantage comes from visibility. Visibility into trends before they peak. Visibility into sentiment before it becomes churn. Visibility into unauthorized sellers before they damage brand equity. Web scraping is no longer just a technical capability - it’s a competitive moat. It transforms: - Pricing chaos into margin control - Ingredient speculation into data-backed R&D - Review noise into conversion messaging - Marketplace risk into brand protection - Regulatory surprises into proactive compliance In beauty, speed compounds. The brand that detects a signal 6–12 months earlier owns the demand curve. The brand that monitors daily protects its margin. The brand that listens at scale builds products consumers were already asking for. Formulation still matters. Branding still matters. But in 2025, data intelligence determines who scales and who stalls. If skincare is science, then growth is analytics and the brands that operationalize public data will define the next decade of beauty. Ready to turn public skincare data into a real competitive advantage? With [Datahut’s](https://www.datahut.co/?ref=blog.datahut.co) web scraping and retail intelligence solutions, skincare brands can monitor competitor pricing, track ingredient trends, analyze consumer sentiment, and protect marketplace margins - all through reliable, structured data pipelines. Talk to Datahut today and transform scattered web data into actionable beauty market intelligence. ## Frequently Asked Questions (FAQs) ### 1\. How can web scraping help skincare brands stay competitive? Web scraping helps skincare brands collect publicly available data from competitor websites, marketplaces, and consumer forums. This data enables brands to monitor pricing changes, analyze ingredient trends, track consumer sentiment, and identify emerging skincare demands before competitors do. ### 2\. Is web scraping legal for beauty and skincare market research? Yes, web scraping is legal when brands collect publicly available data and comply with website terms of service, privacy laws, and regulations. Responsible scraping focuses on publicly accessible information and avoids personal or sensitive data. ### 3\. What types of data do skincare brands usually scrape? Skincare brands typically scrape product prices, ingredient lists, customer reviews, ratings, SKU information, retailer listings, and competitor product launches from e-commerce platforms and brand websites. ### 4\. Can web scraping help identify new skincare ingredient trends? Yes. By monitoring search trends, forums, product launches, and discussions across communities like Reddit, brands can detect emerging ingredients such as bakuchiol, PDRN, or tranexamic acid months before they reach mainstream demand. ### 5\. How often should skincare brands collect competitor and marketplace data? Most brands track competitor pricing and marketplace listings daily or weekly. Ingredient trends and consumer sentiment analysis are usually monitored weekly or monthly to identify emerging opportunities. ### Why Most Cannabis Brands Are Losing Market Share (And What the Data Says to Do Instead) URL: https://www.blog.datahut.co/post/why-cannabis-brands-lose-market-share/ Last updated: 2026-09-07T09:43:19.000Z According to Research and Markets, the cannabis-infused products market is projected to[ ](https://www.researchandmarkets.com/reports/6076245/cannabis-infused-products-market-report?ref=blog.datahut.co)[grow from $33.62 billion in 2025 to $41.44 billion](https://www.researchandmarkets.com/reports/6076245/cannabis-infused-products-market-report?ref=blog.datahut.co) in 2026, a compound annual growth rate of 23.2%. That trajectory is already visible on the platform: 112 active brands are competing across 18,707 product listings, with a category-average sale price of $29.64 and an average customer rating of 4.57 out of 5.0. With 112 active brands battling for visibility across \*With 18,707 product listings, the digital shelf is incredibly crowded on Weedmaps. Shoppers are presented with a wide range of pricing and product variety, anchored by an accessible category-average sale price of $29.64 and a strong average rating of 4.57 out of 5.0. This report breaks down the data to show where the opportunities are. ​Methodology: This report analyzes publicly available marketplace data scraped from Weedmaps via Datahut's web scraping services, covering 112 active brands across 18,707 product listings. Brand-level analysis is limited to brands with over 1,000 customer reviews to ensure statistical significance and focus on established market trends. In a market this fragmented, capturing consumer attention requires more than just a good product. Category leaders like STIIIZY, West Coast Cure, and Eighth Brother currently dominate the space, leveraging[ deep product assortments](https://www.blog.datahut.co/post/web-scraping-to-extract-product-data-from-e-commerce-sites/) and massive review volumes to win customer trust. For mid-tier and emerging brands, understanding how pricing, promotions, and customer sentiment shape conversions is the only proven way to[ maximize product visibility](https://developers.google.com/search/docs/appearance/structured-data/product?ref=blog.datahut.co) and successfully steal market share shaping the marketplace. Let's dive into details of why Cannabis Brands Are Losing Market Share: ## Key Findings from Weedmaps Category Analysis: Cannabis Brands Are Losing Market Share Our analysis of 18,707 product listings across the Weedmaps cannabis category surfaces three defining patterns in how this market actually works: - A small group of brands commands disproportionate attention. The top 8% of listings account for nearly 80% of all customer reviews — a concentration that signals just how difficult it is for newer or smaller brands to break through without a deliberate visibility strategy. - This is a volume-driven, price-sensitive market. 92% of listings are priced below $60, and the competitive pressure within that range is intense. Winning here isn't about being the cheapest — it's about delivering enough perceived value to convert at scale. - Premium pricing doesn't guarantee better sentiment. Budget products priced at $30 or under carry a slightly higher average rating (4.58) than premium-tier products (4.53). Consumer satisfaction in this category is driven by expectation alignment, not price point. ## Weedmaps Cannabis Category Overview : ![Key Findings from Weedmaps Cannabis Category Analysis ](https://www.blog.datahut.co/content/images/2026/07/img-59.jpg.webp) ## Category Leaders STIIIZY STIIIZY is a clear category leader, maintaining a dominant market share through high review counts and a wide product assortment. Their strategy focuses on visibility and brand recognition, making them a primary driver of shopper engagement across the platform. West Coast Cure West Coast Cure holds a leading position by leveraging consistent ratings and competitive pricing strategies. Their strong market presence is built on high marketing visibility and a diverse range of products that appeal to both budget and mid-range shoppers. ## How Brands Compete for Shelf Space: Product Distribution This analysis shows which brands are physically taking up the most room on Weedmaps. By launching a massive number of products, these brands ensure they are the first thing a shopper sees. ![How Brands Compete for Shelf Space: Product  Distribution](https://www.blog.datahut.co/content/images/2026/07/img-60.jpg.webp) Key Insights - Brands like STIIIZY (1,233 products) and Jeeter (1,216 products) aren't just popular—they are everywhere. - When a brand has over 1,000 products listed, they show up in almost every search, whether a customer is looking for a specific strain, a certain weight, or just browsing a category. - With so many listings from just a few top brands, it becomes much harder for the other 112 brands to get their products seen by shoppers ### Why this matters This is not just about having more stuff. It's a calculated move to own the Search Dominance.If a brand owns 10% of all the products in the category, they effectively own 10% of the customer's attention before a single click even happens—a proven strategy for[ ](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights?ref=blog.datahut.co)[maximizing digital shelf visibility](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights?ref=blog.datahut.co) and driving market share. ## Which brands have the highest average price? Using deep[ ](https://www.blog.datahut.co/post/ecommerce-competitor-analysis/)[competitor pricing analysis](https://www.blog.datahut.co/post/ecommerce-competitor-analysis/), this section compares the average sale price across different brands to see how they position themselves on the value vs. luxury scale. Each bar represents what a customer pays on average for a product from that brand. ![Which brands have the highest average price?](https://www.blog.datahut.co/content/images/2026/07/img-61.jpg.webp) Key Insights - Brands like Sunday Goods ($72.50) and House Party ($66.88) sit at the top of the chart. These brands are successfully maintaining a "Premium" status, charging significantly more than the category average. - A large cluster of popular brands like 710 Labs ($62.17) and Blem ($54.88) compete in the mid-to-high range. This shows that many successful brands aim for a "Premium-but-accessible" price point to capture both quality-seekers and daily users. - Notice how many brands (like Heavy Hitters and Canabotanica) are grouped tightly around the $43 mark. This indicates a "sweet spot" where brands feel they must price themselves to stay competitive without losing their profit margins. Most of the market volume is happening at lower price points. Brands that can keep their average price closer to the $30–$40 range have a much easier time appearing in more "Budget" and "Mid-range" search results. ## Price Band Analysis – How Cannabis Products Are Distributed Across Budget, Mid-Range, and Premium Tiers ![Price Band Analysis – How Cannabis Products Are Distributed Across Budget, Mid-Range, and Premium Tiers ](https://www.blog.datahut.co/content/images/2026/07/img-62.jpg.webp) This table explains how prices are distributed across the category. Understanding price bands helps to identify where most products are positioned, how brands target different customer groups, and whether higher-priced products deliver better customer satisfaction. The table summarizes how many products fall into each band, along with the average rating, average price, and the dominant brands for that group. Key Insights - The Budget segment (≤ $30) is by far the largest group, with 12,098 products. This means most cannabis products on Weedmaps fall into the low-price category. Despite being the cheapest tier (average price $17.60), it actually has the highest average rating of 4.58, showing that customers are highly satisfied with budget-friendly options. Brands like STIIIZY and Jeeter heavily dominate this space. - The Mid-range segment ($31–$60) has 5,209 products. This segment strikes a balance between accessibility and premium positioning. With an average price of $42.85 and a strong average rating of 4.54, this tier is highly competitive, led by brands like COLDFIRE Extracts and 710 Labs. - The Premium segment ($61+) has only 1,400 products, making it the smallest, most exclusive group. Even though these products are priced much higher (average price $84.60), the average rating is 4.53, slightly lower than the budget tier.This indicates that paying more does not necessarily translate to higher customer satisfaction, reinforcing[ ](https://hbr.org/topic/pricing?ref=blog.datahut.co)[established principles of consumer price sensitivity](https://hbr.org/topic/pricing?ref=blog.datahut.co) where a higher price tag must be met with exponentially higher perceived value. 710 Labs and Heavy Hitters own this niche. - Pricing does not strongly influence ratings: All segments have very similar, high ratings (4.53–4.58), showing that customer satisfaction is not strictly dependent on paying a premium. In fact, the Budget tier scores the highest. The price band analysis shows that the cannabis category on Weedmaps is heavily skewed toward budget products, with consistently high customer satisfaction across all segments. Budget products emerge as the sweet spot for consumer value and sheer volume, while premium offerings remain a niche segment that does not necessarily deliver higher ratings despite the steep jump in price. ## Review Concentration Analysis – How Consumer Attention is Evenly Distributed Across the Catalog ![Review Concentration Analysis – How Consumer Attention is Evenly Distributed Across the Catalog ](https://www.blog.datahut.co/content/images/2026/07/img-63.jpg.webp) This table shows how customer ratings are distributed across top-performing products, helping us understand visibility and competition within the category. Key Insights - Competitiveness (Extreme Fragmentation): The Weedmaps cannabis category is incredibly fragmented. With the Top 10 products controlling only 3.64% of all customer reviews, there is absolutely no "winner-take-all" monopoly. The market is highly competitive, but the engagement is evenly spread out across the catalog. - Search Visibility & Dominance: Because no single product hoards the ratings, search visibility is not locked down by a handful of mega-hits. While top brands (like STIIIZY) dominate search by listing thousands of items, no individual product has a large enough "Review Moat" to permanently block competitors from the top search results. - Barrier to Entry for New Products: Breaking into this category is significantly easier than in traditional retail. Because the market is completely decentralized, new entrants do not need to spend massive promotional budgets trying to outrank a "hero" product with 100,000 reviews. A high-quality new drop can quickly capture its own share of the market. - Demand Distribution: Consumer demand is widely spread out across a massive variety of options rather than being hyper-focused on bestsellers. This proves that cannabis shoppers are explorers—they actively seek out new strains, localized drops, and variety rather than just blindly repurchasing the [#1](https://www.blog.datahut.co/blog/hashtags/1) overall item. ## Voice of the Customer - What Are the Ratings Telling Us? ## Average Review Count Across Brands ![Voice of the Customer - What Are the Ratings Telling Us? Average Review Count  Across Brands ](https://www.blog.datahut.co/content/images/2026/07/img-64.jpg.webp) Key Insights - Brands like Froot (384 average reviews) and Sparkiez (277) get the most reviews per product. This shows their items aren't just sitting on dispensary menus—people are actively buying and reviewing them much faster than the competition. - Even though the overall market is very spread out, top brands like Froot, Sparkiez, and ST IDES (178) still stand out. Because their individual products get hundreds of reviews, they naturally show up higher in search results and build customer trust. This proves a brand can get great visibility just by having well-liked products, without needing to take over the whole market. - Customers deeply engage with brands that focus on a few great items rather than those that just flood the market with options. Shoppers actively seek out brands like PLUGPLAY™ (128) and Uncle Arnies (112). On the other hand, massive brands like STIIIZY spread their reviews thin across thousands of items, dropping their average to just 47 reviews per product. This proves that having a focused lineup builds stronger customer interest than just having a massive inventory. ## Average ratings Where Brands Stackup ![Average ratings Where Brands Stackup](https://www.blog.datahut.co/content/images/2026/07/img-65.jpg.webp) Key Insights - The Highest-Rated Leaders: Froot (averaging 4.8 stars) and Uncle Arnies (4.5 stars) are the clear winners for customer satisfaction. They are the absolute best-rated brands in the entire category by a wide margin. - Consistent Quality Across the Board: Because this chart only includes brands that sell at least five different items, we know Froot and Uncle Arnies aren't just getting lucky with one single good product. They consistently deliver high quality across their entire lineups, proving they have enough good products to matter and build real customer trust. - The Steep Drop in Satisfaction: Keeping customers happy across multiple products is clearly difficult in this market. After the top two brands, ratings drop off fast—the third-place brand, LEVEL, falls to 3.8 stars, and well-known names like Cannabiotix drop near 3.1\. Because true 4.5+ star consistency is so rare, Froot and Uncle Arnies are the exact brands we should study further when[ ](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/)[scraping customer reviews for sentiment analysis](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/) to see what they are doing right. ## Revenue Drivers of the Category: What Influences Customer Decisions ![Revenue Drivers of the Category: What Influences Customer Decisions](https://www.blog.datahut.co/content/images/2026/07/img-66.jpg.webp) This chart highlights the most frequently sought-after cannabis effects, showing exactly what shoppers want to feel when making a purchase decision. "Relaxed" leads by a clear margin, followed closely by "Happy," indicating that customers prioritize stress relief and mood elevation over everything else. "Euphoric" also plays a strong role, reinforcing the desire for a highly positive experience, while functional effects like "Sleepy" and "Energetic," though mentioned less often, remain important for targeted, specific needs. Overall, the chart shows that physical recovery and mental unwinding are the absolute strongest drivers of customer intent and conversion in this category. ![Emotional drivers](https://www.blog.datahut.co/content/images/2026/07/img-67.jpg.webp) ## Final Thoughts — What the Data Really Tells Us The Weedmaps cannabis marketplace is highly fragmented, but the data reveals a clear blueprint for winning. Success here splits along two distinct paths: brands that dominate through sheer scale, and brands that win through deep consumer trust. Mega-brands like STIIIZY flood the market with SKU depth, making themselves impossible to ignore through sheer accessibility. But volume alone is not the whole story. Brands like Froot and Uncle Arnies have built a different kind of competitive advantage — one grounded in product excellence. Near-perfect ratings combined with massive review counts generate the kind of organic platform visibility that paid placement simply cannot replicate. In a crowded marketplace, that accumulated social proof functions as a structural moat: hard to build, and even harder to displace. The data also reveals something important about how cannabis consumers actually make decisions. Today's shopper is not just selecting a product — they are selecting an outcome. "Stress Relief" and "Mood Elevation" consistently rank as the dominant intent signals driving purchase behavior, and brands that align their positioning directly with these emotional drivers convert at measurably higher rates. This dynamic plays out across every price tier, from the $43 entry-level items that offer accessible relief to the $72+ premium products where the promise of a precise, high-quality experience justifies the spend. When a product's claimed effect matches a shopper's specific need, price resistance drops — and the review depth built by category leaders is what closes the remaining gap in confidence. What this ultimately means is that competing effectively on Weedmaps is not a creative challenge — it is a data challenge. Brands and dispensaries that operate with clean, structured, and timely market intelligence are not guessing at what shoppers want. They are responding to real behavioral signals in real time. Datahut provides the enterprise-grade data infrastructure needed to translate these raw marketplace signals into actionable competitive advantage — powering smarter decisions across pricing, assortment, and product str ## Your competitors are already watching this data. Are you? Every day without reliable market intelligence is a day your pricing is off, your assortment has gaps, and a competitor is filling shelf space that should be yours. Datahut delivers the structured, real-time cannabis marketplace data that turns those blind spots into your most defensible advantages. No scraping headaches. No stale exports. No guesswork. See exactly what your category looks like through clean, enterprise-grade data — before your next product decision, pricing move, or market expansion. Already know what you need? Talk to a data strategist today. Frequently Asked Questions (FAQs) 1\. Is the Weedmaps cannabis market too crowded for new brands to succeed? Not at all. While the digital shelf is crowded with 112 active brands, the Weedmaps cannabis category is incredibly fragmented. The top 10 products control only 3.64% of all customer reviews, meaning there is no "winner-take-all" monopoly. Because no single product hoards search visibility, a high-quality new drop can quickly capture its own share of the market, making the barrier to entry significantly easier than in traditional retail. 2\. Do premium-priced cannabis products get better customer reviews? Surprisingly, no—higher prices do not buy better sentiment. Our analysis shows that budget products ($30 or less) maintain a slightly higher average rating (4.58) than premium products priced over $61 (4.53) . This indicates that customer satisfaction is not strictly dependent on paying a premium, and the budget segment emerges as a sweet spot for both consumer value and high ratings. 3\. How do mega-brands like STIIIZY dominate search results? STIIIZY and similar brands list 1,000+ products, ensuring they appear across nearly every search query. More listings means more touchpoints before a shopper ever clicks a competitor. 4\. How can smaller or mid-tier brands compete against massive inventories? Volume isn't the only path to victory. Brands like Froot and Uncle Arnies have built an unbeatable competitive advantage through product excellence and targeted consumer trust. By focusing on generating hundreds of reviews per product and maintaining near-perfect ratings, these brands secure highly valuable organic visibility without needing to flood the market with thousands of options . 5\. What is the most popular pricing "sweet spot" for new products? Most of the market volume (92%) competes below the $60 price point. Within that range, a large cluster of popular brands groups tightly around the $43 mark, indicating a highly competitive "sweet spot" where brands feel they must price themselves to stay relevant without losing profit margins . However, brands that can keep their average price closer to the $30–$40 range have an easier time appearing in high-volume "Budget" and "Mid-range" search results. 6\. What are the strongest psychological drivers making customers click "buy"? The modern cannabis shopper is buying an emotional outcome, not just a product . "Relaxed" and "Happy" lead by a clear margin as the most sought-after effects. This proves that physical recovery and mental unwinding (Stress Relief and Mood Elevation) are the absolute strongest drivers of customer intent and conversion. 7\. How can my brand or dispensary track competitor strategies on Weedmaps? Navigating the complex matrix of pricing strategies and consumer sentiment is impossible without a clear view of the market. Datahut provides enterprise-grade, structured data delivery that allows you to track competitors' pricing in near real-time, benchmark your share of voice, and identify hidden revenue leaks . You can talk to a Datahut Data Strategist to get a free 10-minute data audit and start turning raw marketplace signals into a definitive competitive edge . ### How to Scrape Data from Noon’s Fragrance Store? URL: https://www.blog.datahut.co/post/how-to-scrape-data-from-noon-s-fragrance-store/ Last updated: 2026-09-07T09:43:21.000Z Have you ever wondered how to collect product information from online stores without copying everything by hand? In this blog, I’ll walk you through a simple project where we gather data from Noon, a well-known shopping website. We’ll be focusing on fragrance products—and by the end, you’ll see how we can collect, clean, and make sense of that data using a bit of[ Python](https://docs.python.org/3/library/logging.html?ref=blog.datahut.co) code. [Web scraping](https://www.blog.datahut.co/post/python-web-scraping-tutorial/) is just a way of telling the computer, “Hey, go to this website and bring me back the information I need.” Instead of manually going through hundreds of pages and copying prices or product names, we can write a small program that does it for us—faster and more accurately. This technique is useful for anyone who wants to track prices, compare brands, or study how the market changes over time. We chose[ Noon](https://www.noon.com/uae-en/?ref=blog.datahut.co) because it’s one of the top online stores in the Middle East, and it has a wide variety of fragrances to explore. That makes it a great example to practice with. Fragrance products also show interesting trends when we look at them closely—things like brand popularity, price ranges, and customer ratings. To keep things simple, we’ll split the scraping process into two main steps: 1. Collecting the product links – First, we’ll go through the fragrance category and grab the individual product page links. These links are important because they’ll guide us to where the real details are. 2. Extracting the product details – Next, we’ll visit each of those product pages and pull out the information we care about, like the brand, price, rating, and description. Breaking it into two steps makes our code cleaner and easier to manage if something goes wrong along the way. [By the time you finish this guide](https://www.blog.datahut.co/post/noon-skincare-category-analysis/), you’ll not only have working code but also a clear idea of how to scrape product data from Noon—or even any other website with a similar structure. So, let’s dive in and see how we Scrape Data from Noon’s Fragrance Store and can put all the pieces together to build a simple but [powerful data scraping workflow](https://www.blog.datahut.co/post/top-10-web-scraping-companies-in-2026/)! ## Scrape Data from Noon’s Fragrance Store: Link Harvesting Now let’s turn to the first part of our web scraping code. This code is designed to scrape product URLs from the fragrance category of Noon’s website, [using Python](https://www.blog.datahut.co/post/web-scraping-in-python/) as well as modern technology such as asynchronous programming followed by browser automation. Essentially, the code moves through all the various fragrance categories, collecting all the product links, regardless of how many pages of product links there may be. In addition, the code includes error logging for tracking any issues, and it saves all the links in a SQLite database for future retrieval in the next part of the data collection process. Next, we’ll move on to the part where we visit these product pages and collect the data we’re really after. ### Setting Up the Foundation: Imports and Database Configuration ``` import asyncio from playwright.async_api import async_playwright import sqlite3 from bs4 import BeautifulSoup import logging # Setup logging logging.basicConfig(filename="scraping_errors.log", level=logging.ERROR) ``` Before we dive into writing the scraping logic, the script starts by importing a few important Python libraries. Think of these like tools in a toolbox—each one has a specific job to help us get things done. First, we have asyncio. This lets us run several tasks at the same time, which means we can scrape many pages faster instead of waiting for each one to finish before starting the next. Next is[ ](https://playwright.dev/?ref=blog.datahut.co)[playwright](https://playwright.dev/?ref=blog.datahut.co). This tool allows us to control a real web browser with code. It’s especially helpful for websites like Noon, where some content loads only when you interact with the page—just like a human would. Then comes sqlite3, which helps us store all the product links we collect in a small, local database. This makes it easy to keep our data organized and ready to use later. We also use [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co), a simple but powerful library for reading and picking out specific parts of a web page. It’s like using a magnifying glass to zoom in on the exact bits of text or data we need. Finally, logging is there to keep track of anything that goes wrong. If something fails while the script is running, logging will make a note of it so we can understand and fix the issue later. Together, these libraries give us everything we need to build a scraper that’s both efficient and dependable. ``` # SQLite database setup def setup_db(): """ Set up the SQLite database connection and create the product_urls table if it doesn't exist. Returns: tuple: A tuple containing (connection, cursor) objects for database operations. The database schema includes: - id: Auto-incrementing primary key - category: The fragrance category (women, men, unisex) - product_url: The complete URL to the product page """ conn = sqlite3.connect("product_urls.db") cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, product_url TEXT ) """) conn.commit() return conn, cursor ``` Now that we’ve got our tools ready, the next step is to set up a place to store the product URLs we’ll be collecting. For that, we’re using our small database SQLite. This helps us keep our data safe and organized, even after the script finishes running. We create this setup using a simple function that builds a table named product\_urls. Think of this table like a spreadsheet with three columns: 1. ID – A unique number that goes up automatically each time a new product is added. This helps us keep track of how many links we’ve collected. 2. Category – Whether the product is from the men's or women’s fragrance section. 3. Product URL – The actual link to the product page. By storing everything in this format, it becomes much easier to sort, search, or filter the data later—especially if we want to look only at certain categories. Also, keeping the scraping and analysis parts separate like this makes our workflow cleaner. We collect the data first, store it safely, and then later we can use it for analysis without having to scrape again. ### Browser Initialization: Setting Up Playwright ``` # Function to initialize the browser and page async def init_browser(): """ Initialize Playwright browser and page with custom headers to mimic a real user. Returns: tuple: A tuple containing (browser, page) objects for web automation. If initialization fails, returns (None, None). The function: 1. Starts a Playwright instance 2. Launches a non-headless Chromium browser (visible UI) 3. Creates a new page with custom headers to avoid detection as a bot Exceptions are logged to the error log file. """ try: playwright = await async_playwright().start() browser = await playwright.chromium.launch(headless=False) page = await browser.new_page() # Set extra HTTP headers for all requests await page.set_extra_http_headers({ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Connection": "keep-alive", }) return browser, page except Exception as e: logging.error(f"Error initializing browser: {str(e)}") return None, None ``` Next, we move on to setting up the browser that will do the actual visiting and clicking around on the Noon website. This is done through a function called init\_browser(), which uses Playwright to launch a browser we can control with code. Playwright is especially helpful because it behaves just like a real web browser. That means it can load pages that rely on JavaScript—just like Noon’s product listings, where the content doesn’t appear right away but loads after the page finishes rendering. Inside this function, we also customize the HTTP headers—specifically the User-Agent. This is a small detail that tells the website what kind of browser is visiting. By setting it to mimic a real user’s browser (like Chrome or Firefox), we reduce the chances of being blocked or flagged as a bot. The function gives us back two things: - The browser – this is like the full window of the browser we launched. - The page – this is like a single tab where we’ll open and interact with different product pages. Lastly, the function includes error handling, so if something goes wrong while setting up the browser, we’ll get a clear message in the logs. That way, we can troubleshoot easily without guessing what failed. ### Resource Management: Properly Closing Browser Sessions ``` # Function to close the browser async def close_browser(browser): """ Safely close the Playwright browser instance. Args: browser: The Playwright browser instance to close. Exceptions during closing are logged to the error log file. """ try: await browser.close() except Exception as e: logging.error(f"Error closing browser: {str(e)}") ``` Once we’re done scraping, we need to shut things down properly. That’s where the close\_browser() function comes in—it closes the Playwright browser when we’ve finished collecting all the data. While closing a browser might sound like a small task, it’s actually quite important. If we leave the browser open, it keeps using up memory and system resources, which can slow things down—especially if the script runs for a long time or visits hundreds of pages. To handle this properly, the function uses a try-except block. This means it tries to close the browser as expected, but if something goes wrong, it catches the error and logs it. That way, we’re not left guessing why the script didn’t finish cleanly, and we can go back later to fix any issues if needed. Managing resources like this is a smart habit to build. It keeps our script efficient, avoids system slowdowns, and makes sure everything runs smoothly even with large or long-running scraping tasks. ### The Heart of the Operation: Scraping Category URLs ``` async def scrape_category_urls(page, category_url, category, cursor): """ Scrape product URLs by navigating through all pages of a category. Args: page: Playwright page object for web interaction category_url (str): The URL of the category to scrape category (str): The name of the category being scraped (women, men, unisex) cursor: SQLite cursor for database operations Returns: list: A list of all product URLs scraped from all pages of the category The function: 1. Navigates to the initial category page 2. Extracts product URLs from the current page 3. Saves URLs to the database 4. Clicks the 'Next' button to navigate to the next page 5. Repeats steps 2-4 until no more pages are available For each page, the function scrolls down to ensure all products are loaded, with a short pause between scrolls to allow content to load. """ await page.goto(category_url, timeout=60000) pages = 1 total_urls = [] while True: print(f"\nScraping Page {pages}...") try: # Scroll down to load all products for _ in range(10): await page.evaluate( "window.scrollTo(0, document.body.scrollHeight)" ) await asyncio.sleep(2) content = await page.content() extracted_urls = extract_product_urls(content) if not extracted_urls: print("No products found on this page. Ending.") break save_urls_to_db(cursor, extracted_urls, category) total_urls.extend(extracted_urls) print(f"✅ Scraped and saved {len(extracted_urls)} URLs from page {pages}") except Exception as e: logging.error(f"Error scraping page {pages}: {str(e)}") break # Try to click the 'Next' button try: next_button = await page.query_selector('#catalog-page-container > div > div.ProductListDesktop_container__08z7c > div.ProductListDesktop_content__3KHXe > div.PlpPagination_paginationWrapper__1AFsm > div > ul > li.next > a') if next_button: await next_button.click() await page.wait_for_timeout(3000) # wait for next page to load pages += 1 else: print("🚫 No 'Next' button found. Reached last page.") break except Exception as e: print(f"❌ Exception while clicking next: {str(e)}") break return total_urls ``` Now let’s talk about the scrape\_category\_urls() function. This part of the script does the real legwork—it goes through a fragrance category on the Noon website and collects the product links one page at a time. Here’s how it works behind the scenes: First, the function opens the category page and scrolls to the bottom. This step is important because many websites, including Noon, load more products only when you scroll down. So this ensures we’re seeing everything. Next, it pauses briefly to give the page enough time to finish loading. Then, it grabs the HTML content of the page and passes it to another function, which is in charge of pulling out the product URLs from that HTML. Once we have the links, they’re saved into our SQLite database, which keeps everything neatly stored for later use. The function then looks for a "Next" button on the page. If it finds one, it clicks it to move to the next page and repeats the process. This loop continues automatically until there are no more pages left to visit. We don’t have to tell the script how many pages to expect—it just keeps going until it reaches the end. What makes this function solid is that it includes error handling and shows live progress messages. So if something goes wrong, we’ll know exactly where and why—and we can come back and fix it without too much trouble. Overall, this approach makes the scraper flexible, smart, and capable of handling changes on the site without hardcoding anything. ### Parsing Product URLs: Extracting What Matters ``` # Function to extract product URLs from the page content def extract_product_urls(content): """ Extract product URLs from the HTML content using BeautifulSoup. Args: content (str): The HTML content of the page to parse Returns: list: A list of complete product URLs extracted from the page The function uses a CSS selector to find product link elements and constructs full URLs by prepending the base domain. """ soup = BeautifulSoup(content, "html.parser") product_links = soup.select("#catalog-page-container > div > div.ProductListDesktop_container__08z7c > div.ProductListDesktop_content__3KHXe > div.ProductListDesktop_layoutWrapper__Kiw3A > div.ProductBoxLinkHandler_linkWrapper__b0qZ9 > a.ProductBoxLinkHandler_productBoxLink__FPhjp") return [f"https://www.noon.com{link.get('href')}" for link in product_links if link.get('href')] ``` The extract\_product\_urls() function plays a key role in our scraping process. This is the part where we actually pull the product links out of the webpage and prepare them to be saved. Here’s what it does, step by step: First, it uses BeautifulSoup to read the HTML of the page. Think of this like scanning a document for certain words or phrases—we’re just telling the program what to look for in the HTML. Next, it uses a CSS selector to find only the specific parts of the page that contain product links. This is important because it avoids picking up random or unrelated links—so we stay focused only on what we need. Often, the links we get from the page are just partial URLs, like /product/xyz123, which by themselves aren’t usable. So the function adds the full Noon website URL to the front, turning it into a complete link like [https://www.noon.com/product/xyz123](https://www.noon.com/product/xyz123?ref=blog.datahut.co). This approach is simple but very effective. It keeps the process clean and fast by focusing only on the relevant elements—and nothing extra. Even though this part of the code might seem small, it’s actually critical to the entire project. If the product links aren’t collected correctly here, then none of the later steps will work. So this function may be small, but it does a big job. ### Data Persistence: Storing URLs in the Database ``` # Function to save URLs to the SQLite database def save_urls_to_db(cursor, urls, category): """ Save the scraped URLs to the SQLite database. Args: cursor: SQLite cursor for database operations urls (list): List of product URLs to save category (str): The category name (women, men, unisex) The function inserts each URL with its category into the product_urls table and commits the transaction. Exceptions are logged to the error log file. """ try: for url in urls: cursor.execute("INSERT INTO product_urls (category, product_url) VALUES (?, ?)", (category, url)) cursor.connection.commit() print(f"Saved {len(urls)} URLs for category: {category}") except Exception as e: logging.error(f"Error saving URLs for category {category}: {str(e)}") ``` Once we’ve collected the product links, the next step is to save them somewhere safe—and that’s exactly what the save\_urls\_to\_db() function does. It stores each product URL in our SQLite database, along with its category. Here’s how it works: The function loops through the list of URLs, and for each one, it adds a new row to the database. Along with the link itself, it also saves the fragrance category (like men's or women’s), so we always know where the link came from. If something goes wrong during this process—like a connection issue or a duplicate entry—the function catches the error and writes it to the log. This way, the script doesn’t crash, and we can look back later to see what happened. Using a database has some clear advantages over simply saving the links to a text file or keeping them in memory. For example: - It organizes the data in a way that’s easy to search, filter, or sort. - It keeps the data safe—even if the script stops halfway, nothing is lost. - It lets us build smarter features later, like skipping links we’ve already saved so we don’t waste time re-scraping. All in all, this function helps us stay organized, efficient, and ready for the next stage in our project. ### Orchestration: Managing the Scraping Process ``` # Function to handle the scraping of one category async def scrape_category(category, category_url, cursor): """ Handle the complete scraping process for a single category. Args: category (str): The category name (women, men, unisex) category_url (str): The URL for the category's product listing page cursor: SQLite cursor for database operations The function: 1. Initializes a browser and page 2. Scrapes all product URLs from the category 3. Closes the browser when finished If browser initialization fails, the error is logged. """ browser, page = await init_browser() if browser and page: print(f"Starting to scrape {category}...") await scrape_category_urls(page, category_url,category, cursor) await close_browser(browser) else: logging.error(f"Failed to initialize browser for {category}") ``` Now let’s look at one of the key functions that manages the actual scraping process—the scrape\_category() function. This function is designed to handle everything needed to scrape one fragrance category, from start to finish. Here’s what it does: - First, it starts the browser using Playwright, so we’re ready to navigate the website. - Then, it runs the scraping steps to collect product URLs from that specific category. - Once that’s done, it closes the browser properly to free up memory and resources. The beauty of this function is how neat and reusable it is. If you only want to scrape, say, women’s perfumes or just one specific category, you can simply call this function without touching the rest of the code. It keeps things modular and avoids the need for major changes when switching between tasks. So whether you’re scraping one section or planning to cover the entire site, this function keeps your workflow clean and flexible. ``` # Main function to manage all categories async def main(): """Orchestrate the entire scraping process for all fragrance categories. The function: 1. Sets up the database connection 2. Defines the category URLs to scrape 3. Iterates through each category and scrapes its product URLs 4. Closes the database connection when finished Category URLs are constructed with filters for: - women's fragrances - men's fragrances - unisex fragrances """ conn, cursor = setup_db() CATEGORY_URLS = { "women": "https://www.noon.com/uae-en/beauty/fragrance/?f[is_fbn][]=1&f[fragrance_department][]=women&sort[by]=popularity&sort[dir]=desc&limit=50&page=1&isCarouselView=false", "men": "https://www.noon.com/uae-en/beauty/fragrance/?f[is_fbn][]=1&f[fragrance_department][]=men&sort[by]=popularity&sort[dir]=desc&limit=50&page=1&isCarouselView=false", "unisex": "https://www.noon.com/uae-en/beauty/fragrance/?f[is_fbn][]=1&f[fragrance_department][]=unisex&sort[by]=popularity&sort[dir]=desc&limit=50&page=1&isCarouselView=false" } # Loop through categories and scrape URLs for category, category_url in CATEGORY_URLS.items(): await scrape_category(category, category_url, cursor) # Close the SQLite connection conn.close() ``` The main() function is the final piece of the puzzle—it’s what actually starts the whole scraping process. Here’s what it takes care of: - It first connects to the database, so we’re ready to store the links we collect. - Then, it defines the fragrance categories we want to scrape. Each category (like men’s or women’s) is paired with its corresponding URL. This setup makes it super easy to update later—just add a new entry to the list if Noon adds a new category. - It then loops through each category and calls the scraping function to collect product links. - Once everything is done, it closes the database to make sure all data is saved properly and resources are released. This function keeps things tidy and gives us a clear starting point for the entire script. And because we’re using asynchronous programming, there’s room for improvement too. In the future, you could easily modify the script to scrape multiple categories at the same time—making the whole process even faster and more efficient. ### Putting It All Together: The Script Execution ``` if __name__ == "__main__": # Entry point: Run the main async function asyncio.run(main()) ``` This line ensures that the main() function is executed only when the variables are executed directly, rather than imported from a different file. It tells the script to initiate the asynchronous scraping process by calling asyncio.run(). Using asyncio also allows the script to accomplish other tasks while waiting on pages to load or waiting on network responses to return, thus making the scraping process much more efficient. ## Deep Data Extraction In the second section of our web scraping project, we are going to go a step further than simply collecting the product URLs. Now it’s time to extract detailed product information from each of the product URLs we previously collected. While the first script was all about finding and saving links, this one will actually visit each product page and collect some useful data, such as:Product title, Pricing, Ratings, Descriptions and Specifications. This step in the web scraping operation is usually known as the "data enrichment" part of the project. We’re further transforming a set of basic URLs into a dataset with rich, structured data that can be analyzed for price comparisons, trends, or market analysis and everything in between. Let's dive into how the code works to achieve this outcome. ### Foundational Setup: Imports and Logging ``` import asyncio import sqlite3 import logging from playwright.async_api import async_playwright from bs4 import BeautifulSoup import random import json logging.basicConfig(filename="scraping.log",level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") ``` At the very start of our script, as usual we import all the Python libraries we need to make everything work smoothly. We bring in: - asyncio – to handle asynchronous tasks, allowing our script to do multiple things at once. - sqlite3 – to manage our local database where all the product URLs will be stored. - playwright and BeautifulSoup – these are the core tools we use to load webpages and extract the information we care about. - random – to help us rotate user agents, which makes our scraper look more like a real person browsing the site and reduces the chance of being blocked. But one of the most important additions to our script is logging. We’ve set up logging so that every message it prints includes a timestamp. This helps us track the script’s progress, figure out where something might go wrong, and even see how long each part of the process takes. When you’re scraping a lot of data, like we did here, having logging is essential. It’s like having a behind-the-scenes logbook that tells you exactly what happened, when it happened, and what might need fixing. With all of this in place, our script isn’t just functional—it’s stable, reliable, and ready for real-world use. Whether you’re working on a side project or building something for production, these small details make a big difference. ### User Agent Rotation: Avoiding Detection ``` def load_user_agents(path="user_agents.txt"): """ Load a list of user agents from a text file. Args: path (str): Path to the text file containing user agents, one per line Returns: list: A list of user agent strings The user agents will be used to randomize the browser's identity for each request, helping to avoid detection and blocking by the website's anti-scraping measures. """ with open(path, "r") as f: return [line.strip() for line in f if line.strip()] def get_random_user_agent(user_agents): """ Select a random user agent from the provided list. Args: user_agents (list): List of user agent strings Returns: str: A randomly selected user agent string """ return random.choice(user_agents) ``` One valuable improvement in this script is the use of user agent rotation. This small but powerful change helps the scraper behave more like a real human browsing the web—rather than a robot making repeated requests. Here’s how it works: every time the script loads a new page, it randomly selects a user agent from a list. A user agent is basically a short message your browser sends to a website saying, “Hi, I’m Chrome on Windows” or “I’m Safari on an iPhone.” Without rotation, the scraper would always look like the same browser and device, which makes it easy for websites to notice unusual patterns. To make rotation happen, the script uses two simple functions: - One that reads user agents from a file (like "user\_agents.txt") - And another that picks one at random whenever a new request is made This approach helps the scraper stay under the radar. Many websites track requests from the same user agent, and if they see too many too quickly, they might block or slow down access. By changing how the scraper “introduces itself” on each page, we make it seem like different people are visiting the site—which is much closer to normal traffic. ### Database Management: Tracking Progress and Storing Results ``` def connect_db(): """ Connect to the SQLite database containing product URLs. Returns: Connection: SQLite database connection object """ return sqlite3.connect("product_urls.db") def ensure_scraped_column(cursor): """ Ensure the product_urls_old table has a 'scraped' column to track progress. Args: cursor: SQLite cursor for database operations The function checks if the 'scraped' column exists and adds it if not. The column stores a boolean flag (0/1) indicating whether a URL has been processed. """ cursor.execute("PRAGMA table_info(product_urls_old)") columns = [col[1] for col in cursor.fetchall()] if "scraped" not in columns: cursor.execute("ALTER TABLE product_urls_old ADD COLUMN scraped INTEGER DEFAULT 0") def fetch_unscraped_urls(cursor): """ Fetch all product URLs that have not been scraped yet. Args: cursor: SQLite cursor for database operations Returns: list: A list of tuples containing (url, category) for unscraped products """ cursor.execute("SELECT url, category FROM product_urls_old WHERE scraped = 0") return cursor.fetchall() def save_product_data(cursor, data): """ Save extracted product data to the database. Args: cursor: SQLite cursor for database operations data (tuple): Product data tuple containing: - url: Product URL (primary key) - category: Product category (women, men, unisex) - name: Product name - rating: Product rating score - total_rating: Number of ratings/reviews - sale_price: Current sale price - price: Original price - discount: Discount percentage - description: Product description - specifications: JSON string of product specifications The function creates the product_data table if it doesn't exist and inserts or updates the product data using the URL as the primary key. """ cursor.execute(''' CREATE TABLE IF NOT EXISTS product_data ( url TEXT PRIMARY KEY, category TEXT, name TEXT, rating TEXT, total_rating TEXT, sale_price TEXT, price TEXT, discount TEXT, description TEXT, specifications TEXT ) ''') cursor.execute(''' INSERT OR REPLACE INTO product_data (url, category, name, rating, total_rating, sale_price, price, discount, description, specifications) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?) ''', data) def mark_url_as_scraped(cursor, url): """ Mark a URL as having been scraped in the database. Args: cursor: SQLite cursor for database operations url (str): The URL to mark as scraped This prevents re-scraping the same URLs if the script is run multiple times. """ cursor.execute("UPDATE product_urls_old SET scraped = 1 WHERE url = ?", (url,)) ``` This script takes a much smarter and more thoughtful approach when it comes to handling the database. It’s no longer just collecting links and hoping for the best—it actually keeps track of what’s been scraped and what still needs to be done. To make this happen, the script introduces several helpful functions: - connect\_db() sets up a connection to our database. - ensure\_scraped\_column() checks that we have a way to mark whether a URL has been scraped or not. - fetch\_unscraped\_urls() pulls out only the links we haven’t visited yet. - save\_product\_data() saves the detailed product information we scrape. - And mark\_url\_as\_scraped() updates the database to show that we’ve already collected data from a particular link. These functions work together to build a scraping process that’s smart and efficient. If something goes wrong—like your internet cuts out or the script crashes—it doesn’t mean starting over. The script can simply resume from where it left off, saving time and preventing duplicate work. The database itself has also been improved. There’s now a new table called product\_data, where we store everything we extract from each product page. This includes basic details like the product URL and category, as well as more specific info like:Product nam, Price and discount, Customer rating, Description and Technical specifications Some of the more detailed info—like product specs—is stored in JSON format, which keeps the data tidy and flexible. Since different products can have different types of specs, this format makes it easier to handle that variety. To avoid duplicates or errors, the script uses a special SQL command: INSERT OR REPLACE. This means if we happen to scrape a product again, it simply updates the existing data instead of creating a messy duplicate. The result? A clean, reliable dataset that’s easy to maintain and ready for deeper analysis later. This kind of setup makes the script feel more like a serious data pipeline—something that’s robust enough for large scraping projects and smart enough to handle real-world interruptions without falling apart. ### Parsing Functions: Extracting Structured Data ``` def parse_name(soup): """ Extract the product name from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The product name if found, None otherwise Uses a CSS selector to locate the product name heading on the page. """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > h1") return tag.get_text(strip=True) if tag else None def parse_rating(soup): """ Extract the product rating score from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The product rating (e.g., "4.5") if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > div.CoreDetails_offerContainer__BNBZp > div.CoreDetails_offerItemCtr__ONued > div.CoreDetails_ratingsAndVariantsCtr__rLNx8 > a > div > div.RatingPreviewStar_starsCtr__cXGit > span") return tag.get_text(strip=True) if tag else None def parse_no_rating(soup): """ Extract the total number of ratings/reviews from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The number of ratings (e.g., "128") if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > div.CoreDetails_offerContainer__BNBZp > div.CoreDetails_offerItemCtr__ONued > div.CoreDetails_ratingsAndVariantsCtr__rLNx8 > a > div > div.RatingPreviewStar_ratingsCountCtr__VHpPi > div > span") return tag.get_text(strip=True) if tag else None def parse_saleprice(soup): """ Extract the current sale price from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The sale price (e.g., "AED 329.00") if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > div.CoreDetails_offerContainer__BNBZp > div.CoreDetails_offerItemCtr__ONued > div.CoreDetails_priceCtr__ZfUY9 > div > div > div > div > span.PriceOffer_priceNowText__08sYH") return tag.get_text(strip=True) if tag else None def parse_price(soup): """ Extract the original price from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The original price (e.g., "AED 455.00") if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > div.CoreDetails_offerContainer__BNBZp > div.CoreDetails_offerItemCtr__ONued > div.CoreDetails_priceCtr__ZfUY9 > div > div > div.PriceOffer_oldAndNewPricesCtr__yhHvc > div.PriceOffer_priceWasCtr__qwKoN > div > span:nth-child(2)") return tag.get_text(strip=True) if tag else None def parse_discount(soup): """ Extract the discount percentage from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The discount percentage (e.g., "28% OFF") if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div.ProductDetailsDesktop_fullWrapper__DGQAo.ProductDetailsDesktop_noGap__qQjap > div:nth-child(2) > div > div.ProductDetailsDesktop_primaryDetails__6r9u9 > div.ProductDetailsDesktop_coreCtr__ZVN_b > div > div.CoreDetails_offerContainer__BNBZp > div.CoreDetails_offerItemCtr__ONued > div.CoreDetails_priceCtr__ZfUY9 > div > div > div.PriceOffer_savingPriceCtr__DRd7p > div.PriceOffer_priceSaving__ajbD4 > span") return tag.get_text(strip=True) if tag else None def parse_description(soup): """ Extract the product description from the page HTML. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str or None: The product description text if found, None otherwise """ tag = soup.select_one("#catalog-page-container > div > div:nth-child(2) > div:nth-child(1) > section > div > div > div.OverviewTab_container__2ewCs > div > div.OverviewTab_overviewDescriptionCtr__d5ELj") # Adjust selector as needed return tag.get_text(strip=True) if tag else None ``` One of the things that makes this script easy to understand and maintain is how it handles parsing—that is, pulling specific bits of information from each product page. Instead of writing one long block of code to grab everything at once, the script is neatly divided into small, focused functions. Each one has a clear name and a single job. Each of these functions uses a CSS selector to locate exactly the part of the page it needs. This setup is based on a solid understanding of how Noon’s website is structured, so the code can move quickly and accurately through each page. All of these functions follow the same simple pattern:They search for the element, pull out the text if it exists, and if not, they just return None. This keeps things clean and predictable—and it avoids unnecessary errors if a certain detail isn’t available on the page. This approach has two major benefits: 1. It’s easy to update. If Noon changes the layout of their site, you only need to tweak the one function related to that change—without having to rewrite your whole script. 2. It’s resilient. If a product page is missing some info, the scraper doesn’t crash. That field is just left empty, and the script keeps going, smoothly collecting the rest of the data. In short, this modular design makes the scraper more reliable and much easier to maintain, especially when working with a large number of products or handling sites that occasionally change their structure. ``` def parse_product_specifications(soup): """ Extract product specifications from the page HTML and format as JSON. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML Returns: str: A JSON string containing the product specifications as key-value pairs The function: 1. Locates the specifications table on the product page 2. Extracts each specification row (key-value pair) 3. Builds a dictionary of specifications 4. Converts the dictionary to a JSON string with indentation If no specifications are found or an error occurs, an empty JSON object is returned. All exceptions are logged for troubleshooting. """ try: specifications_section = soup.select( "#catalog-page-container > div > div:nth-child(2) > div:nth-child(1) > section > div > div > div.SpecificationsTab_container__uBaMs > div > div > table > tbody > tr" ) if not specifications_section: logging.warning("No product specifications found in the given HTML.") return json.dumps({}) specifications = {} for tr in specifications_section: try: tds = tr.find_all("td") if len(tds) >= 2: key = tds[0].get_text(strip=True) value = tds[1].get_text(strip=True) specifications[key] = value else: logging.warning("Missing elements in while parsing specifications.") except Exception as e: logging.error(f"Error processing a in specifications: {e}") continue return json.dumps(specifications, indent=4) except Exception as e: logging.error(f"Error parsing product specifications: {e}") return json.dumps({}) ``` The parse\_product\_specifications() function is a little different from the other parsing functions in the script—and that’s because it deals with more complex, detailed data. Instead of pulling out just one piece of text, this function goes through a full table of product specifications. These are the technical details about a product—like size, material, scent family, and more. Each row in the table usually has a feature name (like “Brand” or “Weight”) and a value that goes with it. The function loops through each row and builds a dictionary, where each feature becomes a key, and its value is stored alongside it. This creates a structured collection of all the technical details for that product. Once that dictionary is complete, it’s turned into a JSON string—a flexible format that works really well for storing this kind of semi-structured data. The script then saves this JSON string in the database. This method has two big advantages: - It preserves the structure of the information, so we can easily analyze or display it later. - It’s tolerant to missing or messy data. If one part of the specs table has an issue, the rest of the function still works just fine—the scraper won’t crash. Because different products often have different sets of specifications, this flexible approach is a perfect fit. It handles all that variation smoothly and ensures we don’t lose valuable data, even when things aren’t perfectly formatted. ### Core Scraping Logic: Browser Automation and Extraction ``` async def scrape_product_page(url, category, user_agents): """ Scrape a single product page to extract all product data. Args: url (str): URL of the product page to scrape category (str): Category of the product (women, men, unisex) user_agents (list): List of user agent strings to choose from Returns: tuple or None: A tuple containing all extracted product data if successful, None if an error occurs The function: 1. Selects a random user agent for the browser 2. Launches a browser with the chosen user agent 3. Navigates to the product URL 4. Extracts all product information using the parsing functions 5. Returns the compiled data as a tuple If any error occurs during scraping, it is logged and None is returned. """ try: ua = get_random_user_agent(user_agents) async with async_playwright() as p: browser = await p.chromium.launch(headless=False) context = await browser.new_context(user_agent=ua) page = await context.new_page() await page.goto(url, timeout=60000) await page.wait_for_timeout(3000) content = await page.content() soup = BeautifulSoup(content, "html.parser") name = parse_name(soup) rating = parse_rating(soup) total_rating = parse_no_rating(soup) sale_price=parse_saleprice(soup) price = parse_price(soup) discount=parse_discount(soup) description = parse_description(soup) specifications=parse_product_specifications(soup) await browser.close() return (url, category, name, rating, total_rating, sale_price, price, discount, description, specifications) except Exception as e: logging.error(f"Failed scraping URL: {url} - {str(e)}") return None ``` The scrape\_product\_page() function is really the core of the entire scraping process—it’s where everything comes together. The function begins by launching a browser using Playwright. Before visiting the page, it picks a random user agent from our list. This small trick helps the scraper act more like a real user and reduces the chances of being blocked by the website. Once the browser is up, it navigates to the product’s URL and waits for the page to fully load. This includes all the content that’s generated by JavaScript—something many modern websites, like Noon, rely on heavily. Waiting ensures we don’t miss any important details. After the page is ready, the HTML is captured and sent to BeautifulSoup, which parses it and passes it through the parsing functions we created earlier. These functions pull out all the key product information: name, price, ratings, description, and technical specifications. When all the data is collected, the browser is closed properly to free up system memory—especially important when scraping many products in a row. This function is powerful because it combines the browser automation of Playwright (for handling dynamic content) with the simplicity of BeautifulSoup (for extracting clean data). On top of that, it includes error handling, so even if one product fails to load or parse correctly, the script keeps going smoothly. ### Orchestration: The Main Function and Program Flow ``` async def main(): """ Main function to orchestrate the product data scraping process. The function: 1. Loads user agents for browser fingerprint randomization 2. Connects to the database and ensures required columns exist 3. Fetches all unscraped product URLs 4. Scrapes data from each product URL 5. Saves successful results to the database 6. Marks URLs as scraped to avoid reprocessing Progress is logged throughout the process, with success and error messages to aid in monitoring and troubleshooting. """ user_agents = load_user_agents() conn = connect_db() cursor = conn.cursor() ensure_scraped_column(cursor) conn.commit() urls = fetch_unscraped_urls(cursor) logging.info(f"Total unscraped URLs: {len(urls)}") for url, category in urls: logging.info(f"Scraping: {url}") product_data = await scrape_product_page(url, category, user_agents) if product_data: save_product_data(cursor, product_data) mark_url_as_scraped(cursor, url) conn.commit() logging.info(f"✅ Saved data for: {url}") else: logging.warning(f"⚠️ Skipped URL due to error: {url}") conn.close() logging.info("🎉 Done scraping all products.") if __name__ == "__main__": # Entry point: Run the main async function asyncio.run(main()) ``` The main() function is the central hub of the entire scraping process. It’s where everything gets started and tied together. It begins by loading the list of user agents and connecting to the database. This setup step makes sure the database is ready with the right tables and structure to store all the product details we’re about to collect. Then, it pulls a list of product URLs that haven’t been scraped yet, so we only work on fresh data. This avoids wasting time on duplicates or re-scraping old products. From there, the function goes through each URL one by one. For each product page, it: 1. Calls the scrape\_product\_page() function to collect the data 2. Saves the product info to the database 3. Marks that URL as scraped, so it won’t be touched again next time This step-by-step process may sound simple, but it’s very intentional. By processing the URLs one at a time, the script avoids overloading the website with too many requests. It’s a respectful and responsible approach to scraping. Throughout the run, the function uses logging messages to show what’s happening in real time. These logs include emojis—like ✅ checkmarks for successful scrapes and ⚠️ warnings when something goes wrong. This makes the updates more readable and even a little fun, which is especially nice when you're scraping hundreds of pages and want to keep track without reading dry messages. ## Conclusion Together, these two scripts form a smart and efficient scraping system—one script collects the product links, and the other dives into each link to gather detailed product information. What makes this setup strong isn’t just the tools it uses, but the way it’s built. By rotating user agents, the script avoids getting flagged or blocked by the website. It’s also designed with error handling, so if one product page fails, the rest of the process keeps running without a hitch. All the data gets stored neatly in a database, and the script keeps track of what’s already been scraped—saving time and avoiding duplicate work. Another great aspect is how it’s designed to respect the website. It doesn’t overload the servers by making too many requests at once, which makes it more sustainable for long-term use. In the end, this system collects rich, well-structured data that can be used for all kinds of insights—like analyzing market trends, tracking pricing strategies, or simply exploring what kinds of fragrance products are being offered. Whether you’re a beginner exploring web scraping or someone working on a more advanced project, this setup gives you a strong foundation for collecting and working with real-world data. ## FAQ SECTION ### 1) Is it legal to scrape data from Noon? Web scraping legality depends on how the data is collected and used. Publicly available product information can generally be collected for research or competitive analysis, but you should always review the website’s terms of service and robots.txt file. Avoid scraping personal data and follow ethical scraping practices. ### 2) Why use Playwright instead of BeautifulSoup alone? BeautifulSoup is great for parsing static HTML. However, Noon loads product listings dynamically using JavaScript. Playwright automates a real browser, allowing the page to fully render before extracting data, making it ideal for dynamic ecommerce websites. ### 3) What data can be extracted from Noon’s fragrance store? You can extract product names, brand names, prices, ratings, number of reviews, product descriptions, availability status, and category information. This data can be used for price tracking, brand analysis, and market research. ### 4) Why store scraped URLs in a SQLite database? Using SQLite helps organize and preserve scraped data efficiently. It prevents data loss, supports filtering by category, and makes it easier to scale the scraping workflow for future analysis without re-scraping the same pages. ### 5) How can scraped fragrance data be used for business insights? Fragrance data can reveal pricing trends, brand popularity, discount patterns, rating distribution, and assortment gaps. Businesses can use these insights for competitive intelligence, pricing optimization, and product strategy decisions. AUTHOR I’m Shahana, a Data Engineer at Datahut, where I design and develop scalable data pipelines that transform messy, complex web content into clean, structured datasets—especially for e-commerce, pricing intelligence, and product tracking. In this blog, I shared a step-by-step walkthrough of a real-world scraping project where we collected detailed product data from Noon’s fragrance category. Using tools like Playwright, BeautifulSoup, and SQLite, we built a robust solution that handles dynamic content, rotates user agents to avoid detection, and stores the data in a clean, organized format for analysis. At Datahut, we focus on building practical, responsible web scraping systems that are ready for real-world challenges—whether it’s managing large-scale data collection or ensuring the scraper can recover gracefully from interruptions. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. ### The Upsell You're Overlooking: Micro Data Products Hidden Inside Your Existing Accounts URL: https://www.blog.datahut.co/post/micro-data-products-hidden-inside-your-existing-accounts/ Last updated: 2026-09-07T09:43:23.000Z Most software services firms are hunting for the next big transformation program. However, they are overlooking an opportunity already [sitting within the accounts you have](https://cloud.google.com/discover/what-is-data-as-a-product?ref=blog.datahut.co). ## The Middle Ground Nobody Proposals Account expansion has a familiar playbook: land a project, deliver value, propose the next phase. The problem is that "next phase" almost always means another large program, which means executive buy-in, long sales cycles, and a budget that's never guaranteed. That leaves a [massive middle ground unexplored](https://www.thoughtworks.com/en-in/radar/techniques/data-product-thinking?ref=blog.datahut.co). Inside every account, there are dozens of smaller, high-value problems your team can see clearly but never proposes. Not because they aren't real problems. Because they were never economically viable to solve. Too narrow in scope as a project. Too costly to build as a system. Too hard to justify in an ROI spreadsheet when only five or six people feel the pain. [Foundation models changed that math.](https://www.thinslices.com/insights/foundation-models-redefining-lean-prototyping-for-tech-startups?ref=blog.datahut.co) ## What Changed [Before GenAI, even a "small" workflow meant weeks of engineering](https://www.blog.datahut.co/post/why-ai-web-scraping-fails-at-enterprise-scale/): data mapping, extraction, rules, edge-case handling, and a UI. The work wasn't conceptually hard; it was just expensive relative to the size of the problem. So niche use cases died before they became projects. Now, connecting a model to data already in an account and wrapping it in a lightweight workflow takes days, not months. Problems that could never be approved as standalone projects can be shipped quickly and cheaply. That's what makes micro data products commercially viable, not a new concept, but a new cost structure. ## What a Micro Data Product Actually Is A microdata product solves one specific problem extremely well. It has four parts: - An existing data source - something the client already has - A logic layer - model, rules, classification, extraction, matching - A simple workflow - review, validate, route, alert. - A clear output — exception list, weekly summary, dashboard tile, alert No grand platform. No multi-quarter roadmap. One job: turning a messy data stream into [always-on intelligence](https://www.blog.datahut.co/post/category-price-index/)[ for a specific team](https://www.blog.datahut.co/post/category-price-index/). It might not be a million-dollar project, but it could be worth between $30K-$50 per year. If you find similar [micro data opportunities ](https://atlan.com/know/what-is-a-data-product/?ref=blog.datahut.co)inside the accounts, it can accumulate to a million dollars fast. ![What a Micro Data Product Actually Is ](https://www.blog.datahut.co/content/images/2026/07/img-230.png-1.webp) A Concrete Example [We’re working with ](https://www.datahut.co/?ref=blog.datahut.co)an IT services company whose customer, a retailer, needed to track ingredient changes across thousands of SKUs - detecting when product formulations shifted as new items launched. In the old world, this would have been rejected. Too narrow. No clear home for it in the project portfolio. But the signal was already in their existing product dataset. The team connected it to a foundation model API, added lightweight validation logic, and shipped a simple workflow: when a meaningful formulation change is detected, the relevant team gets an alert with a plain-language explanation - what changed, which SKU, why it matters. Total build time: under three weeks. Ongoing cost: minimal. Value: a previously invisible risk, now impossible to ignore. That's the pattern. Take a change that's hard to notice and make it impossible to miss. ## Why Niche Problems Are the Best Upsell Surface The problems that never get funded are often the ones worth solving. A problem affecting five or six people won't survive a normal approval process - it can't justify a broad user base, a big business case, or procurement involvement. But those same people might be spending hours each week on manual workarounds, missing signals that cost real money, or making decisions on stale information. Microdata products sidestep the approval problem entirely. Small enough to be approved at the team or department level. Fast enough to deliver value before the budget conversation stalls. Cheap enough to be a line item, not a program. And once one is live, something important happens: other teams start asking for theirs. ## The Multiplier Effect A micro product in production — not a demo, but an always-on signal — changes how a team thinks about what's possible. The questions start multiplying fast: "Can you detect this change too?" "Flag this risk for my category." "Surface this trend weekly." "Tell us what moved and why." Merchandising, category management, supply chain, finance ops, compliance — each team has its own version of the same underlying need: their data, filtered into decisions they can act on today. Each of those requests is a small, scoped engagement. Each one expands your footprint. Instead of waiting for a large platform program to be approved, you're shipping small products frequently, embedding yourself in day-to-day decisions across multiple teams. ## What These Look Like in Practice Three patterns show up repeatedly across industries: 1. Change detection — ingredient or formulation changes across SKUs; pack size changes that break comparisons; supplier or spec sheet updates; policy changes affecting compliance. Output: a weekly "meaningful changes" feed with alerts for anything critical. 2. Anomaly explanation — sales spikes or drops with likely drivers; sudden cost changes in procurement lines; unusual refund or return patterns; outlier equipment downtime. Output: an exception list with plain-language explanations and supporting evidence. 3. Data quality monitoring — fields going missing; attribute consistency breaking across sources; duplicate entities creeping in; low-confidence matches increasing over time. Output: a drift dashboard and a daily review queue before problems compound silently. [None of these requires a full platform build. ](https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-missing-data-link-five-practical-lessons-to-scale-your-data-products?ref=blog.datahut.co)They require a dependable data stream and a thin product layer that turns it into decisions. ## The Commercial Model That Follows Service firms that build this motion well don't just earn project revenue—they earn recurring revenue embedded in client accounts. Because once these micro products exist, they need maintenance: alert tuning, threshold adjustments, model drift checks, edge case refinement, and expansion into adjacent teams. That's not scope creep. That's a natural engagement model — one where the relationship shifts from transactional projects to continuous value delivery. ## The Real Advantage The firms that win here will combine two things: deep account context and delivery credibility, plus a data layer that supports multiple micro products without repeated reinvention. Put those together, and account expansion stops being about selling phases. It becomes about shipping signals — small, frequent, and tied to decisions teams make every day. Your next upsell is probably already sitting in a dataset your client already owns. The question is whether you're looking for it. ## FAQ SECTION ### 1\. What are micro data products? Micro data products are lightweight, focused analytics solutions built on top of data a client already owns. Instead of launching large, standalone data platforms, they solve specific business problems using a thin product layer powered by models, automation, and workflow logic. ### 2\. How are micro data products different from traditional data projects? Traditional data projects are infrastructure-heavy, time-consuming, and expensive. Micro data products, on the other hand, leverage existing systems and datasets, require minimal new infrastructure, and can be built and deployed in weeks instead of months. ### 3\. Why are micro data products considered an upsell opportunity? They unlock new revenue within existing accounts. Since the client’s data already exists, businesses can introduce targeted intelligence layers that solve niche but high-value problems—creating expansion revenue without needing new customer acquisition. ### 4\. What kind of problems can micro data products solve? They are ideal for use cases such as anomaly detection, pricing alerts, competitor monitoring, inventory triggers, risk signals, and performance diagnostics—especially problems too niche to justify a full standalone project. ### 5\. Do micro data products require new infrastructure or major system changes? No. Most micro data products integrate with a client’s existing data stack. They connect to current data sources, apply model-driven logic, and deliver actionable signals through dashboards, alerts, or workflow integrations. ### Stop Tracking Just Competitors: Build a Category Price Index Instead URL: https://www.blog.datahut.co/post/category-price-index/ Last updated: 2026-09-07T09:43:25.000Z Most companies ask us for [competitor pricing data](https://www.blog.datahut.co/post/how-competitor-data-transforms-category-management-pro-tips-best-practices/). Usually, they begin with a shortlist and say, “Track these five competitors.” We usually push back. Why those five? Why not the entire category? That question gets the same response every time: Why would we do that? Because [tracking competitors](https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/) only shows what a few companies did today. Tracking the whole category shows what the market is doing and helps you see if a change is just a one-time event or a real shift. This blog explains the difference. ## The Problem With Pricing in E-commerce This is a familiar scene. You log in on a Tuesday, check a key competitor, and see they’ve cut prices across their electronics aisle. Instantly, you’re forced into a decision: - Do you match them and protect conversion? - Do you hold and protect margin? - Do you undercut and try to take a share? But the real question is the one nobody can answer from a single price check: Is this a one-off move or the market shifting under your feet? Most teams default to gut[ feel](https://www.blog.datahut.co/post/6-signs-your-business-has-pricing-fatigue-and-how-to-fix-it/), a few screenshots, and a Slack thread. The problem is that pricing is noisy. A seller clears inventory. A marketplace runs a flash sale. A brand tests a promo. If you react to every alert, you end up reacting to randomness. Most retailers underestimate the discipline [dynamic pricing requires.](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-dos-and-donts-of-dynamic-pricing-in-retail?ref=blog.datahut.co) A category price index solves this problem. It combines scattered competitor changes into a single, clear market signal, giving you context rather than just notifications. ## What Is a Category Price Index? A category price index is a simple way to answer a hard question: what’s happening to pricing in my market, overall? Instead of watching individual SKUs, which can be noisy and easy to misread, you track a group of similar products and combine them into a single index number over time. This works like the Consumer Price Index (CPI): one product can be misleading, but a group shows the real trend. In practical terms, a category price index tells you: 1. Is this category getting more expensive or cheaper over time? 2. Are my prices moving with the market — or drifting out of position? You can define the category broadly, like “running shoes,” or more narrowly, such as “men’s trail running shoes, size 9 to 12.” Narrower definitions usually give a clearer signal, as long as your group is large enough to be representative. You can also create useful variants: - Competitor index: only your top competitors - Channel index: marketplaces vs D2C - Price-tier index: entry / mid / premium baskets No matter how you slice it, the idea is the same: stop treating pricing as a set of isolated SKUs and start measuring it as a market trend you can manage. ## How Category Price Indexes Are Built A category price index isn’t complicated. It’s just a disciplined way to turn messy SKU-level pricing into a clean signal you can track. Here’s the basic process: 1. Define the category (and the rules). Decide what belongs in the category and what doesn’t. Be explicit about filters like brand, specs, size range, condition (new vs refurbished), seller type (1P vs 3P), and whether you include out-of-stock items. 2. Build a representative basket. Choose enough products to cover the full range of the category, including entry-level, mid-tier, and premium, so that no single brand or price point dominates the results. 3. Collect prices consistently over time. Pull the same set of fields at a regular cadence (daily / weekly). This can come from web scraping, retailer feeds, or a third-party provider like Datahut. 4. Normalize to a baseline. Pick a reference period (e.g., January 2026 = 100), so changes are easy to interpret. If the index moves from 100 to 92, the category is \~8% cheaper than the baseline. 5. Weight where it matters. Not every SKU should count the same. You can use sales volume, review count, bestseller rank, or other signals to make sure the index reflects what customers actually buy, not just what is listed in the catalog. ![How Category Price Indexes Are Built ](https://www.blog.datahut.co/content/images/2026/07/img-285.png.webp) The result is simple: a single number and trend line that shows where the category is heading, without getting lost in the noise of individual SKU changes. ## Why It Matters for [E-commerce Sellers](https://www.blog.datahut.co/post/the-invisible-e-commerce-profit-killers-spot-the-flaws-your-standard-audits-fail-to-catch/) A category price index might sound like just another analytics tool. In reality, it changes how you make pricing decisions because it replaces isolated reactions with a broader market context. ### 1\. It Separates Signal From Noise SKU prices are naturally volatile. One seller clears inventory, a marketplace runs a 48-hour promo, or a brand tests a discount, and suddenly your dashboards light up. A category price index dampens one-off events and shows the market's underlying direction. If the category index drops 8% over six weeks while your prices stay flat, you’re not looking at “a competitor having a bad week.” You’re looking at a category-level shift you can’t ignore. ### 2\. It Reveals Seasonal and Structural Pricing Cycles Most categories move in cycles, but the pattern of these cycles changes every year. The index helps you see these changes. You can compare this year vs last year and answer questions like: - Are discounts starting earlier this season? - Is the Q4 dip deeper than normal? - Are prices recovering more slowly post-holiday? These patterns are easy to miss when you focus on individual SKUs. At the index level, they become clear. ### 3\. It Benchmarks Your Positioning Against the Market Instead of asking, “Am I cheaper than Competitor X on Product Y?” you can ask the more strategic question: Is my overall price positioning moving toward or away from the market? ![price positioning and market index](https://www.blog.datahut.co/content/images/2026/07/img-286.png.webp) This is especially important when your catalog is large. You can’t manually monitor thousands of SKUs, but you can track your index against the market index and spot changes early, whether you’re becoming too expensive or losing margin without noticing. ### 4\. It Lets You Go Granular With Price-Tier and Basket Trends One index gives you the direction of the category. But real decisions often live in the details. So you split the category into baskets and track each one: - Entry / mid/premium price tiers - Top brands vs long-tail brands - Key spec buckets (e.g., storage sizes, screen sizes, active ingredients) - Hero SKUs vs the rest ![Product distribution by price band ](https://www.blog.datahut.co/content/images/2026/07/img-287.png.webp) This approach helps you spot things the overall index might hide, such as “premium is holding price while entry is collapsing,” or “one spec bucket is being aggressively discounted.” ### 5\. It Powers Smarter Repricing Rules Most repricers are SKU-first: match the lowest price, stay within X% of Buy Box, follow competitor A. Indexes enable more advanced strategies. You can adjust your rules based on where the category is in its pricing cycle: - When the category index is falling, price more defensively to protect volume. - When it’s rising, pull back discounts and protect margin. You stop repricing in response to chaos [and start repricing based on the overall market situation.](https://www.bain.com/how-we-help/retailers-are-you-getting-the-full-value-of-your-dynamic-pricing-strategy?ref=blog.datahut.co) ### 6\. It Informs Procurement and Inventory Decisions [Pricing intelligence](https://www.blog.datahut.co/post/why-ai-web-scraping-fails-at-enterprise-scale/) shouldn’t stay trapped inside a pricing dashboard. When the category index trends up, it can signal that it's time to lock inventory earlier, before costs move. When it trends down, it can signal the need to negotiate harder, reduce buys, or shift mix. The index becomes a shared language for pricing, buying, and inventory planning, not just another report for the pricing team. ## Getting Started With Category Price Indexes You don’t need a data science team to start. You need a clean external data feed and a simple operating rhythm. 1. Start with 3–4 revenue-driving categories. Pick categories where pricing volatility directly impacts margins and conversion rates. 2. Define a basket of 20–30 representative SKUs per category. Don’t aim for “complete.” Aim for “representative.” Cover entry, mid, and premium tiers. 3. Get the data layer right. Use a provider like Datahut to collect prices consistently across your chosen sources and deliver a standardized dataset you can trust (not a fragile scraper you babysit). 4. Start by tracking weekly. Consistency is more important than perfection. A simple weekly index is better than a complex system that no one uses. 5. Share it cross‑functionally. The index is most valuable when pricing, merchandising, and buying teams reference the same signal. ## Why Now (and Why Datahut) Category price indexes only became practical — and urgent — in the last few years. ### Why now - Margin compression is structural. When costs, promos, and competitor behavior swing faster, you can’t manage pricing with occasional checks. [You need a market-level signal that updates predictably](https://www.bcg.com/publications/2024/overcoming-retail-complexity-with-ai-powered-pricing?ref=blog.datahut.co). - Pricing is no longer “a few competitors.” It’s a category-wide algorithmic game. Marketplaces, repricers, and promo engines move prices across hundreds of SKUs at once. If you only watch a shortlist, you miss the real trend. Also, keep in mind that regulators[ are paying attention.](https://www.reuters.com/legal/litigation/ftc-investigating-instacarts-ai-pricing-tool-source-says-2025-12-17/?ref=blog.datahut.co) - Information asymmetry compounds. Teams with a reliable category index know when the market is rising, falling, or stabilizing — and they tune discounts, promo depth, and price positioning accordingly. Everyone else reacts late (or overreacts) to noise. ### Why Datahut Building a category index is straightforward. Maintaining the data pipeline behind it is not. Retailers rarely struggle because they can’t calculate an index. They struggle because they don’t have consistent, structured, decision-ready price data across the sources that matter — with stable identifiers, clean normalization, and enough coverage to keep the basket representative. Datahut makes category indexing usable by delivering: - Standardized price feeds across retailers/marketplaces (with timestamps, seller type, availability, discount flags) - Clean, stable SKU mapping so your basket doesn’t break every week - Category and tier coverage that keeps the index representative (entry / mid / premium) - BI-ready datasets so you can compute and track indexes inside the tools you already use You don’t have to rebuild your entire data stack overnight. Add the category-index layer first and start making pricing decisions with market context rather than isolated competitor alerts. ## The Bigger Picture Category price indexes represent a shift in how ecommerce teams think about pricing — from a product-by-product activity to a market-level capability. In a world where algorithms can reprice thousands of products per day, advantage doesn’t come from faster reactions. It comes from better interpretation. Indexes give you the context to act with intent instead of reflex. That’s not just better pricing. That’s a durable business advantage. About the author I’m Tony Paul, founder of [Datahut](https://datahut.co/?ref=blog.datahut.co). I’ve spent 15+ years in the web scraping industry helping retailers and enterprises build reliable external data pipelines. My belief is simple: most teams waste time reinventing the commodity layer (scrapers, proxies, maintenance) when they should focus on the value layer (insights, decisions, execution). If you’re exploring web scraping, pricing intelligence, retail data, or AI-ready external datasets, reach out — you can connect with me on LinkedIn: [Tony Paul.](https://www.linkedin.com/in/tonypauldh/?ref=blog.datahut.co) ## FAQ’s ### 1) What is a category price index in e-commerce? A category price index tracks the overall price movement of a product category over time using a representative basket of SKUs. It turns thousands of price changes into a single trend line so you can see whether the market is rising, falling, or stabilizing. ### 2) How is a category price index different from competitor price tracking? Competitor tracking shows what a few sellers did. A category price index shows what the market is doing. That difference matters because a single competitor drop can be noise, while an index reveals whether pricing pressure is broad-based and persistent. ### 3) Can I build a category price index using web scraping? Yes. Web scraping is one of the most common ways to collect consistent category pricing across retailers and marketplaces — as long as you maintain a reliable pipeline (stable identifiers, clean normalization, consistent cadence, and monitoring). This is exactly where teams often prefer a managed provider like Datahut instead of maintaining fragile scrapers internally. ### 4) What does Datahut deliver for category price indexing? Datahut typically delivers a standardized, BI-ready dataset that includes prices over time across your chosen sources, plus the fields you need to keep an index stable: timestamps, seller type (1P/3P), availability, discount flags, and clean SKU mapping so your basket doesn’t break every week. ### Data Readiness: Fix Your Data Before You Invest In AI URL: https://www.blog.datahut.co/post/data-readiness-fix-your-data-before-you-invest-in-ai/ Last updated: 2026-09-07T09:43:27.000Z Retailers are investing heavily in AI initiatives; some are hiring dedicated Chief AI officers or VP AI roles to lead them. Some are bringing in external consultants to help with it, but most of those initiatives won’t hit their original goal because they are having a [data maturity problem](https://sapinsider.org/map/ai-readiness-is-defined-by-data-readiness/?ref=blog.datahut.co). What we’re seeing now is that [Pilots perform well, but in production, it goes haywire.](https://www.blog.datahut.co/post/why-ai-web-scraping-fails-at-enterprise-scale/) Models perform well in pilot but fail at scale. Executives who were once [cheerleaders of AI are now turning ](https://www.ibm.com/think/insights/data-quality-issues?ref=blog.datahut.co)into skeptics. The main issue is not budget or talent. Most often, it is a lack of [data readiness](https://insights.fusemachines.com/data-readiness-in-retail-the-first-step-to-winning-the-ai-race/?ref=blog.datahut.co). ## What data Immaturity looks like in practice When [AI programs ](https://www.blog.datahut.co/post/y-combinator-2025-how-ai-is-reshaping-startups-and-markets/)or data initiatives underperform within a company, the postmortems often suggest model limitations or friction from the team members to adopt the solution. However, the deeper issue is the misalignment between ambition and infrastructure. Common signs of immaturity. - Inaccurate data costs organizations an average of [$12.9 million per year, according to Gartner](https://firsteigen.com/blog/10-common-data-quality-issues-and-how-to-solve-them/?ref=blog.datahut.co). - Your AI models are fed obsolete data . In high-velocity retail, a price recommendation based on 24-hour-old data is often obsolete, and by the time a decision is made, the opportunity is long gone. - Pricing engines are trained on inconsistent data. - Automation initiatives are deployed without clear domain ownership. - People talk to each other across departments, but systems and data do not. - Teams does not have access to data from other departments that you’re dependent on for making decisions. In such environments, AI initiatives will fail no matter how much investment you make. The only exception is for totally isolated deployments of a single function that doesn't have any internal data dependencies. One of our customers said, “ The distance between our ambition and infrastructure is the distance between Hawaii and California, which is about 4000 kilometers. For this to work, both of them should be together in California. ## The Illusion of a Big Leap We’re a web scraping company, and our data powers many retail initiatives. While talking to retail leaders, one thing is clear. They want to move from their current fragmented reporting stage to a fully autonomous AI-driven stage. I understand this goal. Margin pressure is high, tariffs are causing new problems, and AI is changing how consumers behave. But there are no shortcuts to data maturity. Companies must move through five stages to get there. These stages happen in order. Each one builds key abilities in data quality, speed, access and integration, timeliness, governance, ownership, and alignment with business goals. ![6 dimensions of data maturity ](https://www.blog.datahut.co/content/images/2026/07/img-305.png.webp) ## Data Readiness Maturity Model: How to read it Look at the data readiness maturity model below. It shows two variables. The horizontal axis tracks how mature your data capabilities are, moving from low to high as systems, governance, integration, and ownership get better. The vertical axis represents business readiness, the degree to which data actively supports operational and strategic decisions in retail. ![Data capability maturity model](https://www.blog.datahut.co/content/images/2026/07/img-306.png.webp) ### Fragmented This is the first stage, where data readiness and capabilities are at their lowest. Data lives in disconnected systems, such as spreadsheets, Product information management software, and Shopify. It does not communicate at all, and there is no ownership, no accountability, and the decision-making is mostly reactive without knowing the full story, often on false positives or false negatives. At this stage, the organization is in survival mode. The Latency is high (weeks/months). ### Controlled In the controlled stage, maturity begins to improve . Usually, a central team establishes the standards, schemas, and data reporting structure. This definitely improves data readiness; however, the business readiness remains moderate because the reports live in isolation. Latency is moderate (days to overnight). ### Integrated This stage reflects further progress, platforms are standardised, and shared data products and dashboards replace isolated reports. Business functions operate with cross functional visibility and improved coordination. Latency is low (days/ hours) ### Domain owned. In the domain owned stage, the business teams assume accountability for their data assets, enabling automation and predictive decision-making. The data becomes embedded in production workflows. Maturity changes from merely reporting to actually using it. Near-real time or Zero-latency. AI acts on events as they happen (e.g., a competitor price drop triggers an immediate automated response). Data as Infrastructure Finally, at the far right of the model sits Data as Infrastructure. Here, data maturity and business readiness are both high. Internal systems are seamlessly connected, enriched with external data, and capable of supporting autonomous, AI-driven decisions at scale. At this stage, the data is ready enough that you can move from AI pilots to production without many major bottlenecks. The model's trajectory is deliberately sequential. Each stage builds structural strength across quality, latency, integration, timeliness, governance, and alignment. ## The Cost of Skipping the Foundation Organizations that try to use AI without improving their data infrastructure face predictable risks. - [Capital and resources wasted](https://www.blog.datahut.co/post/build-vs-buy-web-scraping-in-2026-the-definitive-guide-for-data-teams/) - Wrong decisions - More margin pressure And at the end, AI initiatives become an expensive experiment that yields negative results. More importantly, you don’t stand a chance against competitors who are doing it right. ## Who Should Care Data readiness is not a technology problem; it is fundamentally a P&L issue and should be seen as one. Pricing Leaders: Depend on trusted external data to protect margins, and even minor inconsistencies can compound into a material financial impact. Strategy and executive teams: Capital allocation is based on enterprise-wide visibility, and if that visibility depends on immature data foundations, the allocation will be inappropriate. [Merchandising teams](https://www.blog.datahut.co/post/want-to-fix-your-unit-economics-do-what-nestl%C3%A9-did-start-saying-no-to-more-skus/) need complete, timely product attributes and demand signals to optimize assortments. Without integration, missed trends and excess inventory become problems that affects cashflow. At advanced stages of readiness, these functions operate with synchronized, trusted data flows. At early stages, they operate with friction and blind spots. The difference is not incremental. It is systemic. ### Why now (and why us) The rush to become data-ready is intensifying for three reasons. - Margin compression is structural. Retail volatility, from tariffs to supply shocks to shifting consumer demand, requires faster reaction cycles. Delayed insight translates directly into lost margin. - AI‑led discovery is reshaping how customers find and evaluate products. If product and pricing data are inconsistent or poorly structured, brands risk invisibility in algorithmic retail. - Information asymmetry compounds. Retailers that reach higher maturity stages accelerate decision cycles and reduce operational friction. Those that remain fragmented experience increasing drag. In such an environment, incremental improvement is insufficient. Structural readiness determines strategic positioning. This is precisely where [Datahut](https://www.datahut.co/?ref=blog.datahut.co) operates. [Retailers](https://www.blog.datahut.co/post/data-for-fashion-retailers-the-four-problems/) do not always fail because they lack internal data. They struggle because they lack structured, timely, and decision‑ready external market signals , competitive pricing shifts, assortment gaps, stock‑out patterns, demand signals across marketplaces, and product attribute inconsistencies that weaken AI visibility. [Datahut helps close that gap.](https://www.datahut.co/?ref=blog.datahut.co#contact) By delivering structured, high‑quality external data feeds, we strengthen the six core dimensions of readiness: improving data quality, enhancing integration, accelerating timeliness, reinforcing governance through standardised datasets, and aligning data directly to pricing, merchandising, and strategy use cases. We do not ask organizations to rebuild their entire infrastructure overnight. Instead, we provide the market intelligence layer that enables faster margin protection, [sharper assortment decisions](https://www.blog.datahut.co/post/product-assortment-for-retailers-in-laymans-terms/), and improved AI discoverability, without requiring a full internal platform overhaul. In other words, while internal maturity must progress sequentially, external intelligence can accelerate the journey. Retailers that combine disciplined internal data evolution with reliable external market signals move from reaction to anticipation. Those who do not risk operating with incomplete visibility in an increasingly algorithmic market. ### About the author: I’m Tony Paul, founder of Datahut, with over 15 years of experience working in the web scraping industry. We provide Data as a Service" (DaaS) to global retailers and enterprises My core belief is that most organizations waste resources "re-inventing the wheel" when they should be focusing on the Value Layer (insights and decision-making) rather than the Commodity Layer (maintaining scrapers and infrastructure). If you’re looking for help with web scraping, pricing, retail data, or anything in between - hit me up. You can connect with me on LinkedIn here: [Tony Paul](https://www.linkedin.com/feed/?ref=blog.datahut.co) ## Frequently Asked Questions ### 1\. Can’t we improve data readiness while simultaneously investing in AI? Yes — but only if AI initiatives are sequenced appropriately. When foundational gaps in quality, integration, and ownership are ignored, AI projects tend to stall or require costly rework; when readiness improvements and AI deployment are aligned deliberately, each reinforces the other. ### 2\. How do we know which maturity stage we are in? The clearest signal is operational behavior, not technology inventory. If teams debate whose numbers are correct, rely heavily on manual reconciliation, or struggle to embed insights into workflows, the organization is likely operating in Fragmented or Controlled stages rather than Integrated or Domain‑Owned. Or you can talk to [us](https://datahut.co/?ref=blog.datahut.co) and figure out. ### 3\. What is the first practical step toward becoming AI‑ready? Start by identifying one high‑impact business domain — pricing, merchandising, or supply chain — and assess it across the six readiness dimensions: data quality, integration, timeliness, governance, and strategic alignment. Strengthening one domain structurally creates a repeatable blueprint for scaling readiness across the enterprise. ### Data for Fashion Retailers: The Four Problems Nobody Talks About (But Everyone Feels) URL: https://www.blog.datahut.co/post/data-for-fashion-retailers-the-four-problems/ Last updated: 2026-09-07T09:43:29.000Z ## The Impact of Tariffs on the Fashion Industry As fashion retailers look toward 2026, they are adjusting to a fundamentally new reality. The US tariffs have forced brands and their suppliers to adapt quickly. Major brands like Nike, [Hermès](https://www.businessoffashion.com/articles/global-markets/the-state-of-fashion-2026-report-tariffs-trade-us-market/?ref=blog.datahut.co), and Ralph Lauren have already indicated or implemented price increases. Consumers are also changing their spending habits. They are reprioritizing what to buy and where to spend their money. According to a Fashion–McKinsey [State of Fashion Executive Survey](https://www.mckinsey.com/industries/retail/our-insights/state-of-fashion?ref=blog.datahut.co), 46 percent of fashion executives expect conditions to worsen in 2026, compared to 39 percent in last year's survey. Source: [McKinsey](https://www.mckinsey.com/~/media/mckinsey/industries/retail/our%20insights/state%20of%20fashion/2026/stateoffashion-ex1.svgz?cq=50&cpy=Center&ref=blog.datahut.co) Let’s simplify this. Most [fashion](https://www.icaew.com/library/industry-guides/fashion-industry?ref=blog.datahut.co) retailers face more margin pressure than ever before. They struggle against competition, primarily due to a lack of visibility into the market. When a quarter doesn't perform well, the reasons are often complex. Marketing may miss its goals, demand could drop, or customers might become more price-sensitive. Most people just read the quarterly reports and move on. We have worked with many fashion retailers, and the same four problems consistently emerge: 1. They leak revenue without realizing it. 2. Margins are getting squeezed, and the situation is worsening each quarter. 3. [They fail to capture demand](https://www.blog.datahut.co/post/find-ecommerce-competitors/) when competitors run out of stock. 4. They either understock or drown in excess inventory. None of these issues happen overnight. They develop slowly through small decisions made without a full market context. And that’s where better data changes everything. ## 1\. Revenue Leakage You Cannot See Imagine this scenario. A competitor drops the price on a best-selling SKU by 12%. You don’t notice for five days. During those five days, conversion shifts. Not dramatically, but just enough. That’s revenue leakage. It usually happens when: - Pricing updates lag behind competitors. - Promotions aren’t aligned with market timing. - Premium products appear overpriced next to aggressive discounters. No one makes a bad decision on purpose. The problem lies in delayed awareness. A 2% conversion shift on a hero SKU portfolio doesn’t trigger alarms. But over a season, across categories, and across regions, it compounds into millions. Because it happens gradually, it rarely receives executive scrutiny. [Pricing](https://www.blog.datahut.co/post/h-m-s-pricing-strategy-detailed-analysis-data/) teams often rely on internal sales signals to decide when something is “wrong.” By the time internal data reacts, the market has already moved. They check the conversion dashboard, and if they see a decline, they manually check competitors to find out what happened. Revenue leakage isn’t loud. It’s incremental. And incremental losses compound quickly in fashion. ## 2\. Compression from Blind Discounting When sales slow, the default reaction in fashion is simple: discount and clear as much inventory as possible. However, discounting without understanding the market is costly over time and puts additional pressure on margins. Margin compression occurs when: - You discount while competitors hold prices steady. - You match competitor markdowns without understanding their strategy. - You underestimate where you could maintain premium positioning. Without outside benchmarks, pricing decisions become defensive. The better question isn’t “Should we discount?” It’s “Where are we actually misaligned with the market?” Data provides pricing teams with confidence. It’s not about discounting more; it’s about discounting smarter or sometimes not discounting at all. ## 3\. Assortment Gaps That Cost Market Share This is one of the most overlooked growth opportunities in fashion. When a competitor sells out of a popular product, demand doesn’t vanish. It shifts to other retailers. The real question is: are you ready to capture that demand? If you don’t know they are out of stock, you can’t react. You don’t: - Increase bids on substitute products. - Highlight alternatives on your site. - Adjust pricing strategically. - Push the right SKUs in email or paid campaigns. That’s a missed opportunity. Retailers who track competitor inventory by size and color can turn stock-outs into growth opportunities. The goal isn’t to game the system but to stay aware of the market as it changes. ## 4\. Inventory Imbalance - Overstock vs Understock Inventory mistakes are costly. If you understock, you lose revenue. If you overstock, you face markdown pressure. Both problems usually arise because teams lack enough context beyond their own sales data. You might ask: - Are we underrepresented in a growing price band? - Are competitors expanding faster in certain silhouettes? - Are we over-invested in styles that are cooling off? Internal forecasting shows what happened in the past. External benchmarking reveals where the market is heading. Balanced inventory requires both. ## Who Actually Uses This Data? This isn’t just for “the data team.” In most fashion organizations, four teams benefit immediately. ### 1\. Pricing & Revenue Teams Pricing and revenue teams operate under constant pressure from two directions: revenue leakage on one side and margin compression on the other. To handle this pressure, they need real-time competitor pricing, historical discount data, visibility into market promotions, and cross-geography comparisons. Without these inputs, even experienced pricing leaders make consequential decisions in the dark. They react to moves that have already happened rather than anticipating the ones about to happen. That's where [competitor pricing data](https://www.blog.datahut.co/post/how-to-leverage-web-scraping-to-create-a-competitor-price-monitoring-strategy/) changes the equation. With it, pricing teams can shift from reactive firefighting to proactive positioning. They can protect margins with confidence, identify when premium pricing will hold, and avoid unnecessary markdown spirals that erode brand equity and train customers to wait for discounts. For these teams, this kind of intelligence isn't a reporting exercise or a quarterly review artifact. It's decision fuel, consumed daily, and its absence is felt immediately in the numbers. ### 2\. Merchandising & Assortment Teams Merchandising teams are caught between two costly extremes: too much inventory in the wrong categories and too little in the right ones. Solving for this requires more than internal sell-through data; it requires a clear view of the competitive landscape. That means category depth comparisons, price band distribution mapping, attribute benchmarking across fabric, fit, and sustainability, and visibility into how competitors are rotating new arrivals. With that intelligence, the nature of merchandising decisions changes. Teams can spot assortment gaps before they become lost sales, avoid overbuying into quietly declining segments, and make smarter seasonal bets grounded in where demand is actually moving. None of this replaces the instinctive experience of seasoned merchandisers. It sharpens it. ### 3\. [eCommerce](https://www.blog.datahut.co/post/ecommerce-competitor-analysis/) & Growth Teams Growth teams face a deceptively simple problem: demand that should be theirs is quietly going elsewhere. When a competitor goes out of stock, there's a window—sometimes hours, sometimes days—where their customers are actively looking for alternatives. Capturing that demand window requires knowing it's open in the first place. That means having real-time visibility into competitor stock levels, variant availability, and sell-through velocity signals. With that intelligence, growth teams can act before the moment passes. Paid media can be redirected instantly toward high-intent audiences, substitute SKUs can be promoted to intercept redirected demand, and merchandising can be adjusted in real time to put the right product in front of the right customer. For growth teams, competitor inventory data translates directly into incremental revenue. ### 4\. Strategy, Data & AI Teams Data science and analytics teams typically focus on the long game. They build forecasting models, refine pricing algorithms, and find structural optimizations that compound over time. However, even the most sophisticated models have a ceiling when they are trained exclusively on internal data. They capture what the business has done; they can't see what the market is doing. ![Inventory trap](https://www.blog.datahut.co/content/images/2026/07/img-439.png.webp) That changes when structured competitive datasets enter the picture. Historical pricing data, inventory trend patterns, assortment expansion signals, and market-wide attribute shifts give models something they can't generate on their own: external context. When competitive data feeds into forecasting and pricing infrastructure, the models stop operating in isolation. They become market-aware, and market-aware models consistently outperform those built on internal data alone. ## The Solution for These Four Problems Is Not Dashboards Most retailers already have dashboards. They show revenue, conversion, sell-through, and margin by category. ![retailer dashboards](https://www.blog.datahut.co/content/images/2026/07/img-440.png.webp) But dashboards are retrospective. They explain what happened inside your business. They rarely answer: - Who moved pricing first? - Where did competitors expand assortment depth? - Which price bands are gaining share? - Which variants are disappearing fastest? Infrastructure-level competitive data does something different than dashboards. It feeds directly into pricing engines, informs merchandising buys, and retrains forecasting models. Dashboards describe the past. Infrastructure changes the next decision. Performance improves when we treat competitive fashion data as infrastructure—something that feeds directly into pricing tools, merchandising workflows, and AI systems. When pricing, merchandising, growth, and strategy teams operate from the same external market reality, alignment improves. Revenue leakage shrinks. Margins stabilize. Stock-out opportunities turn into growth. Inventory decisions become smarter. Data stops being a report; it becomes coordination. ## Let Us Help You Solve These Four Problems Fashion retail is moving faster than ever. Discovery is algorithmic, pricing is dynamic, and demand shifts overnight. The four main profit killers—revenue leakage, margin compression, missed demand, and inventory imbalance—aren't unavoidable. They are symptoms of operating without full visibility. Retailers that embed structured competitive data into everyday decisions don't just react to the market; they move with it. And sometimes, ahead of it. [Datahut](https://www.datahut.co/?ref=blog.datahut.co) helps fashion retailers achieve this. Whether you need real-time competitor pricing, inventory signals, assortment benchmarks, or market-wide attribute trends, Datahut delivers the structured data your teams need to make faster, smarter, and more confident decisions. Get in touch to see what full market visibility looks like for your business. ### About the Author I’m Tony Paul, founder of Datahut, with over 15 years of experience in the web scraping industry. We provide Data as a Service (DaaS) to global retailers and enterprises. My core belief is that most organizations waste resources "re-inventing the wheel" when they should focus on the Value Layer (insights and decision-making) rather than the Commodity Layer (maintaining scrapers and infrastructure). If you’re looking for help with web scraping, pricing, retail data, or anything in between, please connect with me on LinkedIn here: [Tony Paul](https://www.linkedin.com/in/tonypauldh/?ref=blog.datahut.co). ### FAQ Section #### 1\. Why Do Fashion Retailers Need Competitive Data? Fashion retailers need fashion retail competitive data to understand how competitors price products, manage inventory, and adjust assortments. This helps brands make smarter pricing, merchandising, and promotional decisions in a rapidly changing retail market. #### 2\. How Does Competitive Pricing Data Reduce Revenue Leakage? Fashion retail competitive data allows retailers to monitor price changes across the market in real time. This helps pricing teams quickly adjust strategies and avoid losing conversions when competitors lower prices on similar products. #### 3\. Can External Market Data Improve Inventory Planning in Fashion Retail? Yes. Fashion retail competitive data provides insights into demand trends, price band shifts, and competitor assortment strategies, enabling retailers to avoid overstocking slow-moving items or understocking high-demand products. #### 4\. How Can Retailers Capture Demand When Competitors Run Out of Stock? Retailers who use fashion retail competitive data to track competitor inventory levels can identify stock-outs and respond by promoting substitute products, adjusting pricing, or increasing marketing bids to capture redirected customer demand. #### 5\. Which Teams in Fashion Retail Benefit from Competitive Intelligence? Pricing, merchandising, eCommerce growth, and strategy teams benefit from fashion retail competitive data by using market insights to improve pricing decisions, assortment planning, campaign targeting, and forecasting models. ### Amazon vs Argos Smartwatch Pricing Analysis: What Ecommerce Brands Can Learn from Marketplace Data (2026) URL: https://www.blog.datahut.co/post/amazon-vs-argos-smartwatch-pricing-analysis-what-ecommerce-brands-can-learn-from-marketplace-data/ Last updated: 2026-09-07T09:43:30.000Z ## What Marketplace Structure Reveals About Competitive Strategy Most brands focus on price analysis, while few examine the underlying marketplace structure. Smartwatch brands frequently monitor [competitor pricing](https://www.blog.datahut.co/post/how-popular-price-comparison-websites-grab-data/); however, significantly fewer assess how marketplace structure influences these prices. However, [marketplace structure](https://www.statista.com/topics/1556/wearable-technology?ref=blog.datahut.co) plays a critical role. The economic logic of each platform is usually reflected through their Price dispersion, discount frequency, segmentation patterns and inventory velocity. So to study and explore this, we analyzed smartwatches listings of two major marketplaces, Amazon and Argos. ![marketplace pricing analysis ](https://www.blog.datahut.co/content/images/2026/07/img-356.png.webp) So usually the focus of analysis would be finding platforms that give lower prices compared to others but that’s where this [analysis stands out](https://www.blog.datahut.co/post/top-10-web-scraping-companies-in-2026/). Here in our Amazon vs Argos Smartwatch Pricing Analysis, we address a more strategic approach: What do pricing patterns reveal about the competitive strategies of each marketplace? ## Amazon vs Argos Smartwatch Pricing Analysis Dataset This analysis is done from the structured datasets collected and processed by the Datahut team. For [Amazon](https://www.blog.datahut.co/post/the-secret-weapon-of-successful-amazon-sellers/), we [tracked approximately 60-65 smartwatch listings](https://www.blog.datahut.co/post/how-to-scrape-amazon-s-smart-watch-data-for-time-series-analysis/) daily over a one-week period, capturing: ![Amazon dataset](https://www.blog.datahut.co/content/images/2026/07/img-357.png.webp) - Sale and retail prices - Discount percentage - Rating signals - Availability shifts - Brand distribution [For Argos, we analyzed](https://www.blog.datahut.co/post/how-argos-sells-watches-what-the-data-reveals/) structured listings segmented by audience categories, including men’s, women’s, and children’s smartwatches, capturing: ![Argos dataset](https://www.blog.datahut.co/content/images/2026/07/img-358.png.webp) - Price levels - Promotional frequency - Feature distribution - Brand representation The result is structured, comparable marketplace data rather than anecdotal observation. ## Price Dispersion: Two Very Different Architectures On Amazon, smartwatch pricing spans a remarkably wide band: - Lowest sale price: \~₹699 - Highest sale price: \~₹27,995 - Average sale price: \~₹7,800 - Median sale price: \~₹2,800 Average retail prices were materially higher (\~₹13,000), indicating systematic [discounting across listings.](https://www.deloitte.com/us/en/insights/industry/retail-distribution/retail-distribution-industry-outlook.html?ref=blog.datahut.co) The distribution indicates a marketplace optimized for breadth, where budget, mid-range, and premium products coexist. However, sales volume appears concentrated in lower price tiers. Argos presents a contrasting pattern: - Lowest sale price: \~£3.99 - Highest sale price: \~£829 - Average sale price: \~£77 - Median sale price: \~£44.99 Retail and sale prices are closely aligned (average retail \~£78), which indicates significantly less reliance on promotional markdowns. This pattern reflects underlying structural differences: Amazon tolerates, and perhaps even incentivizes, wide price dispersion. Argos operates within tighter pricing corridors. ## Discounting as Competitive Infrastructure The most pronounced divergence emerges in promotional intensity. Amazon listings showed median discounts of 58–66% early in the observation period, with maximum discounts of 88-93%. Discount depth shifted across the week, suggesting dynamic repricing. Only 6.4% of Argos smartwatch listings were discounted. This difference is substantive and reflects two distinct competitive logics: - Amazon appears to compete on liquidity, with high transaction velocity supported by discount elasticity. - Argos appears to compete on positioning, maintaining price integrity, and structured segmentation. In this context, discounting functions as an integral component of marketplace design rather than solely as a marketing tool. ## Inventory Signals and Demand Intensity Amazon data also revealed consistently sold-out listings alongside frequent introduction of new SKUs. This pattern suggests rapid inventory turnover, active consumer demand in specific price tiers, ongoing competitive repositioning The catalog size remained stable despite inventory turnover, indicating ongoing replenishment rather than contraction. In contrast, Argos displayed more static assortment patterns during the observation window, consistent with its more stable pricing approach. ## Segmentation as Strategy Argos’ smartwatch assortment reveals explicit demographic segmentation: - Men’s watches emphasize GPS, durability, and performance metrics. - Women’s models emphasize health tracking and connectivity. - Children’s devices prioritize safety and ease of use. In this case, product features serve as the primary basis for segmentation rather than discounting. Amazon’s segmentation appears less structurally defined and more price-layered, with clustering in lower and mid-tier bands. Therefore, the two marketplaces differ not only in [pricing strategies](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/how-we-help-clients/dynamic-pricing?ref=blog.datahut.co) but also in how they organize consumer demand. ## Brand Power and Price Integrity Across both platforms, premium brands such as Samsung, Garmin, Fossil, and Citizen consistently occupy higher price tiers. However, even premium brands participate in discount cycles on Amazon. On Argos, premium positioning appears more insulated from deep promotional erosion. Pricing power remains brand-dependent; however, marketplace architecture determines how this power is expressed. ## What This Reveals About [Marketplace Economics](https://www.gartner.com/reviews/market/digital-commerce?ref=blog.datahut.co) Taken together, the findings suggest two distinct competitive systems: ![Amazon vs Argos architecture](https://www.blog.datahut.co/content/images/2026/07/img-359.png.webp) Amazon’s architecture: - Wide price dispersion - High promotional volatility - Rapid inventory cycling - Elastic demand stimulation Argos’ architecture: - Structured segmentation - Narrower price bands - Low promotional dependency - Greater price stability For brands, this distinction holds strategic significance. Entering Amazon without a dynamic repricing capability risks margin compression. Entering Argos with a discount-led strategy risks diluting the brand. Marketplace selection is not a neutral decision; it actively shapes pricing behavior. ## Implications for E-commerce Leaders Many e-commerce teams regard pricing as a static decision, whereas marketplace data indicates it should be managed as adaptive infrastructure. Brands that continuously monitor competitor price movement, discount frequency, tier clustering, assortment turnover are better positioned to protect margins, identify white-space segments , avoid reactive discounting and align positioning with platform logic. Without structured monitoring, pricing decisions become speculative. ## Data Transparency The datasets used in this study were extracted and structured by Datahut’s internal web scraping systems from publicly available marketplace listings. To enable independent exploration: - Download the cleaned Amazon and Argos smartwatch datasets. - Review the full exploratory data analysis for both platforms. 📥 Download the cleaned smartwatch dataset ( [Amazon downloadable dataset link here](https://tally.so/r/EkQjBq?ref=blog.datahut.co)) ([Argos downloadable dataset link here](https://tally.so/r/A7Lo6N?ref=blog.datahut.co)) 📊 Explore the full Exploratory Data Analysis (EDA) ([ Amazon EDA link here](https://tally.so/r/EkQjBq?ref=blog.datahut.co)) ([Argos EDA link here](https://tally.so/r/A7Lo6N?ref=blog.datahut.co)) The objective is not only to ensure transparency but also to demonstrate how structured marketplace data can inform pricing strategy. ## The Broader Lesson Pricing differences are not merely tactical; they are systemic. Marketplaces embed competitive logic within their structure. Brands that recognize this dynamic can adapt effectively, while those that overlook it often default to margin-eroding responses. The advantage does not come from lowering prices faster. Competitive advantage arises from understanding the system within which pricing decisions are made. Category-level datasets and automated monitoring allow businesses to track pricing, discounts, and assortment continuously and make faster decisions. To explore competitor pricing datasets or marketplace analysis for a specific category, one can begin with a sample dataset to observe how pricing intelligence operates in practice. For further information or to request a consultation, please contact Datahut at[ ](https://www.datahut.co/?ref=blog.datahut.co)[datahut.co](http://datahut.co/?ref=blog.datahut.co). ### Why AI Web Scraping Fails at Enterprise Scale? URL: https://www.blog.datahut.co/post/why-ai-web-scraping-fails-at-enterprise-scale/ Last updated: 2026-09-07T09:43:32.000Z People often see AI web scraping as fast and simple. But when it comes to [large-scale enterprise](https://www.blog.datahut.co/post/how-popular-price-comparison-websites-grab-data/) use cases, relying just on large language models (LLMs) introduces new risks that you may not notice. When large language models (LLMs) became common, many believed that web scraping problem had finally been solved. The idea was logical at first sight. If AI could read and understand language, it should also handle web pages, extract data, and adapt as sites change, which are some of the hardest [web scraping problems](https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/). For teams frustrated with fragile scripts and constant maintenance, LLMs looked like a new beginning. Even the [market projections](https://www.technavio.com/report/ai-driven-web-scraping-market-industry-analysis?ref=blog.datahut.co) were through the roof. In reality, things have become more complicated. This gap between early expectations and real-world performance is exactly why AI web scraping fails when organizations try to use it at enterprise scale. As my friend Arjun said, you can hack together an AI tool to work at 70% accuracy in a weekend; from there onwards, each percentage gain is extremely difficult. ## The Illusion of Effortless Scale in AI Web Scraping In industries such as retail, real estate, travel, and finance, data engineers find that [LLM-based web scraping ](https://www.researchgate.net/publication/399707730%5FBeyond%5FBeautifulSoup%5FBenchmarking%5FLLM-Powered%5FWeb%5FScraping%5Ffor%5FEveryday%5FUsers?ref=blog.datahut.co)works well in early tests but struggles in large-scale production environments. This leads to a gap between what works in demos and what succeeds in real, high-volume use. LLMs are really good at finding patterns. They can work around messy HTML, summarise text, and make different terms consistent. For small projects or early tests, this is a real improvement. At the enterprise level, large-scale web scraping is more than just reading HTML. It is an operational system that must deliver reliable, [accurate data from websites](https://www.blog.datahut.co/post/noon-skincare-category-analysis/), often in places designed to block automated tools. This is where depending mostly on LLMs begins to fail. ## Three Structural Limitations of LLM-Only Web Scraping ### 1\. Cost grows nonlinearly with volume LLMs charge based on computing power, so AI-powered web scraping gets expensive as volume increases. But web pages are not made for machines. They have a lot of repeated code, scripts, and extra content that must be processed repeatedly. What looks affordable for thousands of pages quickly becomes too costly at millions. Retrying failed pages, updating content, or keeping records only adds to the expense. For organizations needing regular data updates, like price checks or catalog monitoring, this pricing model creates long-term uncertainty. ### 2\. Reliability is probabilistic, not deterministic [Enterprise web scraping](https://www.blog.datahut.co/post/top-10-web-scraping-companies-in-2026/) systems must be predictable to ensure accurate and reliable data extraction. A field is either present or missing; a price is either right or wrong. LLMs, however, give results based on probability. This creates subtle risks: - silent [hallucinations](https://www.lakera.ai/blog/guide-to-hallucinations-in-large-language-models?ref=blog.datahut.co) - website layout drift over time - inconsistent schema - difficulty tracing root causes when errors occur. On their own, these errors might seem minor. But at scale, they multiply and reduce trust in analytics, AI models, and decision-making systems. ### 3\. Access remains the unsolved bottleneck One of the biggest overlooked limits in LLM web scraping is that LLMs do not fix the problem of access. Anti-bot systems check network behaviour, browser details, session patterns, and traffic consistency. No matter how advanced your extraction model is, it cannot collect data it cannot reach. In real-world production, most failures happen before parsing even starts: - [blocked requests](https://www.cloudflare.com/learning/bots/what-is-bot-management/?ref=blog.datahut.co) - degraded coverage - partial crawls - region-specific denials LLMs operate above this layer. They do not replace it. The access problem is extremly hard and needs continues adaptation often requireming a dedicated R&D team. ## [Enterprise Web Scraping](https://www.blog.datahut.co/post/build-vs-buy-web-scraping-in-2026-the-definitive-guide-for-data-teams/) Is an Infrastructure Problem The main mistake is to see web scraping as just a single technical skill. In reality, enterprise web scraping and scalable [data extraction](https://www.casestudy.datahut.co/post/amazon-seller-case-study?ref=blog.datahut.co) are more like other core infrastructure: - It must be resilient to change - observable and auditable - cost-predictable - compliant by design - adaptable without constant rework This means you need a system with multiple layers, not just one model. ## Where AI Actually Creates Durable Value in Scalable Web Scraping This does not make AI less important. In fact, AI is essential when used wisely. Effective enterprise web scraping systems separate concerns across access, extraction, and validation: - Access and crawl control are handled through engineered, deterministic mechanisms. - Stable data fields rely on rule-based or ML-assisted extraction. - Unstructured or ambiguous content is selectively routed through LLMs. - Quality and drift detection are continuously monitored, not inferred. - Human feedback loops correct edge cases and retrain models where it matters. In this setup, LLMs enhance the system’s abilities instead of being a single point of failure. ## A Strategic Question for Leaders For executives considering AI-driven data projects, the real question is not: “Can AI extract this data?” It is: “Can this system deliver trustworthy data, every day, at scale, under real-world constraints?” Organizations that miss this point often have to rebuild their data pipelines after early wins, but now with more urgency and higher costs. ## Conclusion: Scale Is a Discipline, Not a Feature in Enterprise Web Scraping LLMs have changed how we experiment, but they have not replaced the basics of good system design. Teams that treat web scraping as a strategic skill, not just a quick fix, are building systems that combine strong engineering with focused AI support. The result is not just faster and more scalable data extraction, but also a lasting advantage for enterprise data. In today’s economy, where external data is more important than ever, that difference truly matters. ## Build infrastructure or buy data from us You could spend engineering time battling proxies, fingerprints and more web scraping headaches Or you could plug into Datahut's battle-tested extraction layer—get structured, data feeds delivered daily—and reserve your AI budget for actual value-add tasks like enrichment and analysis. ## Frequently Asked Questions About AI Web Scraping 1. Why do AI web scrapers fail at web scraping at scale?AI web scrapers struggle because they are probabilistic systems and they can't crack the access problem at scale. While they work well for small experiments, they introduce reliability, cost, and consistency issues in production-grade web scraping environments, especially when combined with anti-bot protection and frequent website changes. 2. Is AI web scraping scalable for enterprise use?AI web scraping can be scalable when used selectively within a broader enterprise web scraping architecture. Deterministic crawling, access management, and validation layers must be in place alongside LLMs to ensure reliability, cost control, and long-term scalability. 3. How should enterprises think about LLMs vs traditional web scraping?LLMs vs traditional web scraping is not an either-or decision. Traditional scraping provides deterministic control and cost predictability, while LLMs add value in unstructured data extraction and enrichment. The most effective systems combine both approaches. ### Scraping Condo Listings from Homes.com Using Playwright: A Complete Engineering Guide URL: https://www.blog.datahut.co/post/scraping-condo-listings-from-homes-com-using-playwright-a-complete-engineering-guide/ Last updated: 2026-09-07T09:43:35.000Z Is buying a [condo](https://www.investopedia.com/terms/c/condominium.asp?ref=blog.datahut.co) in California really more expensive than owning one in New York, or does the story change once the numbers are laid out side by side? To explore this question, real [condo listings](https://en.wikipedia.org/wiki/Condominium?ref=blog.datahut.co) were carefully scraped from [Homes.com](https://www.homes.com/?ref=blog.datahut.co), with [California and New York](https://www.blog.datahut.co/post/california-vs-new-york-condo-prices-2025-1-400-homes-com-data-insights-revealed/) collected separately to keep the comparison clear and fair, using publicly available listing pages such as the California condos section on Homes. Rather than relying on assumptions or headlines, this approach builds a solid data foundation straight from the source, making it easier to understand how these two iconic housing markets [differ in cost](https://www.investopedia.com/terms/c/condominium-fee.asp?ref=blog.datahut.co), scale, and overall buying pressure. The sections that follow document this data scraping journey step by step, showing how structured code can quietly turn thousands of online listings into meaningful, ready-to-analyze housing data. ## Scraping Condo Listings from Homes.com Using Playwright Here is the step-by-step step guide on Scraping Condo Listings from Homes.com Using Playwright ## STEP 1 : URL Collection from Condo Listings When working with Homes.com , the first real challenge is not extracting prices or property details, but reliably finding every individual property URL spread across many result pages. Homes.com loads listings dynamically and divides them across multiple pages, which means important links are not always visible at once and can easily be missed if the page is rushed. To solve this, a structured scraping flow was built using Python and browser automation tools to carefully move through the listings, starting from the California condos page and then repeating the same process separately for New York. The scraper opens the site like a normal visitor, waits for listings to load, scans each page for property cards, and collects only clean, valid links before moving to the next page. Small pauses are added between actions so the browsing pattern looks natural, while logging quietly records progress in the background. Instead of saving links in loose files, California condo URLs are stored in one database table and New York URLs in another, keeping both regions clearly separated and easy to manage later. This step works much like writing down the addresses of houses in two different cities before visiting them—slow, careful, and organized—ensuring that the foundation is solid before moving on to deeper property-level scraping. ## STEP 2 : Structured Data Extraction from Individual Property Pages Once the list of property URLs was safely stored in the database, the next step was to visit each link one by one and gently extract the details that truly describe a condo listing. The property pages often load important information only after the page settles, so each URL was opened carefully, allowing the content to fully appear before anything was read. The scraper fetched the page, parsed the HTML, and then looked for clear, reliable markers to capture essentials such as the property image, listed price, and address, similar to slowly reading a property brochure instead of skimming it. Each extracted record was saved back into a structured database table, with California listings written to one table and New York listings to another, keeping both markets cleanly separated and easy to manage. This stage completes the journey that began with URL collection, transforming simple links into meaningful, organized housing data that is ready for comparison and deeper understanding. ## STEP 3 : Data Cleaning with OpenRefine After scraping condo listings from Homes.com and storing California and New York data in separate tables, the next important step was cleaning the raw data so it could be read, compared, and analyzed with confidence. The scraped files were uploaded into OpenRefine, a tool designed to make large datasets easier to inspect and fix, where missing values were first standardized by filling empty fields with clear markers like “N/A” to avoid confusion later. Duplicate property URLs, which can quietly appear during repeated scraping or pagination, were identified and removed to ensure each condo appeared only once. Price fields often contained currency symbols and formatting that looked fine to the eye but caused issues during analysis, so these symbols were carefully removed and the values normalized into a consistent numeric format. To improve readability, some combined fields were split into separate columns, making details like location and pricing clearer at a glance. This cleaning stage acts like tidying handwritten notes into a neat spreadsheet, turning raw scraped content into a reliable and well-structured dataset that is ready for meaningful comparison between California and New York housing markets. ## Advanced Tool-sets and Libraries for Efficient Data Extraction This project brings together a carefully chosen set of Python libraries that make the process of scraping condo listings from Homes.com clear and reliable, even when [California](https://www.homes.com/california/condos-for-sale/?ref=blog.datahut.co) and [New York](https://www.homes.com/new-york-ny/condos-for-sale/?ref=blog.datahut.co) data are handled separately. [Playwright](https://www.lambdatest.com/playwright?ref=blog.datahut.co), accessed through playwright.sync\_api, is responsible for opening real browser pages and navigating listings the same way a human would, while playwright\_stealth quietly adjusts browser signals so the activity looks natural rather than automated. Alongside browser automation, requests is used for fetching individual pages through an API-based approach when needed, and [BeautifulSoup](https://beautiful-soup-4.readthedocs.io/en/latest/?ref=blog.datahut.co) helps gently read the HTML and pull out meaningful details like prices, images, and addresses, similar to carefully scanning a printed brochure for key information. Data handling is kept simple and lightweight using [sqlite3](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co) for storing URLs and tracking progress locally, and json for exporting clean, shareable results. Supporting libraries such as time and [random](https://www.w3schools.com/python/module%5Frandom.asp?ref=blog.datahut.co) introduce small pauses and variation between actions to keep browsing behavior realistic, while logging acts like a detailed activity journal that records what happens at each step, making errors easier to understand and fix. Utilities like Path from [pathlib](https://www.geeksforgeeks.org/python/pathlib-module-in-python/?ref=blog.datahut.co) help manage files and folders cleanly, and List from typing improves code clarity without adding complexity. Together, these tools form a smooth workflow where browsing, data extraction, storage, and monitoring flow naturally, making the entire scraping process approachable and easy to follow for interns or beginners exploring real-world web data collection for the first time. ## STEP 1: Extracting Condo Listing URLs from the [Homes.com](http://homes.com/?ref=blog.datahut.co) ### Importing Libraries ``` Imports import time import random import json import sqlite3 import logging from pathlib import Path from typing import List from playwright.sync_api import sync_playwright, Page from playwright_stealth import stealth_sync ``` The code begins by importing a small set of essential [Python libraries](https://citrusbug.com/blog/python-libraries/?ref=blog.datahut.co) that work together to make the scraping process smooth and reliable. Playwright is used to open and interact with dynamic pages on Homes.com, playwright\_stealth helps reduce bot detection, while standard libraries like json, sqlite3, and logging handle data storage, structure, and execution tracking so the California and New York condo data can be collected and saved in an organized way. ### Setting the Starting Point and Storage Paths ``` Configuration START_URL = "https://www.homes.com/california/condos-for-sale/" BASE_URL = "https://www.homes.com" USER_AGENTS_FILE = "/home/anusha/Desktop/DATAHUT/chewy-and-petco/user_agents.txt" SQLITE_DB = "/home/anusha/Desktop/DATAHUT/condos/Data/homes_urls_california.db" JSON_OUT = "/home/anusha/Desktop/DATAHUT/condos/Data/homes_urls_california.json" LOG_FILE = "/home/anusha/Desktop/DATAHUT/condos/Log/homes_scraper_california.log" ``` This configuration section defines the foundation for scraping condo listings from Homes.com in a clear and controlled way, starting with the California condos page using the START\_URL, while the BASE\_URL helps convert relative links into complete, usable URLs as the scraper moves through multiple pages. To reduce blocking and appear more like a real browser, a list of rotating [user agents](https://www.blog.datahut.co/post/scraping-gucci-made-simple-a-beginner-s-guide-to-extracting-fashion-data/) is loaded from an external file, a common and beginner-friendly practice explained in many web scraping guides such as the Playwright documentation. The scraped condo URLs are stored safely in both a [SQLite](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co) database and a [JSON file](https://blog.hubspot.com/website/json-files?ref=blog.datahut.co), making the data easy to reuse later for analysis or sharing, while a dedicated log file records each step of the process so errors or interruptions can be traced and fixed without guesswork. By keeping California and New York data in separate configuration setups, the scraping process remains organized, scalable, and easy to understand, much like keeping files in clearly labeled folders rather than mixing everything together. ### Using Browser-Like Headers to Avoid Blocking ``` Headers to mimic a real browser EXTRA_HEADERS = { "Content-Type": "text/html; charset=utf-8", "x-frame-options": "SAMEORIGIN", "x-content-type-options": "nosniff", "Content-Security-Policy": "frame-ancestors 'self' https://*.homes.com https://auth.homes.com;", "Permissions-Policy": "browsing-topics=()", } ``` This section adds extra [request headers](https://playwright.dev/docs/network?ref=blog.datahut.co) that help the scraper behave more like a real web browser when visiting Homes.com, which is especially important when collecting condo listings separately from California and New York. Headers such as Content-Type and security-related fields quietly tell the website how the page should be handled, reducing the chances of the request being flagged as automated and improving page stability during loading. Using [headers](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers?ref=blog.datahut.co) in this way is a widely accepted best practice in responsible web scraping. By keeping these settings simple and consistent, the scraper can access listing pages more smoothly while staying aligned with standard web behavior. ### Managing Cookies for Stable Page Access ``` Cookies COOKIES = [ {"name": "sb", "value": "%255b%257b%2522pt%2522%253a4%252c%2522gk%2522%253a%257b%2522key%2522%253a%25227969zft86ltlw%2522%257d%252c%2522gt%2522%253a1%257d%255d", "domain": "www.homes.com", "path": "/", "httpOnly": True, "secure": True}, {"name": "ls", "value": "%2Fcalifornia%2Fcondos-for-sale%2F", "domain": "www.homes.com", "path": "/", "httpOnly": True, "secure": True}, ] ``` This cookies section helps the scraper maintain a steady and predictable browsing session while visiting condo listings on Homes.com, whether the data comes from California or New York, by storing small pieces of information that the website normally saves in a real user’s browser. In simple terms, [cookies](https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Cookies?ref=blog.datahut.co) act like a bookmark and preference note combined, allowing the site to remember visited paths and session details so pages load correctly without repeated interruptions or redirects. By predefining these cookies, the scraper avoids unnecessary reloads and behaves more like a regular visitor calmly browsing listings, rather than repeatedly knocking on the door as a new guest each time. ### Adding Delays for Human-Like Browsing ``` Delay between page loads DELAY_MIN = 1.0 DELAY_MAX = 4.0 MAX_PAGES = None # Set limit if needed ``` This section introduces a small random delay between page loads so the scraper moves through website at a calm, natural pace, similar to how a real person would browse condo listings in different locations. By waiting a short, varied amount of time between requests, the process becomes more stable and respectful to the website. ### Logging: Keeping Track of Every Step ``` Logging setup def setup_logger(): """Sets up a logger to record all scraper activity""" log_dir = Path(LOG_FILE).parent log_dir.mkdir(parents=True, exist_ok=True) logger = logging.getLogger("homes_scraper_california") logger.setLevel(logging.DEBUG) formatter = logging.Formatter("%(asctime)s [%(levelname)s] %(message)s", "%Y-%m-%d %H:%M:%S") file_handler = logging.FileHandler(LOG_FILE, mode="w", encoding="utf-8") file_handler.setLevel(logging.DEBUG) file_handler.setFormatter(formatter) console_handler = logging.StreamHandler() console_handler.setLevel(logging.INFO) console_handler.setFormatter(formatter) logger.addHandler(file_handler) logger.addHandler(console_handler) return logger logger = setup_logger() ``` This logging setup creates a clear activity record for the [Homes.com](http://homes.com/?ref=blog.datahut.co) condo scraper, making it easy to understand what happens while collecting listings from California and New York, especially when the two regions are handled separately. The logger acts like a travel journal for the scraper, writing detailed messages to a log file for later review while also showing important updates on the screen, and commonly used in real-world scraping projects to quickly spot errors, interruptions, or unexpected page behavior without guessing what went wrong. ### Utility Function for Loading User Agents ``` Utilities def read_user_agents(path: str) -> List[str]: """Reads user-agent strings from a text file""" p = Path(path) if not p.exists(): logger.error(f"User agents file not found: {path}") raise FileNotFoundError(f"User agents file not found: {path}") lines = [l.strip() for l in p.read_text(encoding="utf-8").splitlines() if l.strip()] logger.debug(f"Loaded {len(lines)} user agents.") return lines ``` This utility function safely reads browser user-agent strings from an external text file and prepares them for use during scraping, which helps the [Homes.com](http://homes.com/?ref=blog.datahut.co) condo crawler appear more like a real visitor while collecting listings separately from California and New York. By checking whether the file exists, logging useful messages, and returning only clean, non-empty entries, the function adds a small but important layer of reliability. ### Initializing the SQLite Database for Storing URLs ``` Initialize SQLite DB and create table if not exists def init_db(db_path: str): """Creates (or connects to) an SQLite database and sets up the 'urls' table""" conn = sqlite3.connect(db_path) cur = conn.cursor() cur.execute(""" CREATE TABLE IF NOT EXISTS urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT NOT NULL UNIQUE, inserted_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP ) """) conn.commit() logger.debug(f"Initialized database at {db_path}") return conn ``` This database setup function creates a simple and reliable place to store condo listing URLs, keeping data neatly organized even when they are scraped separately. Using SQLite allows the scraper to save each unique URL with a timestamp, much like writing entries into a small digital notebook that automatically avoids duplicates, making the data easy to manage and revisit later without additional complexity. ### Saving and Deduplicating Property URLs ``` Save a URL to the database, return True if new, False if duplicate def save_url_to_db(conn: sqlite3.Connection, url: str) -> bool: """Inserts a property URL into the database if it’s new""" try: cur = conn.cursor() cur.execute("INSERT INTO urls (url) VALUES (?)", (url,)) conn.commit() logger.debug(f"Inserted new URL: {url}") return True except sqlite3.IntegrityError: logger.debug(f"Duplicate URL skipped: {url}") return False ``` This function handles the careful task of saving each URL into the [SQLite database](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co#sqlite3.IntegrityError) while automatically skipping duplicates, which is especially helpful when collecting data separately for California and New York. It works like a checklist that marks a URL only once—new links are stored and logged, while repeated ones are quietly ignored— ensuring the dataset stays clean, accurate, and easy to work with as scraping progresses. ### Exporting Collected URLs to a JSON File ``` Dump all URLs from DB to JSON file def dump_db_to_json(conn: sqlite3.Connection, out_path: str): """Exports all stored URLs from the database into a JSON file""" cur = conn.cursor() cur.execute("SELECT url FROM urls ORDER BY id") rows = [r[0] for r in cur.fetchall()] with open(out_path, "w", encoding="utf-8") as f: json.dump(rows, f, indent=2) logger.info(f"Wrote {len(rows)} unique URLs to {out_path}") ``` This function takes all the condo listing URLs stored in the SQLite database and neatly exports them into a [JSON file](https://docs.python.org/3/library/json.html?ref=blog.datahut.co), creating a clean and portable output that is easy to read, share, or use in later analysis. It acts like packing a well-organized folder from a notebook into a digital file, ensuring the final scraped data is structured, and reusable. ### Extracting Property Links from Each Listings Page ``` Scraper def extract_property_urls(page: Page) -> List[str]: """Extracts all property listing URLs from the current page""" script = """ () => { const nodes = Array.from(document.querySelectorAll('div.description-container')); const urls = []; for (const el of nodes) { let a = el.closest('a'); if (!a) a = el.querySelector('a') || el.parentElement?.querySelector('a'); if (a && a.href) urls.push(a.href); } return Array.from(new Set(urls)); } """ try: urls = page.evaluate(script) except Exception as e: logger.exception("evaluate failed: %s", e) urls = [] normalized = [BASE_URL.rstrip("/") + u if u.startswith("/") else u for u in urls] return normalized ``` This scraper function focuses on carefully finding and collecting individual condo listing links by looking for the page sections that describe each property and then tracing them back to their clickable links. I scans the page the same way a person’s eyes would move across listing cards, gathers only valid and unique URLs, fixes any partial links using the base website address, and safely handles errors if the page structure changes. ### Moving Through Pagination to Reach More Listings ``` Navigate to the next page if available def click_next(page: Page, current_page: int) -> bool: """The function looks for the 'Next Page' button or link using its page number""" try: logger.info(f"Looking for next page link (current page {current_page})...") next_link = page.query_selector(f'a.text-only[title="Page {current_page + 1}"], a[data-page="{current_page + 1}"]') if not next_link: logger.info("No next page link found.") return False href = next_link.get_attribute("href") if not href: logger.info("Next link has no href.") return False next_url = BASE_URL.rstrip("/") + href logger.info(f"Navigating to next page: {next_url}") page.goto(next_url, timeout=60000) page.wait_for_selector("div.description-container", timeout=15000) return True except Exception as e: logger.warning(f"Failed to navigate next page: {e}") return False ``` This pagination function helps the scraper[ move smoothly from one results page to the next](https://playwright.dev/docs/navigations?ref=blog.datahut.co) by looking for the link that points to the following page number and navigating to it when available, which is essential when collecting condo listings spread across multiple pages. It works like turning the next page of a catalog—checking that the page exists, opening it safely, and waiting until the listings load—while logging each step and handling errors gracefully. ### Starting the Scraper and Preparing the Browser ``` Main Scraper def run_scraper(): """The main function that runs the entire scraping process from start to finish""" user_agents = read_user_agents(USER_AGENTS_FILE) conn = init_db(SQLITE_DB) with sync_playwright() as p: browser = p.chromium.launch(headless=False, args=["--start-maximized"]) chosen_ua = random.choice(user_agents) logger.info(f"Using User-Agent: {chosen_ua}") context = browser.new_context( user_agent=chosen_ua, viewport={"width": 1366, "height": 768}, extra_http_headers=EXTRA_HEADERS, ) ``` This first part sets the stage for the entire scraping process by loading the list of user agents, connecting to the SQLite database, and launching a real browser using Playwright. A random user agent is selected to make each run look like a normal browsing session, similar to how different people use different devices or browsers, which helps reduce blocking. The browser context is then created with a standard screen size and custom headers so the Homes.com pages load as expected, forming a stable base before visiting the condo listings. ``` try: context.add_cookies([{**c, "url": BASE_URL} for c in COOKIES]) except Exception: logger.warning("Failed to add cookies; continuing without them.") page = context.new_page() stealth_sync(page) logger.info(f"Opening start URL: {START_URL}") page.goto(START_URL, timeout=60000) time.sleep(random.uniform(DELAY_MIN, DELAY_MAX)) ``` This section focuses on opening the Homes.com listing page in a calm and realistic way by first adding cookies, if possible, and then enabling stealth mode to reduce automated detection. Cookies help the site remember session details, while stealth settings adjust small browser signals so the page behaves as if a real person is browsing. After opening the California condos start URL, the scraper pauses briefly using a random delay, much like a person taking a moment to read the page before scrolling. ``` page_num, total_new = 1, 0 while True: logger.info(f"Scraping page {page_num}") try: page.wait_for_selector("div.description-container", timeout=15000) except Exception: logger.warning("Listings not detected quickly; proceeding anyway.") urls = extract_property_urls(page) new_count = sum(save_url_to_db(conn, u) for u in urls) total_new += new_count logger.info(f"Page {page_num}: Found {len(urls)} URLs ({new_count} new)") ``` Once the page is loaded, this part handles the core task of finding condos links on each results page and saving them safely into the database. The scraper waits for listing containers to appear, extracts all property URLs, and stores only new ones while skipping duplicates, keeping the dataset clean and organized. Each page is logged with the number of links found, which helps track progress clearly when scraping large regions like California or New York. ``` if MAX_PAGES and page_num >= MAX_PAGES: logger.info("Reached max pages limit. Stopping.") break if not click_next(page, page_num): logger.info("No more pages. Scraping complete.") break time.sleep(random.uniform(DELAY_MIN, DELAY_MAX)) page.wait_for_load_state("networkidle", timeout=20000) time.sleep(random.uniform(1, 2)) page_num += 1 dump_db_to_json(conn, JSON_OUT) logger.info(f"Scraping finished: {page_num} pages, {total_new} new URLs collected.") context.close() browser.close() conn.close() ``` The final part controls how the scraper moves across multiple pages and knows when to stop, either by reaching a page limit or when no next page is available. After each successful page turn, the scraper waits for the network to settle and adds short pauses to keep the browsing pattern natural and stable. When all pages are completed, the collected URLs are exported into a JSON file, the browser and database connections are closed properly, and a final summary is written to the logs. ### Script Entry Point and Safe Execution ``` Entry point if __name__ == "__main__": """Runs the scraper when this script is executed directly""" try: run_scraper() except Exception as e: logger.exception(f"Fatal error: {e}") ``` This entry point ensures the scraper runs only when the file is executed directly, acting like a clear “start here” sign for the program while keeping the code safe if it is imported elsewhere. By wrapping the main scraping logic in a try–except block, any [unexpected issue ](https://docs.python.org/3/tutorial/errors.html?ref=blog.datahut.co)during the Homes.com condo collection for California or New York is captured and written to the logs instead of stopping silently, making the overall workflow easier to understand and more reliable. ## Step 2: Extracting Complete Property Details from Individual [Homes.com](http://homes.com/?ref=blog.datahut.co) Listings ### Importing Libraries ``` Imports import sqlite3 import requests import random import time import logging from bs4 import BeautifulSoup import json ``` ### Defining Paths for Data, Logs, and Browser Identity ``` Configurations DB_PATH = "/home/anusha/Desktop/DATAHUT/condos/Data/homes_urls_california.db" USER_AGENTS_FILE = "/home/anusha/Desktop/DATAHUT/chewy-and-petco/user_agents.txt" LOG_FILE = "/home/anusha/Desktop/DATAHUT/condos/Log/homes_scraper.log" ``` This configuration block sets clear file paths that tell the scraper where to store collected condo URLs, where to read browser [user-agent](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/User-Agent?ref=blog.datahut.co) details, and where to write log messages while visiting Homes.com listings for California or New York. It is like assigning fixed shelves for notes, tools, and progress records. ### Managing the Scraper API Key Securely ``` SCRAPERAPI_KEY = "********************686" ``` This configuration line represents the API key used to route requests through a [scraping service](https://www.blog.datahut.co/post/is-web-scraping-legal/), which helps access Homes.com listings more reliably when collecting data. This key can be thought of as an access pass that allows the scraper to use an external helper service, and best practice is to keep it private and load it from environment variables instead of writing it directly in the code. ### Basic Logging Configuration for Clear Progress Tracking ``` Logging setup logging.basicConfig( filename=LOG_FILE, level=logging.INFO, format='%(asctime)s [%(levelname)s] %(message)s', datefmt='%Y-%m-%d %H:%M:%S' ) """This section sets up all the basic configurations needed for the scraper to run""" ``` This [logging setup](https://docs.python.org/3/library/logging.html?ref=blog.datahut.co) defines how the scraper records its activity while collecting condo listings, by saving clear, time-stamped messages into a log file. In simple terms, it works like a running diary that notes what happened and when, making it much easier to understand the scraper’s behavior or diagnose issues later. ### Using Browser-Like Headers for Natural Page Requests ``` Headers to mimic a real browser EXTRA_HEADERS = { "User-Agent": "", # will be overwritten "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5" } ``` ### Cookies for Maintaining a Stable Browsing Session ``` Cookies COOKIES = [ {"name": "sb", "value": "%255b%257b%2522pt%2522%253a4%252c%2522gk%2522%253a%257b%2522key%2522%253a%25227969zft86ltlw%2522%257d%252c%2522gt%2522%253a1%257d%255d", "domain": "www.homes.com", "path": "/", "httpOnly": True, "secure": True}, {"name": "ls", "value": "%2Fcalifornia%2Fcondos-for-sale%2F", "domain": "www.homes.com", "path": "/", "httpOnly": True, "secure": True}, ] ``` ### Selecting a Random User-Agent for Natural Browsing ``` Helper function to get a random User-Agent def get_random_user_agent(): with open(USER_AGENTS_FILE, 'r') as f: user_agents = [line.strip() for line in f if line.strip()] return random.choice(user_agents) ``` This helper function simply reads a list of browser user-agent strings from a text file and randomly picks one each time the scraper runs, helping the scraper appear like a regular visitor rather than an automated script when accessing California or New York listings. In easy terms, it is similar to changing the type of browser being used on each visit, which helps pages load more smoothly and reduces unnecessary blocking. ### Handling Requests Safely with Retries and Smart Delays ``` Robust request function with retries and backoff def make_request(url, retries=2, timeout=60): for attempt in range(1, retries + 1): try: user_agent = get_random_user_agent() headers = EXTRA_HEADERS.copy() headers['User-Agent'] = user_agent params = { 'api_key': SCRAPERAPI_KEY, 'url': url, 'country_code': 'us', 'device_type': 'desktop' } response = requests.get("https://api.scraperapi.com/", params=params, headers=headers, timeout=timeout) if response.status_code == 200: return response.text else: logging.warning(f"Request failed with status {response.status_code} for {url}. Attempt {attempt}/{retries}") except requests.exceptions.RequestException as e: logging.warning(f"Request exception for {url} (Attempt {attempt}/{retries}): {e}") # Exponential backoff with jitter sleep_time = min(10, 2 ** attempt + random.random()) time.sleep(sleep_time) logging.error(f"All retries failed for {url}") return None ``` This function is designed to fetch pages in a steady and reliable way, even when the network is slow or the site temporarily refuses a request, which can happen when collecting condo data separately for large regions. Each [request](https://requests.readthedocs.io/en/latest/?ref=blog.datahut.co) is sent with a randomly chosen browser identity and helpful headers, routed through a scraping service, and if something goes wrong, the function patiently retries after waiting for a gradually increasing amount of time, similar to taking short breaks before trying again rather than rushing. This retry-and-wait approach, often called backoff, helps avoid unnecessary failures , making the scraping process calmer, more resilient, and easier to understand for those new to web automation. ### Parsing Individual Property Pages into Structured Data ``` Parse property page HTML to extract data def parse_property_page(html, url): soup = BeautifulSoup(html, 'html.parser') data = { "url": url, "image_url": None, "price": None, "address": None } img_tag = soup.select_one('figure#carousel-primary-photo img') if img_tag: data['image_url'] = img_tag['src'] if img_tag else None price_tag = soup.select_one('span.property-info-price') if price_tag: data['price'] = price_tag.get_text(strip=True) if price_tag else None address_tag = soup.select_one('div.property-info-address') if address_tag: data['address'] = ' '.join(address_tag.stripped_strings) if address_tag else None return data ``` This [parsing function](https://www.octoparse.com/blog/data-parsing-guide?ref=blog.datahut.co) takes the raw HTML of a single Homes.com property page and gently turns it into clean, readable information by locating key elements such as the main image, listed price, and property address. In simple terms, it works like carefully reading a flyer and picking out only the important details, using clear HTML markers to avoid confusion if the page layout changes slightly. By returning the data in a simple dictionary, the scraper keeps California and New York condo details well organized and ready for further use or analysis without adding unnecessary complexity. ### Creating a Table to Store Scraped Property Details ``` Database setup def create_table_if_not_exists(conn): conn.execute(''' CREATE TABLE IF NOT EXISTS scraped_data ( url TEXT PRIMARY KEY, image_url TEXT, price TEXT, address TEXT ) ''') conn.commit() ``` This database setup function prepares a [simple table](https://www.sqlite.org/lang%5Fcreatetable.html?ref=blog.datahut.co) where each scraped Homes.com property can be saved in an organized and reliable way, whether the data comes from California or New York. In everyday terms, it creates a structured notebook with fixed columns for the property link, image, price, and address, and uses the URL as a unique key so the same listing is not stored twice. By setting up the table only if it does not already exist, the scraper can be run multiple times without breaking earlier data, keeping the overall workflow clean and easy to understand. ### Tracking Progress with a processed Column ``` Add 'processed' column if not exists def add_processed_column_if_not_exists(conn): cur = conn.cursor() cur.execute("PRAGMA table_info(urls)") columns = [col[1] for col in cur.fetchall()] if 'processed' not in columns: cur.execute("ALTER TABLE urls ADD COLUMN processed INTEGER DEFAULT 0") conn.commit() ``` This small [helper function](https://www.geeksforgeeks.org/javascript/what-are-the-helper-functions/?ref=blog.datahut.co) prepares the database to track scraping progress by checking whether a processed column already exists in the URLs table and adding it only if it is missing. In simple terms, this column works like a checklist mark, showing whether a condo listing from Homes.com has already been handled or still needs attention, which is especially useful when California and New York data are scraped separately or when a long run stops midway. ### Main Loop for Managing Unprocessed URLs ``` Main loop def main(): conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() add_processed_column_if_not_exists(conn) create_table_if_not_exists(conn) cursor.execute("SELECT url FROM urls WHERE processed=0") urls = [row[0] for row in cursor.fetchall()] logging.info(f"Found {len(urls)} unprocessed URLs") for idx, url in enumerate(urls, 1): logging.info(f"Scraping ({idx}/{len(urls)}): {url}") html = make_request(url) if not html: logging.warning(f"Failed to get HTML for {url}") continue ``` This main loop acts as the control center of the scraper, opening the database, preparing required tables, and identifying which Homes.com condo URLs still need to be processed, a step that is especially useful when California and New York listings are handled separately. In simple terms, it reads a to-do list from the database, logs how many property pages are pending, and then visits each URL one by one, carefully requesting the page content and skipping over any link that fails to load. This structured flow makes long scraping runs easier to pause and resume. ``` try: data = parse_property_page(html, url) cursor.execute(""" INSERT OR IGNORE INTO scraped_data (url, image_url, price, address) VALUES (?, ?, ?, ?) """, ( data["url"], data["image_url"], data["price"], data["address"] )) cursor.execute( "UPDATE urls SET processed=1 WHERE url=?", (url,) ) conn.commit() except Exception as e: logging.error(f"Error parsing or inserting data for {url}: {e}") time.sleep(random.uniform(1, 5)) conn.close() logging.info("Scraping completed.") ``` Once a property page is successfully loaded, this part focuses on turning the page into useful data and storing it safely, while clearly marking progress in the database. The scraper extracts key details like the image, price, and address, saves them into a separate table, and then updates the original URL record to show that it has been processed, much like ticking off a completed task on a checklist. Short pauses between each request keep the browsing pattern natural and stable, and any [errors ](https://www.geeksforgeeks.org/python/beautifulsoup-error-handling/?ref=blog.datahut.co)are logged without stopping the entire run. ### Entry Point: Starting the Scraper Safely ``` Entry point if __name__ == "__main__": main() ``` This entry point tells Python to start the scraping process only when the script is run directly, acting like a clear “start button” for the program. By calling the [main() function](https://docs.python.org/3/library/%5F%5Fmain%5F%5F.html?ref=blog.datahut.co) here, the code stays organized and avoids running unintentionally if the file is reused elsewhere. ## Conclusion By the end of this workflow, the entire scraping journey feels less like a collection of isolated code blocks and more like a smooth, well-guided process. Each part—starting from discovering listing pages, moving through individual property details, and finally storing clean, usable data—fits together naturally, much like following a clear route on a map rather than guessing directions along the way. The careful use of delays, logging, progress tracking, and simple storage ensures that the scraper remains steady even when handling large regions like California and New York separately. For anyone stepping into real-world data collection for the first time, this approach offers a reassuring reminder that with the right structure and patience, complex-looking tasks can be broken down into manageable, understandable steps that quietly do their job in the background. ## Libraries and Versions Used Name: random Version: Built-in Python module Name: sqlite3 Version: Built-in Python module Name: json Version: Built-in Python module Name: logging Version: Built-in Python module Name: BeautifulSoup (bs4) Version: 4.12.3 Name: playwright Version: 1.48.0 Name: playwright-stealth Version: 1.0.6 ## AUTHOR I’m Anusha P O, a Data Science Intern at Datahut, with hands-on experience in building reliable and scalable web-scraping workflows. In this blog, the focus is on extracting structured condo listing data from Homes.com, where California and New York listings were scraped separately and organized into clean, usable datasets. The walkthrough covers how dynamic listing pages were navigated using Playwright, how property URLs and details were stored safely in databases, and how raw web content was cleaned and refined into analysis-ready data using tools like SQLite, JSON, and OpenRefine. At [Datahut](http://homes.com/?ref=blog.datahut.co), the work centers on helping businesses turn public web data into meaningful intelligence for market research, pricing analysis, and location-based insights. If there is interest in understanding real-estate trends, building dependable data pipelines, or working with large, unstructured web datasets, feel free to connect through the chat widget. Raw listings become far more valuable when they are transformed into clear, structured insights that support confident decision-making. ## Frequently Asked Questions (FAQs) 1\. Why use Playwright for scraping condo listings from [Homes.com](http://homes.com/?ref=blog.datahut.co)?Playwright is ideal for scraping modern real estate websites because it can handle JavaScript-heavy pages, dynamic content loading, pagination, and user interactions like filters and scrolling. This ensures accurate extraction of listing details that may not appear in the initial HTML. 2\. What type of data can be extracted from [Homes.com](http://homes.com/?ref=blog.datahut.co) condo listings?You can extract property titles, prices, locations, number of bedrooms and bathrooms, square footage, listing agents, property descriptions, images, and availability status. This data is useful for market analysis, price benchmarking, and real estate trend tracking. 3\. How do you handle pagination and dynamic loading on [Homes.com](http://homes.com/?ref=blog.datahut.co)?Pagination and dynamic loading can be handled in Playwright by simulating user actions such as clicking “Next,” scrolling to trigger lazy loading, and waiting for network requests or DOM elements to load before extracting data. 4\. Is it legal to scrape real estate listing websites like Homes.com?Scraping is generally allowed when done responsibly, respecting the website’s terms of service, robots.txt guidelines, and applicable laws. It’s important to avoid excessive requests, bypassing authentication, or collecting personal data improperly. 5\. What are the common challenges when scraping real estate websites?Common challenges include anti-bot protections, frequently changing page structures, dynamic rendering, rate limits, and image-heavy pages. Using proper request throttling, resilient selectors, and automated monitoring helps maintain stable scraping pipelines. ### Best Web Scraping Services in 2026: Reviewed and Compared URL: https://www.blog.datahut.co/post/best-web-scraping-services-in-2026/ Last updated: 2026-09-07T09:43:37.000Z Web data is now essential for analytics, AI, pricing, and business decisions. Still, [data professionals spend almost 80% of their time on tasks like](https://www.oecd.org/content/dam/oecd/en/publications/reports/2019/06/data-analytics-in-smes%5F1535d46b/1de6c6a7-en.pdf?ref=blog.datahut.co) finding, cleaning, validating, and combining data from different systems instead of actual analysis. In a 40-hour workweek, that means each person spends 32 hours on tasks that don’t involve actual analysis every week. For a whole data team, this inefficiency adds up and leads to: - Slower analytics and AI workflows - Higher engineering and cloud costs - Reduced agility for pricing, assortment, and competitor intelligence teams - Increased compliance risk - Lower reliability across downstream models By 2026, the gap between what companies need from data and how ready that data is will continue to grow. Websites are getting more complex and better protected, so picking the right web scraping company is more important than ever. Today’s web platforms use dynamic rendering, React interfaces, virtualization, ongoing UI changes, and strong anti-bot systems like [Cloudflare](https://www.cloudflare.com/learning/bots/what-is-bot-management/?ref=blog.datahut.co), PerimeterX, Kasada, and DataDome. Meanwhile, companies must also meet higher standards for [GDPR](https://gdpr.eu/what-is-gdpr/?ref=blog.datahut.co), CCPA, DMA, and internal governance. In this environment, the right [web scraping service](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/) becomes a strategic partner instead of just another tool. This guide reviews the Top 10 Web Scraping Companies in 2026, looking at reliability, compliance, delivery quality, scalability, and value for enterprise teams. The analysis is neutral, but we point out where Datahut excels, especially in enterprise web scraping, compliance, and long-term data operations. ## 1\. Why Choosing the [Right Web Scraping Company](https://www.blog.datahut.co/post/how-to-choose-the-best-web-scraping-service/) Matters in 2026 Web scraping in 2026 is very different from just three years ago. Websites now use: - [JavaScript-heavy front-ends](https://react.dev/learn?ref=blog.datahut.co) (React, Next.js, Vue, Angular) - Server-side rendering combined with hydration - Infinite scrolling, dynamic pagination, and lazy loading - Automated experimentation platforms shipping layout changes daily. - Anti-bot systems that fingerprint browser behavior, TLS signatures, and request metadata - Login walls, paywalls, and personalization At the same time: - AI and LLM teams need massive amounts of structured, accurate training data. - Retailers need real-time [competitor insights across thousands of SKUs](https://www.blog.datahut.co/post/how-popular-price-comparison-websites-grab-data/) - Marketplaces need continuous monitoring of supply, pricing, and seller behavior. - Compliance and data governance teams demand higher transparency. - Internal scraping teams are expensive to hire, retain, and maintain Choosing the right partner for enterprise web scraping leads to: ![Wrong vendor vs right vendor for web scraping](https://www.blog.datahut.co/content/images/2026/07/img-227.png-1.webp) This guide is based on public information, customer feedback, and industry trends. Each company listed has its own approach to web data extraction, ranging from fully managed services to API tools and large proxy networks. The goal is not to rank these companies, but to help businesses find the model that best fits their needs. ## 4\. Best Web Scraping Services in 2026 Below is the updated list of the best web scraping companies and web scraping services for 2026. ![Below is the updated list of the best web scraping companies and web scraping services for 2026.](https://www.blog.datahut.co/content/images/2026/07/chart_1-5.webp) ### 4.1 [Datahut](https://www.datahut.co/?ref=blog.datahut.co) ### Years in Business: 15+ years [Datahut is a fully managed, enterprise-level web scraping service](https://www.datahut.co/solutions?ref=blog.datahut.co) for teams that want clean, compliant, ready-to-use datasets without running their own scraping infrastructure. Unlike API-first vendors, Datahut creates custom extraction pipelines for each client. This approach brings higher accuracy, better success rates, and stronger compliance, especially on complex, dynamic, or well-protected websites. Since Datahut doesn’t sell scraping APIs, anti-bot systems can’t reverse-engineer the scraping logic or test public endpoints. Clients get only the data, not the scraping layer, which makes the pipelines harder to detect and block. - Enterprises needing reliable recurring datasets - Retail & ecommerce operations - Marketplaces with large SKU catalogs - Real estate, travel, finance, and alt-data workflows - AI teams need clean, structured training datasets - Companies with strict compliance/governance requirements Strengths - Zero engineering overhead for clients - Custom pipelines designed for complex, changing websites - High accuracy through multi-layer quality checks - Strong GDPR/[CCPA complianc](https://oag.ca.gov/privacy/ccpa?ref=blog.datahut.co)e posture - High reliability with predictable delivery - Excellent for long-term recurring pipelines - Expertise in web scraping for ecommerce - Proven track record with Fortune 500 and large global enterprises Weaknesses - Not suitable for hobbyists - No instant API access - Requires project scoping before onboardingg ### 4.2 [Oxylabs](https://oxylabs.io/?ref=blog.datahut.co) Years in Business: \~10 years Oxylabs is one of the most established infrastructure-focused web scraping companies, offering a powerful suite of scrapers and a large proxy network. It is built for organizations that need to run web data collection at a serious scale and want fine-grained control over how scraping is executed. Oxylabs is best for teams with in-house developers who want strong building blocks like residential and mobile proxies, SERP APIs, and web scrapers, rather than a fully managed, hands-off service. Best Use Cases - Large-scale price and product monitoring - Search engine data collection (SERP) - Market intelligence and competitive benchmarking - High-volume, always-on scraping workloads Strengths - Massive scalability for very large workloads - AI-powered block avoidance and advanced anti-bot tooling - Global residential, mobile, and datacenter proxy coverage - Multiple products (proxies, SERP APIs, scrapers) under one roof Weaknesses - Can be expensive for smaller teams or experiments - An API-first model means your developers still own selectors and data modeling. - Not a fully managed, end-to-end data delivery service by default ### 4.3 [Zyte (Scrapinghub)](https://www.zyte.com/?ref=blog.datahut.co) Years in Business: \~15 years (Scrapinghub) Zyte (formerly Scrapinghub) is a developer-focused web scraping platform known for its Smart Proxy Manager, Zyte API, and Smart Crawler. It focuses on reliability and data quality, especially for teams that want to centralize crawling logic and let the platform handle rendering and blocking. Zyte is a strong choice for engineering teams that want powerful scraping capabilities but are comfortable defining spiders, extraction rules, and handling data flow themselves. Best Use Cases - Complex JS-heavy websites where reliability matters - Long-running crawlers that need stability over time - Teams that value a strong ecosystem (libraries, spiders, support) Strengths - Smart automated crawling through Zyte API and Smart Crawler - Built-in JavaScript rendering to handle modern front-ends. - Strong focus on structured, clean, well-formatted data - Mature ecosystem, documentation, and community tools Weaknesses - Requires solid technical skills to get the most out of the platform - Managed service tiers exist, but are more expensive and not the default. - Still largely a DIY model where your team designs and maintains spidersders ### 4.4 [ScraperAPI](https://www.scraperapi.com/?ref=blog.datahut.co) Years in Business: \~7 years ScraperAPI is a plug-and-play web scraping API designed for developers who want to focus on parsing content rather than managing proxies, blocks, or CAPTCHA. You send a URL; ScraperAPI returns the HTML, handling most of the underlying complexity for you. It’s best for small to mid-sized teams that want a quick, minimal-friction way to add scraping into their applications without running a full scraping stack. Best Use Cases - Rapid prototypes and MVPs - Internal tools needing occasional scraping - Low-to-mid complexity sites with moderate protection Strengths - Very easy to integrate with a simple URL-based API - Automatic proxy rotation, retries, and CAPTCHA handling - Good pricing entry point for mid-level usage - Helpful for teams that don’t want to run their own proxy infrastructure Weaknesses - A general-purpose approach can struggle with very complex, highly protected, or heavily dynamic sites. - An API-only model means no managed data delivery or ownership of custom pipelines. - Costs can scale up quickly on high-volume or very frequent workloads. ### 4.5 [Smartproxy (DecoDo)](https://decodo.com/?ref=blog.datahut.co) Years in Business: \~7 years (DecoDo) Smartproxy became known as an affordable, reliable proxy provider and has expanded into scraping solutions. It’s aimed at users who want cost-effective proxies and simple scraping tools without investing in heavy infrastructure. It’s well-suited for agencies, small businesses, or data teams that have modest scraping needs but still require legitimate proxy networks and basic tools. Best Use Cases - Light-to-moderate price tracking or content monitoring - Small-scale SEO data collection - Experiments and proof-of-concept scraping Strengths - Budget-friendly pricing, especially for smaller workloads - Easy onboarding and simple dashboards - Good balance between quality proxies and cost Weaknesses - Entry-level plans can be limited in terms of bandwidth and features. - Scraping tools are less advanced than dedicated, enterprise scrapers. - No true managed, professional-services layer for end-to-end data delivery ### 4.6 [Apify](https://apify.com/?ref=blog.datahut.co) Years in Business: \~10 years Apify is an automation-focused platform that combines web scraping, browser automation, and workflows using its "Actors" model. Each Actor is a reusable bot that can log in, navigate, extract, transform, and export data in one process. Apify is a great fit for technically inclined teams that want more than just raw HTML—they want to automate entire business processes involving the web. Best Use Cases - Multi-step workflows (login → search → paginate → scrape → export) - Complex browser interactions (forms, filters, dashboards) - Building reusable scraping automations for multiple stakeholders Strengths - Powerful combination of scripting + automation + scraping - Browser rendering support for complex and interactive sites - A large marketplace of ready-made Actors for common tasks - Flexible for both one-off jobs and recurring workflows Weaknesses - Non-technical users may find the Actor model and JS/Node environment challenging. - Monitoring, scaling, and maintaining Actors is still your team’s responsibility. - Managed/consulting help is available, but not the default starting point. ### 4.7 [Bright Data](https://brightdata.com/?ref=blog.datahut.co) Years in Business: \~11 years Bright Data is a well-known enterprise web data platform that combines a large proxy network with advanced unblocking technology and specialized data collection products. It’s built for organizations that see web data as a key part of analytics, AI, and decision-making. With a strong emphasis on compliance, governance, and control, Bright Data is ideal for teams looking to build large-scale scraping and data-gathering infrastructure on top of a mature platform. Best Use Cases - Large-scale market and competitive intelligence - Global price and assortment tracking across many sites - AI training datasets requiring broad, diverse web data Strengths - Huge proxy ecosystem across residential, mobile, and datacenter IPs - Web Unblocker and other tools for difficult, highly protected websites - Strong compliance, KYC, and governance framework for enterprise buyers - Wide range of specialized products (SERP, datasets, proxies, collectors) Weaknesses - Pricing is higher than many alternatives, especially for smaller users. - Product- and infra-first model; managed end-to-end data delivery is not the default - Best value is realized only when you have a capable internal engineering team. ### 4.8 [PromptCloud](https://www.promptcloud.com/?ref=blog.datahut.co) Years in Business: \~16 years PromptCloud is an established managed web scraping company that delivers structured datasets at scale for enterprises. Instead of just providing tools and APIs, PromptCloud acts as a data partner: you tell them what you need, and they deliver it regularly. This makes it attractive for organizations that know their requirements and prefer predictable, SLA-backed deliveries over building internal scraping capabilities. Best Use Cases - Recurring catalog, pricing, or listings data across many sites - Enterprises that want stable feeds instead of tools - Long-term, fixed-structure data projects Strengths - Strong orientation toward managed, customized data pipelines - Reliable long-term delivery with clear formats and schedules - Good fit for enterprises that want a “set it and run” relationship Weaknesses - Limited self-serve or low-commitment options for quick testing - Less flexibility for highly experimental or fast-changing use cases - Engineering teams that want to tinker or iterate rapidly may feel constrained. ### 4.9 [WebScrapingAPI](https://www.webscrapingapi.com/?ref=blog.datahut.co) Years in Business: \~6 years WebScrapingAPI is a fast-growing web scraping API platform for developers who want reliable rendering, proxy rotation, and anti-bot handling without building their own scraping setup. It aims to make complex scraping tasks easier and offers higher success rates on modern, JavaScript-heavy websites. Compared to older scraping APIs, WebScrapingAPI focuses on speed, scalability, and ease of integration, making it a strong alternative to ScrapingBee. Best Use Cases - Small to mid-scale scraping operations - JavaScript-rendered pages using headless browsers - SEO monitoring, SERP tracking, and content extraction - Internal dashboards and automation tools Strengths - Built-in headless browser support - Automatic proxy rotation and CAPTCHA handling - Simple API interface with quick onboarding - Good success rates on dynamic pages - Scales well for developers who need reliability without complexity Weaknesses - Not ideal for very large enterprise scraping programs - No fully managed or custom pipeline service tier - Teams must still write selectors, logic, and validation flows. ### 4.10 [Crawlbase (ProxyCrawl](https://crawlbase.com/?ref=blog.datahut.co)) Years in Business: \~8 years (ProxyCrawl) Crawlbase, formerly known as ProxyCrawl, offers a straightforward scraping and crawling API for developers who want to fetch web content with minimal friction. It focuses on reliability and simplicity at an accessible price point. It’s a good match for smaller teams, indie developers, and internal tools that need basic but dependable scraping capabilities. Best Use Cases - Basic HTML extraction from content or listing pages - Small to mid-scale internal automation and monitoring scripts - Projects where budget and simplicity are more important than deep customization Strengths - Affordable and easy to start using - Simple API surface for common scraping tasks - Handles many standard blocking and retry scenarios for you Weaknesses - Not designed for highly dynamic, heavily protected websites at large scale - No managed or custom pipeline services — you own the full data flow. - Less feature-rich than some larger enterprise scraping platforms ## 5\. Before You Choose a Web Scraping Partner, Ask Yourself These Questions Choosing the wrong vendor is more than just an inconvenience. It can create data blind spots, revenue loss, and growing operational risks that often go unnoticed. Here are the questions every enterprise should ask before selecting a web scraping company: 1\. What happens if a critical crawler breaks during peak business hours? Do you have a vendor who takes responsibility, or does your engineering team suddenly have to handle the emergency? 2\. If a website ships a layout change tonight, will your data pipeline survive tomorrow morning? Or will you wake up to empty dashboards, stale feeds, and missing data? 3\. How confident are you that your vendor can stay undetected on complex, JS-heavy, anti-bot-protected sites? If your data stops because your vendor is blocked, what is the real cost to your business? 4\. Who is liable if your current scraping setup violates GDPR, CCPA, or DMA guidelines? Do you have a partner focused on compliance, or just a tool that leaves you with the risk? 5\. How much engineering time will your team lose every month maintaining scripts, proxies, selectors, and retries? And what could your teams achieve if that time were freed? 6\. If you’re monitoring thousands of SKUs across marketplaces, what happens when product pages become more dynamic or personalized? Can your vendor consistently keep up — or will your data slowly degrade without you noticing? 7\. Are you paying for a tool… or for predictable outcomes? If you are still responsible for selectors, logic, QA, and monitoring, is it really a “service”? 8\. If your competitor invests in better, fresher, cleaner data, how long until that advantage shows up in pricing, assortment, or SEO rankings? Are you comfortable being the one with the slower, noisier data? 9\. Does your vendor give you clean, structured, analysis-ready datasets — or just raw HTML? If you’re still doing the heavy lifting, you’re not getting the value you should. 10\. What is the cost of being wrong? Broken pipelines, incorrect pricing data, stale competitor feeds, or unreliable training datasets can cost far more than the price of the vendor. ## Final Thoughts: The Cost of Bad Data Is Always Higher Than the Cost of Good Data The web is now the world’s largest, fastest-changing data source — and the companies that master it win. But companies that underestimate how complex it is often pay the price without realizing it: - The wrong prices were pushed live. - The wrong assortment decisions were made. - The wrong competitors are monitored. - The wrong training data feeds AI models. - The wrong signals are driving multi-million-dollar outcomes. These are not hypothetical risks. These problems happen when data is incomplete, outdated, unreliable, or blocked. By the time the impact appears in dashboards or revenue, it’s often too late. That’s why choosing the right web scraping partner is no longer just a technical decision. It’s a strategic moat. Each top web scraping company in this guide offers something valuable, such as proxies, APIs, automation, or managed services. In 2026, [the companies that succeed will be those who build on better data](https://www.blog.datahut.co/post/build-vs-buy-web-scraping-in-2026-the-definitive-guide-for-data-teams/), faster, not those still fixing crawlers or guessing what the market is doing. ## Ready to Stop Fighting Scrapers and Start Scaling With Clean, Reliable Data? If you want to avoid expensive crawler failures, compliance risks, and inaccurate datasets—and instead get fresh, structured data delivered the way your business needs it—let’s talk. Get a Free Data Strategy Call With a [Datahut](https://www.datahut.co/?ref=blog.datahut.co) Expert We will review your use case, look at your current data flow, and show you how a fully managed pipeline can remove engineering overhead and help you make better decisions. ## Frequently Asked Questions (FAQs) ### 1\. How do I choose the right web scraping company? Choosing the right web scraping company depends on your requirements, such as data volume, website complexity, compliance needs, and whether you prefer a fully managed service or developer-focused APIs. Enterprises often prioritize reliability, structured data delivery, and compliance support. ### 2\. What is the difference between a managed web scraping service and a scraping API? A managed web scraping service handles the entire process, including extraction, cleaning, validation, and delivery of structured datasets. A scraping API typically provides raw HTML or partially processed data, and your team is responsible for parsing, monitoring, and maintaining the pipeline. ### 3\. Are web scraping services legal to use? Web scraping is legal in many jurisdictions when done responsibly and in compliance with website terms, data privacy laws such as GDPR or CCPA, and ethical data collection practices. Businesses should work with vendors who prioritize compliance and governance. ### 4\. Why is web scraping more difficult in 2026 than before? Modern websites use dynamic JavaScript frameworks, continuous UI changes, login walls, and advanced anti-bot systems. These technologies make scraping more complex and require sophisticated infrastructure, browser automation, and monitoring. ### 5\. What industries benefit most from web scraping services? Industries that rely heavily on web data include e-commerce, travel, real estate, financial services, market research, and AI development. These sectors use web scraping for pricing intelligence, competitor monitoring, trend analysis, and training datasets. ### 6\. What is the best web scraping service for ecommerce? The best web scraping service for ecommerce depends on your business goals, scale, and data requirements. Companies that need reliable product, pricing, inventory, review, and marketplace data often prefer managed web scraping providers that deliver clean, structured datasets rather than raw scraped data. Key factors to consider include data accuracy, scalability, anti-bot handling, compliance practices, delivery formats, and ongoing support. For enterprises, a fully managed solution can significantly reduce the operational burden of maintaining scraping infrastructure and adapting to website changes. ### Scrape Macy’s Sale Section Using Python & Playwright URL: https://www.blog.datahut.co/post/how-to-scrape-macy-s-sale-section/ Last updated: 2026-07-23T07:48:30.000Z When you think of Macy’s, it is not just a department store—it is an American retail legacy established in 1858\. What began as a humble store located in Manhattan has become one of the largest and arguably best-known retail department stores in America. Macy's currently operates in hundreds of locations across the mainland United States, including its world-renowned flagship store located at Herald Square, New York City— a retail space spanning over a million square feet. Macy's is not simply known for its variety; it is also known for the full retail experience. From casual everyday wear to upscale items for special events, Macy’s has something for everyone. With many decades of retail experience behind them, the Macy's brand is compelling to shoppers for essential everyday styles and current fashion trends. ## Scrape Macy’s Sale Section: Data Collection The data collection phase for Macy's clothing consisted of 2 steps: first, obtaining all of the product page links, and then structuring the data for deeper analysis. We narrowed our data collection to women's clothing, then focused further on their "Sales and Clearance" section. ### Step 1: Collecting Product Links from Macy’s “Sales & Clearance” Section To kick off our scraping project, our first goal was to collect all the product links from the women’s clothing sale and clearance section on Macy’s website. Now, like many online shopping sites, Macy’s doesn’t show all its products on a single page. Instead, the items are spread across many pages, and you need to click a “Next” button to see more. That means we couldn’t just grab everything in one go—we had to visit each page one by one. To handle this, we used a tool called Playwright. Think of Playwright as a helper that can use a browser just like a real person would. It opens the website, waits for everything to load properly, and then scrolls down the page naturally so all the products become visible. Since Macy’s uses pages instead of endless scrolling, Playwright clicks the “Next” button to move from one page to the next. As it moves through the pages, Playwright also checks for any pop-up messages—like cookie permissions or sales promotions—and closes them so they don’t block the view or interrupt the scraping. This helps make sure we don’t miss any product links. Once the product links are gathered from a page, they’re saved into an SQLite database along with their category. This way, the data stays neat and easy to manage later. Using a database instead of just saving the links to a basic file makes it easier to search, filter, and analyze them in the future. We also made sure not to store the same link more than once, which helped us avoid duplicates. By the end of this step, the scraper had collected unique product URLs. That gave us a strong foundation to move on to the next part—gathering more detailed information about each item, like names, prices, colors, and any discounts they might have. ### Step 2: Extracting Information from Product Pages Once we’ve saved all the product links from Macy’s women’s clothing sale and clearance section into our SQLite database, the next step is to visit each product page one at a time. But we’re not just collecting links anymore—we’re going deeper. Now, we want to gather all the important details about each item, like the product’s name, its current price, the original price (or what it cost before the sale), how many reviews it has, the brand, the discount offered, available sizes, the SKU number, whether it’s in stock, and even the product description. Now, once we’ve collected this product data, we need to store it somewhere smart. For this task, we’re using MongoDB. It’s a type of database that’s very flexible and can handle different kinds of data easily—perfect for the variety of details we’re collecting. And because scraping large websites can sometimes crash or pause unexpectedly, we’re also saving a copy of the data in backup files called JSONL files. These act as a safety net. If the scraper stops halfway through for some reason, we won’t have to start all over again—we can just pick up where we left off using the saved data. In short, we’re blending smart automation, careful data gathering, and reliable backup methods to collect detailed product information from Macy’s sale section smoothly and efficiently. ## Data Cleaning When you first collect data from a Macy’s, it doesn’t always come in neat and clean. You’ll often find weird symbols in the prices, missing values (like empty cells), or information that’s all over the place. For example, some items might not have a discount or a proper description — in those cases, it's a good idea to replace those blanks with "N/A" so that it's clear there's no data, rather than leaving it empty. Also, the original price sometimes includes a lot of extra symbols, even unwanted text. One of the first things done was remove all those unnecessary symbols so the price looks clean and easy to read — just like how you'd snip off price tags before wearing new clothes. Now, to actually do the cleaning, there are some tools that make life a whole lot easier. One of them is OpenRefine. Don’t worry if you’ve never heard of it — it’s kind of like a super-powered version of Excel. You can use it to remove duplicates, fix inconsistent data (like different spellings of the same brand), and tidy things up with just a few clicks. It’s great for visual learners and super easy to pick up. But sometimes the mess can be a little messier; for example, what if the data has some messy HTML tags, or if the date formats are all weird? That's when Python comes in — and more specifically, a really valuable library in Python called pandas. Pandas allows you to clean and manipulate your data in a more organized way. Think of it like a tiny robot assistant who already knows how you want it all sorted out for you behind the scenes. You know, cleaning your data may not sound very exciting, especially at first, but it is absolutely essential if you want to do anything useful with your project! And once your data is cleaned up, everything else becomes easier and more pleasurable! ## Advanced Tool-sets and Libraries for Efficient Data Scraping In order to scrape a large volume of data from websites in an efficient manner, it is vital you use the proper combination of tools and libraries. This project employs a few powerful Python libraries and modules that work in unison to automate the browsing, data extraction, data storage, and data handling processes in a coherent and reliable way. The first library we are going to take a look at is asyncio. This library allows a program to asynchronously run tasks. Instead of waiting for a task to complete before calling the next task, asyncio allows many of the operations to be proceeding, while pending, at the same time. This is crucial for web scraping because there are a lot of tedious interactions on websites that generally slow down the data collection process. Instead of blocking the entire program during each web page interaction, asyncio allows you to manage multiple web page interactions without blocking your entire program. The scraping automation task is handled by playwright.async\_api. This is a modern browser automation tool that gives you programmatic control over browser instances such as Firefox or Chromium. With [playwright](https://www.lambdatest.com/playwright?ref=blog.datahut.co), you simply specify a web page to open and the library will take care of opening the web page, simulating user actions such as scrolling down a page or clicking buttons, and even take the HTML content away upon completion of those actions. The async refers to playwright's ability to work well in an asyncio library giving it the ability to automate the scraping process more efficiently than in a typical synchronous way. To ensure the browsing looks more human-like and to decrease the chance of being blocked by a website, this project is using playwright\_stealth, a library specifically designed to hide automation fingerprints. Many websites have anti-bot strategies in place to detect automation and scripts, but playwright\_stealth will use the same actions as a typical user would, and modifies properties and behaviors in the browser so it appears as if a regular user is browsing the site. Data storage is another key area. This project will be using sqlite3 to save the scraped product URLs into a simple, local database. SQLite can be simple to set up and use and has no requirement for a server to be created or set up making it an easy way to manage and query URLs during the scraping workflow. As a more currency specific data storage option, for larger data stores or more complicated data storage, pymongo is also included with this project as a way to connect to MongoDB, a NoSQL database particularly used for handling unstructured data. MongoDB is designed very well for saving product information with many detailed fields that may have a non-consistent format such as descriptions, pricing, reviews, etc. MongoDB will allow for flexible queries and is easily extensible. A number of other standard Python libraries also facilitate the entire process: logging provides an overall understanding of what the scraper is doing by logging significant events or errors that can be helpful during debugging and monitoring. It is important to have the random and time libraries for adding some delay and a little randomness between actions, thereby helping to normalize organic browsing activities and to avoid triggering a web site’s anti-bot protection systems. The datetime and pathlib libraries assist in managing file paths and file creation with timestamps to ensure timely organization during data storage. ## STEP 1: Extracting Product URLs from Macy's Women's Clothing Section ### Importing Libraries ``` import asyncio import logging import sqlite3 from playwright.async_api import async_playwright from playwright_stealth import stealth_async ``` Let’s begin by setting the foundation for our web scraping project. When you're trying to collect product details from large e-commerce websites, it's important to use the right tools that can handle the job smoothly and efficiently. That’s exactly what we’ve done in this script. We start by bringing in a tool called asyncio. Think of it like a multitasking manager—it helps the script open multiple pages and collect data at the same time, rather than doing one thing after another. This makes the process much faster and more efficient, especially when dealing with hundreds or thousands of product pages. Next, we use something called logging. This is like keeping a diary of what the script is doing. It records important moments, like when a page is opened or when an error happens. When you're scraping a lot of data, having this kind of record is really helpful to understand what's going on and where things might go wrong. We also use [sqlite3](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co), which allows us to save the product information into a local database file. You can think of this as a storage box where every piece of scraped data is kept safely for later use. Another key tool we include is async\_playwright from the Playwright library. This helps us control a browser through code, letting us visit pages and interact with them as if we were browsing manually—but much faster and more consistently. To avoid getting blocked by websites, we use stealth\_async. This adds a bit of camouflage to our script, making the browser look more like a real person is using it. It imitates human-like behavior, which helps keep the script under the radar. Together, all these tools create a strong starting point for our scraping process. They help us move quickly, stay organized, save data properly, and avoid drawing attention—just like a well-planned mission. ### Tracking the Scraper’s Activity with a Clean Logging Setup ``` # Setup logging logging.basicConfig(    filename="macys_scraper_url.log",    filemode="w",    format="%(asctime)s - %(levelname)s - %(message)s",    level=logging.INFO, ) """ This block sets up logging, which is useful for tracking the script's progress, errors, and important events. Logs are saved in a file named 'macys_scraper_url.log'. """ ``` When you're writing a script that needs to gather thousands of product links spread across multiple web pages, it’s not just about writing code that works. It’s equally important to keep track of what the script is doing while it runs. Think of it like watching over a delivery truck—you don’t just send it out, you also want updates on where it’s been and if anything went wrong. This is where logging comes in. In the code, a simple but effective logging system is set up using Python’s built-in logging module. This system writes updates to a file named macys\_scraper\_url.log. So, each time something happens—like successfully collecting a link, skipping an item, or running into an error—it’s recorded in this file along with the exact time it occurred. The log file is cleared and refreshed every time the script starts, thanks to a setting called filemode="w". This ensures that the log always reflects the most recent run and doesn’t get cluttered with old data. Each log entry includes a timestamp, the type of message (like INFO for regular updates or ERROR if something breaks), and a short description of the event. By setting the log level to INFO, the script will record all important events without being too noisy. This becomes especially useful during long scraping sessions. If something goes wrong or seems off, you can just check the log file to see exactly what happened and when—almost like reading a travel diary your script kept while it was running. ### Efficiently Storing Product URLs with a Lightweight SQLite Database ``` # SQLite setup “”” Create a SQLite database and table for storing product URLs “”” DB_PATH = "macys_products_url.db" ``` Before we start collecting data from a website, it’s important to have a proper place to store everything we gather. Think of it like setting up a clean folder before you begin a big research project—you want everything to go in the right place from the start. In our case, we’re using a simple and lightweight database called SQLite to save the product links (or URLs) we’ll be scraping from the website. Before we start collecting data from a website, it’s important to have a proper place to store everything we gather. Think of it like setting up a clean folder before you begin a big research project—you want everything to go in the right place from the start. In our case, we’re using a simple and lightweight database called SQLite to save the product links (or URLs) we’ll be scraping from the website. We define a variable called DB\_PATH, which tells our program what the database file is named and where to find it. For example, here it’s named "macys\_products\_url.db". You can imagine this file as a mini-warehouse where all the product URLs are safely stored before we do anything else with them. Now, why use SQLite? The reason is simple—it’s easy to set up, doesn’t require a complicated installation, and can smoothly handle large amounts of data, even if you’re collecting thousands or tens of thousands of links. It keeps everything in one file, which helps avoid confusion and keeps the entire process tidy. With this setup, you’re not just collecting data randomly—you’re organizing it from the very beginning, making your job easier as you move on to the next steps.We define a variable called DB\_PATH, which tells our program what the database file is named and where to find it. For example, here it’s named "macys\_products\_url.db". You can imagine this file as a mini-warehouse where all the product URLs are safely stored before we do anything else with them. Now, why use SQLite? The reason is simple—it’s easy to set up, doesn’t require a complicated installation, and can smoothly handle large amounts of data, even if you’re collecting thousands or tens of thousands of links. It keeps everything in one file, which helps avoid confusion and keeps the entire process tidy. With this setup, you’re not just collecting data randomly—you’re organizing it from the very beginning, making your job easier as you move on to the next steps. ### Reliable URL Storage: Keeping Macy’s Data Clean and Organized ``` # database setup def save_url_to_db(url):    """    Saves a single product URL to the SQLite database.    """       try:        conn = sqlite3.connect(DB_PATH)        cursor = conn.cursor()        cursor.execute("INSERT OR IGNORE INTO product_urls (url) VALUES (?)", (url,))        conn.commit()        conn.close()        logging.info(f"Saved URL: {url}")    except Exception as e:        logging.error(f"Failed to save URL: {url} - Error: {e}") ``` Imagine you’re working with a massive list of items—like Macy’s entire collection of women’s clothing on sale, which includes over 4000 product pages. That’s a huge amount of data, and if you’re trying to collect all those web links (URLs), it’s really important to keep everything tidy and safe. You don’t want to lose links you’ve already collected, and you definitely don’t want to collect the same link more than once. That’s where this part of the code comes in—it helps store each product link neatly and reliably. Here’s how it works: every time the scraper finds a product URL, this function takes that link and saves it into something called a SQLite database. Think of this database like a digital notebook—it stores your data in a table, and makes sure nothing gets written twice. To do this, Python uses a built-in tool called sqlite3, which connects to the database file. When it’s time to save a link, the code uses a special command—INSERT OR IGNORE. This little instruction tells the database, “Only save this if it’s not already there.” So if the same link comes up again later, it’ll be skipped without causing any problems. Once the link is saved (or skipped), the code closes the database connection to keep things running smoothly and not use up unnecessary system resources. It also leaves a note in the log—kind of like a quick journal entry—saying whether the URL was saved successfully or if something went wrong. This whole process makes sure your data stays clean, avoids doubles, and most importantly, lets you pause and resume the scraping later without missing a beat. That’s a big win when you’re working with thousands of entries and multiple runs. ### Automating Macy’s Product URL Extraction with Smart Scrolling and Pagination ``` # Main scraping logic async def scrape_macys():    """    Scrapes product URLs from the Macy's women's clothing sale section.    """           async with async_playwright() as p:        # Launch the Firefox browser in visible mode (headless=False)        browser = await p.firefox.launch(headless=False)        context = await browser.new_context()        page = await context.new_page()               # Use stealth mode to avoid getting blocked as a bot        await stealth_async(page)          # Macy's clothing sale base URL        base_url ="https://www.macys.com/shop/sale/womens-sale/womens-clothing-sale?id=338161"        await page.goto(base_url, timeout=60000)        # Try to close newsletter popup if it appears        try:            await page.locator('button:has-text("No, Thank You!")').click(timeout=5000)            logging.info("Closed newsletter popup")        except:            logging.warning("Newsletter popup not found")        # Try to accept cookies if the cookie banner appears        try:            await page.locator('button:has-text("Accept Cookies")').click(timeout=5000)            logging.info("Accepted cookies")        except:            logging.warning("Cookie banner not found")        current_page = 1        max_pages = 999  # You can limit this or detect dynamically              # Loop through pages until there are no more        while True:            logging.info(f"Processing page {current_page}")                     # Scroll multiple times to load all products (simulate lazy loading)            for i in range(10):                await page.mouse.wheel(0, 2000)                await page.wait_for_timeout(1000)            # Extract product URLs            product_links = await page.locator('div.description-spacing a.brand-and-name').all()            for link in product_links:                href = await link.get_attribute("href")                if href:                    full_url = "https://www.macys.com" + href                    save_url_to_db(full_url)            """            Finds all anchor () tags within a container that has the class 'description-spacing'.            """                       # Try to go to next page            try:                next_button = page.locator('li a.next-button')                               if await next_button.is_visible() and await next_button.is_enabled():                    await next_button.click()                    await page.wait_for_load_state("domcontentloaded")                    current_page += 1                    await asyncio.sleep(3)                else:                    logging.info("No more pages. Exiting.")                    break            except Exception as e:                logging.error(f"Error clicking next page: {e}")                break        # Close the browser after scraping is done        await browser.close() ``` This piece of code represents the main logic for scraping product URLs from Macy's women's clothing sale page in order to create a full list of quantity over 17k product links. The code first launches a headfull web browser using Playwright. Playwright is a modern automation library and since we are running in a headful mode we ease debugging in case the code doesn't function as intended. The website could identify the activity as a bot, therefore a 'stealth mode' is invoked that disguises some of the automated indicators showing the activity is automated rather than done by a human. When the browser opens the Macy's sale page, the code attempts to manage common interruptions such as popups. It attempts to close newsletter signups and accept cookie banners, which are quite common in retail sites that might otherwise interfere with the browsing experience. The primary aspect of the scraping process is loading products dynamically via scroll simulation. The Macy's website uses lazy loading, meaning that products only become available as the user scrolls down the page. The process coded for this, would continually scroll down the page, to allow the Macy's site enough time to load all of the available product listings visible at once. Following the scrolling process, the program will scrape product links, by searching for the specific HTML elements—specifically, the anchor tags found within the appropriate container (with certain classes) that describes the product listings. The program also converts each relative link to an absolute link by attaching Macy's base domain to each relative link, in order to ensure the links would be accessible on their own at a later date. To cover the entire catalog, the script navigates through multiple pages by clicking the “Next” button, continuing the process until no more pages remain or a preset maximum page limit is reached. Each time a new page loads, the same scroll-and-extract routine repeats. This loop ensures that product URLs across all sale pages are collected systematically. Finally, once all pages are processed, the browser closes gracefully to free up system resources. In total, this systematic scraping strategy creates a deep set of product URLs that can be utilized for downstream analysis or to scrape high-level product data like prices, descriptions, and reviews. The use of stealth browsing, popup management, dynamic content loading, and gentle pagination management results in a powerful and efficient solution for extracting huge amounts of data from contemporary, interactive e-commerce sites like Macy's. ### Let the Scraping Begin: Running the Async Workflow ``` # ENTRY POINT TO START SCRAPING asyncio.run(scrape_macys()) """ This line actually starts the entire scraping process by calling the 'scrape_macys' function using asyncio """ ``` Every time we build an asynchronous web scraper, we need a way to kick things off—that's where this line comes in: [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(scrape\_macys()). Think of it as pressing the “Start” button for the entire scraping process. When this line runs, it hands control over to the scrape\_macys function, which holds the main steps for getting data from the website. Now, there’s something special about this function: it’s asynchronous. That just means it doesn't run like a regular function. Instead of going step-by-step in a straight line, it can pause and resume—kind of like a multitasker that knows when to wait and when to jump to the next task. But for that to happen smoothly, it needs something called an “event loop,” and [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)() sets that up for us automatically. By doing things asynchronously, our scraper can work much faster and smarter. For example, while it’s waiting for one page to load, it can go ahead and start grabbing data from another. This avoids wasting time just sitting around for slow responses. So in simple terms, that one line of code is what gets everything moving—it tells our scraper, “Alright, go ahead and start working,” and from there, it takes over and begins handling thousands of product pages efficiently and in order. ## Step 2: Scraping Full Product Information from Individual Links ### Importing Libraries ``` import asyncio import json import random import sqlite3 import logging from time import sleep from datetime import datetime from pathlib import Path from pymongo import MongoClient from playwright.async_api import async_playwright import re ``` The initial lines of code may look like a huge list of imports, but each serves a particular function that is crucial in assisting us to extract and store product information without any issue. We kick things off with asyncio. This little tool is what lets our script handle multiple tasks at the same time. Since scraping information from hundreds or even thousands of product pages can take a while, we use something called asynchronous programming. It’s like giving our script the ability to multitask—doing several things at once—so it doesn’t waste time just waiting around for pages to load. Next comes json. Imagine you want to store product details like names, prices, and descriptions in a way that's neat and easy to read. JSON helps with exactly that—it’s a simple way to organize and move around structured data. Then we add a bit of randomness and short pauses between tasks using random and sleep. This makes our script behave more like a human rather than a robot, which helps avoid getting blocked by the website we’re scraping. After that, we bring in sqlite3, which gives us a lightweight database to store all the product links we collect. The best part? It doesn't need any special setup or server—it just works right out of the box. Alongside that, we use logging, which acts like a diary for our script. If anything goes wrong or doesn’t work as expected, we can check the logs to figure out what happened. To keep track of when things happen, we use datetime. It helps us add timestamps to our saved files or log entries, so we always know when something was done. For dealing with files and folder paths smoothly, we bring in Path from the pathlib library—it makes working with file locations cleaner and simpler. Now, while sqlite3 is great for smaller projects, we also include pymongo for more flexibility. It connects us to a bigger, more powerful database called MongoDB. This is especially useful when we’re dealing with data that doesn’t fit neatly into a table—like product info that changes from one item to another. And then, we have the real workhorse of our scraper: playwright.async\_api. Playwright is the tool that allows our script to behave like a person browsing the Macy’s website. It can scroll, click, and wait for content to appear, just like we would. Since we’re using the async version of Playwright, it fits perfectly with asyncio, making the whole process fast and efficient. Lastly, we import re, which stands for regular expressions. Think of it like a smart search tool—it helps us pull out specific pieces of information from messy HTML pages, like finding just the price or product name in a sea of code. So, all these imports together form a strong foundation for our scraper. They help it move through the website, collect data smartly, store it properly, and make sure everything runs as smoothly as possible—even when we’re working with a massive number of products. ## Filtering the Noise: How is\_valid\_price() Keeps Price Data Clean ``` # HELPER FUNCTION: CHECK VALID PRICE FORMAT def is_valid_price(text):    """    This function checks if the provided text looks like a valid price in INR format.    """    return bool(re.search(r'INR\s[\d,]+\.\d{2}', text)) ``` Let’s talk about a small yet very helpful part of our web scraping process—a little [function](https://www.geeksforgeeks.org/javascript/what-are-the-helper-functions/?ref=blog.datahut.co) called is\_valid\_price(). Even though it’s not flashy or complex, it plays a big role in keeping our data clean and accurate. Its job is pretty straightforward: it checks if a piece of text actually looks like a real price, specifically in the Indian Rupee format—like “INR 1,299.00”. Now, you might wonder why we even need this. When we scrape a webpage, we often collect all kinds of text by accident. Along with actual prices, we might also pick up labels like “SALE” or “NOW”, or numbers that aren't really prices at all. This is where is\_valid\_price() steps in. It helps us separate the real price information from everything else. Behind the scenes, this function uses a technique called regular expression (or regex). Think of regex as a smart filter—it lets us describe the exact pattern we’re looking for. In this case, the pattern is designed to spot prices that match the INR style. So, in simple terms, is\_valid\_price() is like a careful editor. It doesn’t just collect anything that looks like a number—it makes sure we’re only keeping the real deal. That way, when we analyze our data later, we’re working with clean, reliable information instead of a messy pile of random text. ### Scraper Setup Essentials: Databases, User-Agent, and Logging in One Place ``` # CONFIGURATION SQLITE_PATH = "/home/anusha/Desktop/DATAHUT/Macys_clothing/macys_products_url.db" USER_AGENTS_PATH = "/home/anusha/Desktop/DATAHUT/Macys_clothing/user_agents.txt" JSONL_PATH = "macys_products_data.jsonl" MONGO_URI = "mongodb://localhost:27017" DB_NAME = "macys" COLLECTION_NAME = "products" LOG_FILE = "macys_scraper.log" """ Configuration variables for the scraper """ ``` At the very beginning of our script, we define a few important file paths and settings. Think of these like setting up your workspace before starting a task—putting your tools where you can reach them easily. We start with something called SQLITE\_PATH. This is simply the path to a local database file that holds all the product URLs we plan to visit. You can imagine it like a to-do list. Instead of writing URLs on a sticky note, we save them in this database so our script knows exactly where to go when it’s time to collect product details. Then there’s USER\_AGENTS\_PATH. This points to a text file that holds a list of different user agents. Now, in simple terms, a user agent is like a mask that tells a website what kind of device and browser is visiting—whether it’s a Chrome browser on a laptop, or Safari on an iPhone. By switching these user agents, our scraper can blend in and avoid raising suspicion or getting blocked. It’s like changing disguises while exploring a secure building—you just want to quietly gather information without being noticed. Next, we have [JSONL\_PATH](https://docs.python.org/3/library/json.html?ref=blog.datahut.co). This is the file where we’ll save all the data we scrape, and it’s stored in a format called JSON Lines. That just means we store one product’s data per line. This way, the file stays organized, even when we’re collecting information about thousands of products. It also makes it easier to go through later, especially if we want to load or process the data again. Now let’s talk about MongoDB settings. MONGO\_URI is like the home address of our database in the cloud. It tells the script where to send the data. Then we have DB\_NAME and COLLECTION\_NAME. These are like the specific apartment and room where we want to drop off the information. For our case, we’re placing all the Macy’s product data into a database called “macys,” inside a collection named “products.” Finally, there's the LOG\_FILE. This file acts like a behind-the-scenes journal. Every time the scraper does something—whether it visits a URL, saves some data, or runs into an error—it writes it all down in this file. That way, if something goes wrong or we want to check what happened during the run, we can just look at this log. It’s incredibly helpful when troubleshooting or tracking performance over time. By gathering all these settings in one place at the start of our script, we make everything easier to manage. If we want to change something—like use a new database, switch out our user agents, or save data in a different place—we only need to update it here. That’s especially useful when you’re working with massive projects like this Macy’s sale scrape, where we’re handling over 4000 product links. A little organization at the beginning goes a long way in keeping everything smooth and under control. ### Tracking the Scraper’s Activity with a Logging Setup ``` # SETUP LOGGING  logging.basicConfig(filename=LOG_FILE,                    level=logging.INFO,                    format="%(asctime)s - %(levelname)s - %(message)s") """ Configure logging to track the script's execution """ ``` Before we start collecting thousands of product links from Macy’s women’s clothing sale section, it’s important to set up a way to keep track of what our script is doing. Think of it like keeping a journal while on a big trip—you want to remember where you went, what worked well, and where things got tricky. That’s exactly what logging helps us do in our Python script. In this project, we’re using Python’s built-in logging module to record important information while the scraper runs. It tells us when the script is working as expected and, more importantly, when something doesn’t go as planned. To set this up, we use a function called logging.basicConfig(), which lets us decide how our logs should be stored and what kind of messages we want to see. We choose a specific file location (called LOG\_FILE) where all the messages will be saved. This is helpful because, later on, you can open that file and see a detailed record of what happened during the scraping. We also tell the logger to include useful messages by setting level=logging.INFO, which ensures we capture events like which page got scraped or if the script skipped a product. Lastly, we set a format for each message that includes the time it happened, the type of message (like INFO or ERROR), and the message itself. With this setup in place, even if we’re scraping over 4000 pages, we’ll have a clear trail of everything the script did—making it much easier to understand what went right and to fix anything that didn’t. ### Mimicking Real Users: A Simple Trick to Bypass Scraper Detection ``` # LOAD USER AGENTS  with open(USER_AGENTS_PATH) as f:    USER_AGENTS = [line.strip() for line in f if line.strip()] """ Load a list of user agents from a text file: """ ``` When building a reliable web scraper for Macy’s women’s apparel sale section, one important challenge was making sure the scraper didn’t get blocked by the website. Websites often have systems in place to detect when a bot—rather than a real person—is trying to access their content. One common way they catch bots is by noticing if every request comes from the same browser or device again and again. To avoid raising any red flags, we used a simple but effective trick. We gave the scraper the ability to pretend it was using different browsers. This is done using something called a user agent. Think of a user agent as a small piece of information your browser automatically sends when you visit a website. It tells the site what kind of browser you’re using (like Chrome or Firefox) and what kind of device you’re on (like a Windows laptop or an iPhone). So, we created a plain text file that had a list of different user agent strings. Then, using Python, we read this file line by line. We cleaned up the lines—removing any empty spaces or blank lines—and stored the final list in a variable. Now, instead of always using the same user agent, our scraper randomly picks one from this list each time it sends a request. This makes it look like different people are visiting the site, which helps us stay under the radar and avoid getting blocked. ### Preparing the Database Table to Store and Manage Macy’s Product Links ``` # DB SETUP conn = sqlite3.connect(SQLITE_PATH) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_urls (    url TEXT PRIMARY KEY,    processed INTEGER DEFAULT 0 ) """) conn.commit() """ Set up SQLite database connection and ensure the product_urls table exists """ ``` When you're trying to collect thousands of product links from a website, things can get messy really fast if you don’t have a proper system to keep track of what you've already collected. That’s where using a small, local database like SQLite comes in handy—it’s lightweight, fast, and doesn’t need any complicated setup. In the code we’re looking at, the first thing it does is open a connection to a database file (that’s what the SQLITE\_PATH is for—it simply tells the program where the database is saved). Once it’s connected, it checks whether a table named product\_urls already exists inside that database. If the table isn’t there yet, it creates one with two simple columns: one called url and another called processed. The url column is used to store the actual product links. It’s also marked as the “primary key,” which is just a way of saying “no duplicates allowed”—if the same URL shows up twice, only the first one will be saved. The second column, processed, acts like a little flag. It keeps track of whether a link has been scraped or not. Every new link starts off as “not yet processed” (which is marked with a zero), and once the scraping is done, the status can be updated. This setup is really helpful, especially if you’re working with a massive list. Having this database means you can stop the process midway and pick up where you left off later, without losing track or doing the same work again. Once the table is created and ready, the code saves the structure so that it’s always available during the scraping process. ### Keeping It Organized: Retrieving Unprocessed Product URLs ``` # FETCH UNPROCESSED URLS  def get_unprocessed_urls():    """    Retrieve URLs from the database that haven't been processed yet.    """    cursor.execute("SELECT url FROM product_urls WHERE processed = 0")    return [row[0] for row in cursor.fetchall()] ``` When you're working on a task like product scraping—basically collecting information from different product pages—it’s important to keep track of where you’ve been and where you still need to go. Otherwise, you might waste time scraping the same page over and over, or miss some pages altogether. That’s where the get\_unprocessed\_urls function comes in. Think of it like a checklist that helps your scraper know which links are still waiting to be visited. This function looks inside a small local database, stored using SQLite, which keeps a list of all the product URLs. Each URL in the database has a label attached to it—a flag that says whether it has been processed or not. If the flag is set to 0, it means the scraper hasn’t touched that URL yet. So, when this function runs, it quickly grabs only those untouched links from the database and hands them over as a neat list. By using this approach, the scraper can pick up right where it left off—even if it was stopped in the middle for some reason. It avoids redoing the same work and keeps things running smoothly, especially when scraping thousands of pages. It’s a simple trick, but incredibly useful for staying organized and efficient during big scraping jobs. ### Marking URLs as Done to Keep Scraping Efficient ``` # Mark a URL as processed def mark_url_processed(url):    """    Mark a URL as processed in the database.    """       cursor.execute("UPDATE product_urls SET processed = 1 WHERE url = ?", (url,))    conn.commit() ``` Once a product page has been successfully scraped, the next thing we want to do is make sure we don’t scrape the same page again. That’s exactly what the mark\_url\_processed function helps us with. Think of it like checking off a task on a to-do list. After we’ve collected the product’s details, this function steps in and updates our local SQLite database to say, “Hey, we’ve already taken care of this one.” It does this by switching a flag in the database from 0 to 1, which simply means the URL has been processed. This small action plays a big role in keeping everything running smoothly. Imagine trying to go through a long list of product pages, but you keep landing on the same ones over and over—it would waste time and energy. That’s what we’re avoiding here. By keeping track of what’s done, the scraper can move forward efficiently without repeating steps. And here’s another benefit: if the script suddenly stops or crashes for some reason, it can pick up right where it left off. Since we’ve marked the completed URLs, there’s no need to start all over again. It’s like leaving a bookmark in a long book—so the next time you open it, you know exactly where to continue. ### Saving Scraped Data the Smart Way with JSONL ``` # SAVE DATA def save_to_jsonl(data):    """    Save product data to a JSONL (JSON Lines) file.    """    with open(JSONL_PATH, "a", encoding="utf-8") as f:        f.write(json.dumps(data) + "\n") ``` When you're collecting product information through web scraping—especially when there are thousands of items and new data coming in all the time—it’s really important to save that information in a smart and organized way. That’s where the save\_to\_jsonl function comes in. This function helps by writing each product’s details into a file called a JSONL file. If you’re not familiar with it, JSONL (which stands for JSON Lines) is just a file format where each line holds one complete product entry, written in a format that computers can easily understand. Think of it like adding new pages to a notebook—each product gets its own line, and nothing gets erased when new information is added. That’s because the function opens the file in “append mode,” which means it simply adds to the end of the file instead of starting over. It also makes sure the product details, which come in as Python dictionaries, are properly turned into neat little JSON strings before saving them. This approach makes life a lot easier when you're dealing with large amounts of data—it keeps things clean, easy to work with, and reliable. Even if something goes wrong in the middle of scraping, you don’t lose the earlier data, and you can pick up right where you left off. That’s the kind of system you want when your project grows bigger and needs to handle more data without breaking. ### Making Data Storage Simple and Scalable with MongoDB ``` # Save product data to MongoDB def save_to_mongo(data):    """    Save product data to MongoDB.    """    client = MongoClient(MONGO_URI)    db = client[DB_NAME]    collection = db[COLLECTION_NAME]    collection.insert_one(data) ``` Once we’ve pulled out the necessary details from a product’s webpage—like its name, price, or brand—we need a place to store that information safely for later use. That’s where the save\_to\_mongo function comes in. Think of it like a digital filing cabinet. This function helps us store all the product data neatly into something called MongoDB, which is a type of database used to keep information organized and easy to access. Whenever we gather new product data, this function steps in to save it in MongoDB as a “document”—which is just a structured way of storing information, similar to a filled-out form. It knows exactly where to put the data, connecting to the correct part of the database, known as a collection. Each time we use the function, it adds a new document with the latest details. This setup isn’t just about keeping things safe—it also makes our future work easier. Want to look up only discounted products? Or filter by brand or price? With the way the data is stored, you can run such searches quickly and easily. And because this function works in a self-contained way—meaning it does its job without needing help from other parts of the code—it stays simple and efficient. Over time, this small but powerful piece becomes a helpful tool in handling large scraping projects, keeping all our data clean, searchable, and ready for deeper analysis. ### How Random Wait Times Help Avoid Bot Detection ``` # RANDOM WAIT  def wait_random_delay():    """    This function introduces a random delay between requests to make the   scraping patterns are less predictable and reduce the risk of being blocked.    """    delay = random.uniform(5,10)    logging.info(f"Waiting for {delay:.2f} seconds...")    sleep(delay) ``` When you're scraping data from a website, it's important not to make your script act too much like a robot. Websites are pretty smart these days—they can spot unnatural behavior, like clicking too fast or loading pages without any pause. That’s where a bit of randomness can help make things look more human. To do this, we use something called the wait\_random\_delay() function. It simply pauses the script for a random amount of time—say, somewhere between 5 and 10 seconds—before moving on to the next step. This small pause gives the impression that a real person is browsing, taking a moment to read or look around before clicking again. It’s a bit like mimicking how you or I might casually scroll through a website. This random wait isn’t just for show—it also gives the website’s servers a breather, especially if we’re collecting data for a long time. That means we’re being more polite to the site and less likely to get blocked or flagged as a bot. To keep track of what’s happening behind the scenes, the script logs how long each delay lasts. That way, if something seems slow or gets stuck, we can look back and understand what happened. These kinds of thoughtful touches may seem small, but they really do help. They make your scraping more stable, more respectful, and more likely to succeed without drawing unwanted attention. ### How the Product Data Scraper Gathers and Organizes Detailed Macy’s Product Information ``` # SCRAPE FUNCTION  async def scrape_product_data(page, url): """ Scrape only: name, brand, mrp, discount from a Macys product page. """ try: await page.goto(url, timeout=60000) await page.wait_for_timeout(3000) # Accept cookie popup """ Handle cookie consent popup if it appears. """ try: await page.wait_for_selector("#onetrust-accept-btn-handler", timeout=7000, state="visible") await page.locator("#onetrust-accept-btn-handler").click() logging.info("Clicked Accept Cookies") await page.wait_for_timeout(2000) except Exception: logging.info("No cookie popup detected.") # Check Access Denied if "Access Denied" in await page.content(): logging.warning(f"Access Denied: {url}") return None # Safe getter helpers async def safe_get(selectors): if isinstance(selectors, str): selectors = [selectors] for selector in selectors: try: text = (await page.locator(selector).inner_text()).strip() if text: return text except: continue return None # Extract only 4 fied name = await safe_get("span.body") brand = await safe_get(".updated-brand-label > a:nth-child(1)") mrp = await safe_get(".extra-price") # original price discount = await safe_get(".font-weight-sm") # If name or brand missing, skip if not name: logging.warning(f"Name missing for: {url}") if not brand: logging.warning(f"Brand missing for: {url}") data = { "url": url, "name": name, "brand": brand, "mrp": mrp, "discount": discount, "scraped_at": datetime.now().isoformat() } return data except Exception as e: logging.error(f"Error scraping {url}: {e}") return None ``` Scraping product details from a website like Macy’s isn’t as easy as just opening a page and grabbing the text. Websites today are more complex—they include pop-ups, interactive elements, and layouts that don’t always follow the same rules. That’s why the scrape\_product\_data function was created. It carefully handles these challenges in a step-by-step way. This function uses a tool called Playwright, which helps us control a web browser automatically. It’s written as an "async" function, which means it can wait for certain tasks—like page loading—to finish before moving on. When the function starts, it opens the product’s page and pauses briefly to make sure everything loads completely. During this wait, it also checks for cookie consent pop-ups that often appear the first time you visit a site. If the "Accept" button shows up, the function clicks it; if not, it just continues. One important thing it looks out for is whether the website has blocked access. Sites sometimes do this if they think a bot is visiting. So the script checks for a message like “Access Denied.” If it finds that message, it skips the page, logs the issue for reference, and moves on without crashing or stopping the whole process. Next, the function starts collecting data using smaller helper functions like safe\_get and safe\_get\_all. These helpers are smart—they try different ways to find the same piece of information, so if one method fails, another might work. They also handle cases where something is missing or not in the expected place. The script then looks at labels on the page to figure out if the product is on clearance or a final sale. This kind of detail helps later when analyzing or filtering products. After that, the scraper collects a lot of useful details about the product: its name, brand, prices (both original and discounted), discount percentage, available sizes and colors, customer ratings, materials, shipping info, and more. It puts all this neatly into a Python dictionary, and it also adds the date and time when the data was collected. What makes this scraping setup really strong is its ability to handle errors without crashing. If something unexpected happens—like a network issue or a sudden change in the webpage layout—it simply logs the problem and returns nothing for that one item. The rest of the process keeps running. This makes it possible to collect clean and structured data from thousands of product pages without constant supervision. Whether you’re building a tool to track discounts, analyze product trends, or recommend items to shoppers, this function gives you a solid and reliable foundation. ### How the Main Function Efficiently Handles Large-Scale URL Scraping ``` #  MAIN RUNNER  async def main():    """    Main function to orchestrate the scraping process.    """    urls = get_unprocessed_urls()    if not urls:        logging.info("No unprocessed URLs found.")        return    logging.info(f"Found {len(urls)} URLs to process.")    for url in urls:        wait_random_delay()        user_agent = random.choice(USER_AGENTS)        async with async_playwright() as p:            browser = await p.chromium.launch(headless=False)            context = await browser.new_context(user_agent=user_agent)            page = await context.new_page()            logging.info(f"Processing URL: {url}")            data = await scrape_product_data(page, url)            if data:                save_to_jsonl(data)                save_to_mongo(data)                mark_url_processed(url)                logging.info(f"Scraped and saved data for: {url}")            else:                logging.warning(f"Skipped: {url}")            await browser.close() ``` The main() function is the central engine that drives the entire scraping workflow. Think of it as the conductor of an orchestra, guiding each step from beginning to end in a smooth and organized way. Its job is to go through a long list of product web pages—sometimes even thousands of them—and collect important details from each one, reliably and efficiently. It all begins by checking a local database (specifically, an SQLite file) to see which product URLs haven’t been scraped yet. If any are still waiting to be processed, the function takes note of how many are left and gets to work. For every product page, it introduces a small, random pause before proceeding. This pause mimics how a human might browse, which helps avoid getting blocked by the website for suspicious activity. Next, a fresh browser window is opened using automation tools, and a random user agent is applied. A user agent is like a browser’s ID badge, and changing it helps disguise the scraper so it doesn’t get recognized or restricted. The scraper then visits the product page and carefully picks out key information such as the product’s name, brand, price, and size options. This information is saved in two places. First, it goes into a lightweight .jsonl file, which is easy to handle and perfect for quick reference. Second, the data is stored in MongoDB, a powerful database that works well when managing huge volumes of information. Once the scraping for that page is successful, the URL is marked as "1" in the database so it won’t be visited again. Finally, the browser window is closed, and the function moves on to the next product page. By handling each page one at a time, the process stays clean, manageable, and much less likely to trigger any website security systems—a key benefit when dealing with such a large collection of pages. ### How the Script Begins and Handles Interruptions Gracefully ``` # Script entry point if name == "__main__":    try:        asyncio.run(main())    except KeyboardInterrupt:        logging.info("Scraping interrupted by user.") ``` Every scraping project needs a place where everything begins—a sort of “start button” for the whole process. In Python scripts, this is usually found right at the bottom of the file. You’ll often see a line that looks like if name == "\_\_main\_\_":. This line has an important job: it makes sure that the script only runs when you open it directly. If someone tries to use this script as a helper in another program, it won’t start running on its own, which is exactly what we want. Now, inside this special block, we usually kick off the main part of our code. Here, it's done with [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(main()), which starts an async function called main(). This is where most of the scraping logic lives—the part that actually visits websites, collects data, and does all the heavy lifting. ## Conclusion Scraping Macy’s sale section using Playwright and Python demonstrates how automation, smart tool selection, and thoughtful data handling can transform a large, complex website into a structured and insightful dataset. Along the way, here used powerful tools like Playwright for browser automation, SQLite and MongoDB for storage, and helpful Python libraries to clean and manage the data. What’s important here is that this isn’t just about scraping a website—it’s about doing it the right way: carefully, responsibly, and in a way that keeps your data clean, organized, and easy to use later. Whether you’re analyzing pricing trends, comparing brands, or studying how discounts work, having a solid dataset makes it all possible. So, if you’re looking to dive into web scraping for real-world e-commerce data, Macy’s is a great place to start—and with the right tools and methods, it’s entirely achievable. With this guide and approach, you’re not just collecting data—you’re building the foundation for insights that can power smarter decisions. ### FAQ SECTION ### 1\. Is it legal to scrape Macy’s sale section? Web scraping public data is generally legal, but it must comply with Macy’s Terms of Service, robots.txt, and local data protection laws. Scraped data should be used responsibly, avoiding personal data collection or excessive request rates that could impact site performance. ### 2\. Why use Playwright instead of BeautifulSoup or Requests for scraping Macy’s? Macy’s sale pages are JavaScript-rendered, meaning product details load dynamically. Playwright can execute JavaScript, handle lazy loading, pagination, and simulate real browser behavior—capabilities that traditional tools like Requests or BeautifulSoup alone cannot provide. ### 3\. What data can be extracted from Macy’s sale section? Using Python and Playwright, you can scrape product names, sale prices, original prices, discount percentages, availability status, product URLs, images, and category tags from the Macy’s sale section. ### 4\. How can I avoid getting blocked while scraping Macy’s? To reduce blocking risks, implement headless browser controls, rotate user agents, add request delays, handle cookies properly, and avoid sending too many concurrent requests. Playwright also helps mimic human browsing behavior, improving scrape reliability. ### 5\. Can this scraping method be scaled for regular price monitoring? Yes. The Playwright-based approach can be scaled using task queues, scheduled runs, proxy integration, and cloud deployment. This makes it suitable for price tracking, discount monitoring, and competitive analysis at scale. AUTHOR I’m Anusha , Data Science Intern at Datahut. I work on automating data collection and transforming unstructured retail data into meaningful insights using tools like Playwright, MongoDB, and pandas etc. At Datahut, we help businesses in retail and e-commerce unlock the power of web data using scalable scraping solutions and smart automation. In this blog, I walk you through how we extracted and structured thousands of product listings from Macy’s women’s clothing sale section—enabling deeper analysis of pricing trends, inventory patterns, and product visibility. If you're looking to scale your data collection efforts or need help turning raw web data into business-ready insights, connect with us through the chat widget on the right. We’d love to collaborate. ### Noon Skincare Category Analysis & Market Insights URL: https://www.blog.datahut.co/post/noon-skincare-category-analysis/ Last updated: 2026-09-07T09:43:39.000Z The skincare category on Noon has evolved into a highly competitive and fragmented marketplace, with hundreds of products competing for consumer attention. With 918 products across 65 brands, success in this category is driven by a careful balance of pricing, discount strategies, customer ratings, and brand scale. Using Datahut’s category-level data, this analysis breaks down how leading brands perform on Noon, where competition is most intense, and what strategies help brands stand out and win in the fast-growing online skincare market. ## Noon Skincare Category Analysis: [Market](https://www.techsciresearch.com/report/uae-cosmetics-market/1406.html?ref=blog.datahut.co) Place Category Overview The skincare category on [Noon](https://www.noon.com/uae-en/beauty/?ref=blog.datahut.co) is highly competitive, with 918 products across 65 brands. Well-known brands like Medicube, Eucerin, and Numbuzin compete alongside many emerging and niche players, making the marketplace crowded and challenging. Using Datahut’s structured, category-level data, this analysis explores pricing, promotions, ratings, and brand presence to understand what drives success. The insights highlight where competition is strongest, what consumers respond to, and how brands can position themselves better in this fast-moving skincare market. ## Noon Skincare Market Overview: Products, Pricing, Ratings & Top Brands This visual snapshot breaks down the Noon skincare category using key performance metrics such as total products, average sale price, discount levels, customer ratings, and brand count. With 918 products across 65 brands, the data highlights a highly competitive market landscape and showcases leading brands like Medicube and Eucerin that dominate category presence on Noon. ![Noon skincare market](https://www.blog.datahut.co/content/images/2026/07/img-463.png.webp) ## How Brands Compete for Shelf Space: [SKU ](https://www.blog.datahut.co/post/want-to-fix-your-unit-economics-do-what-nestl%C3%A9-did-start-saying-no-to-more-skus/)Distribution When customers browse [skincare ](https://www.statista.com/markets/418/topic/481/skincare/?ref=blog.datahut.co)products on Noon, they may not realize it, but the number of products a brand lists plays a major role in what they see. This idea is known as SKU distribution, where each SKU represents a unique product listing. Simply put, the more SKUs a brand has, the more shelf space it occupies in the digital marketplace. ![How Brands Compete for Shelf Space: SKU Distribut](https://www.blog.datahut.co/content/images/2026/07/img-464.png.webp) The chart shows that some brands have far more products listed than others. Medicube leads with 82 products, followed closely by Eucerin with 74\. Because these brands appear more often, shoppers are more likely to notice them, recognize their names, and trust them over time. Brands like Numbuzin and Kiko Milano still have good visibility, but the numbers drop as we go further down. Well-known names such as Cetaphil or La Roche-Posay have fewer listings, which means they show up less often in searches—even if their products are high quality. A simple way to think about this is a store shelf. A brand with many products spread across shelves is easier to spot than one with just a few items tucked away. Online marketplaces work the same way. More products mean more chances to be seen, clicked, and reviewed. Over time, this extra visibility builds familiarity and demand. That’s why brands like Medicube and Eucerin stand out as category leaders—they simply occupy more space in the customer’s journey from search to purchase. ## Brand [Discount ](https://www.blog.datahut.co/post/4-brilliant-ways-retailers-can-offer-price-discounts-still-preserve-margins/)Strategy Snapshot ![Brand Discount Strategy ](https://www.blog.datahut.co/content/images/2026/07/img-465.png.webp) ## Understanding the Pricing Battlefield: How Discounts Shape Competition After product variety, pricing plays a major role in skincare purchase decisions. When shoppers compare similar products, discounts often influence whether they browse or buy. This chart highlights the average discount offered by each brand in the Noon skincare category, showing long-term pricing behavior rather than one-off deals. Beurer stands out with an average discount close to 70%, reflecting an aggressive discount-led strategy. Brands like Ducray and Dear, Klairs also rely heavily on promotions, conditioning shoppers to expect price drops. Mid-range brands such as Nuxe, Banila Co, Milani, and The Ordinary follow a balanced approach, using discounts without making price the primary driver. Meanwhile, brands like Eucerin, Pixi, and Mario Badescu maintain lower discounts, signaling confidence in brand trust and product quality. Overall, while discounts matter, strong brands don’t rely on price alone. ## What the Data Says About Pricing in Skincare on Noon ![price dynamics in skincare ](https://www.blog.datahut.co/content/images/2026/07/img-466.png.webp) ## Understanding Price Dynamics in Skincare: What Customers Really Want Pricing plays a critical role in shaping customer decisions, brand perception, and overall market competitiveness—especially in the skincare category. To truly understand what drives purchases, it’s not enough to look at prices in isolation. We need to analyze how prices are distributed, how brands position themselves, and what customers prefer across different price tiers. ## Typical [Price ](https://www.blog.datahut.co/post/competitive-pricing-strategy-how-products-are-priced/)Ranges in the Category Skincare products generally fall into three broad price tiers: - Budget: Entry-level products designed for affordability and everyday use - Mid-range: Products balancing price and performance - Premium: High-priced offerings focused on brand value, ingredients, or advanced formulations Understanding these ranges helps identify where most products—and [customers](https://www.mckinsey.com/industries/consumer-packaged-goods/our-insights?ref=blog.datahut.co)—are concentrated. A heavy skew toward lower price points often signals price sensitivity, while strong activity in premium tiers suggests brand-driven or results-driven purchasing behavior. ## Top Products in the category and their dominance The goal of this analysis is to understand how much of the category’s total customer engagement is concentrated in the top-performing products. ![Top Products in the category and their dominance ](https://www.blog.datahut.co/content/images/2026/07/img-95.jpg.webp) ## Why this analysis is very important for category reporting Understanding a product category is not just about the number of products available, but about how customer attention is distributed across them. This is where [rating ](https://hbr.org/?ref=blog.datahut.co)distribution analysis becomes valuable for category reporting. The table above, created using data scraped from Noon Skincare, shows how ratings are concentrated among the Top 1, Top 5, and Top 10 products. Instead of analyzing thousands of individual listings, this approach helps us step back and understand the category at a macro level. ## What This Reveals About Competition The data shows that the Top 10 products account for 49.64% of all ratings, meaning nearly half of customer feedback is focused on just ten items. This indicates a moderately concentrated category—not completely dominated by a few products, but far from evenly distributed. Customer attention clearly flows toward top-performing products, making competition strongest at the top. ## Product Visibility and Dominance The Top 1 product alone contributes 9.78% of total ratings, signaling strong dominance. Products with higher ratings gain greater visibility across search results, recommendations, and category listings. This creates a reinforcing cycle where popular products attract even more attention. Identifying such products is crucial, as they often set pricing expectations and influence overall category performance. ## Entry Challenges and Demand Distribution Since the Top 5 products capture 34.74% of total ratings, customer trust is heavily concentrated. This makes entry difficult for new products, which may require aggressive pricing, promotions, or marketing to compete. At the same time, nearly half of ratings still come from products outside the Top 10, showing that demand is not fully locked. This suggests ongoing opportunities for mid-tier products to grow. ### Why This Matters Rating distribution analysis turns raw data into clear insight. It helps teams understand competition, visibility, and growth potential, making category strategies more data-driven and realistic rather than assumption-based. ## Voice of the Customer — What Are the Ratings Telling Us? ## 1\. Visualize the Average Rating Count Across Brand ![Visualize the Average Rating Count Across Brand](https://www.blog.datahut.co/content/images/2026/07/img-467.png.webp) ## Average Ratings — Where [brands ](https://www.blog.datahut.co/post/emotional-branding-and-understanding-customer-emotions/)stackup ![ Average Ratings — Where brands stackup](https://www.blog.datahut.co/content/images/2026/07/img-468.png.webp) ### Where the [Market ](https://www.blog.datahut.co/post/webscraping%5Ffor%5Fmarketing%5F2025/)Leaders Stand The market is currently dominated by high-performers that have managed to maintain high ratings even while selling thousands of units. Here is how the top players currently stack up in the Noon Qatar market: When analyzing these ratings, we see a clear divide in the "pricing battlefield": - The 4.5+ Elite: These brands (like Vichy and La Roche-Posay) often sell at the Premium price point (100+ QAR). Customers are willing to pay more because the rating acts as a guarantee of safety and results. - The 4.0 - 4.3 Middle Ground: This is where most brands sit. It is a highly competitive space where brands like CeraVe and Beauty of Joseon fight for market share using a mix of effective pricing and solid reviews. - The Risk Zone (Below 4.0): In a category where the average is 4.35, a rating below 4.0 is a red flag. On Noon, products in this range often see a sharp drop-off in "Buy Box" visibility, as customers quickly move to higher-rated alternatives. ## Revenue Drivers of the Category: What Truly Influences Customer Decisions ![Revenue Drivers of the Category: What Truly Influences Customer Decisions ](https://www.blog.datahut.co/content/images/2026/07/img-469.png.webp) ![attribute](https://www.blog.datahut.co/content/images/2026/07/img-470.png.webp) ## The Attribute Insights Table ![attribute insights table](https://www.blog.datahut.co/content/images/2026/07/img-471.png.webp) ## Final Thoughts — What the Data Really Tells Us ## Final Insights: Winning in the Noon Skincare Marketplace The analysis of the Noon Skincare category reveals a market that is both highly competitive and deeply driven by trust signals. With an average rating of 4.35 and a dominant price anchor of 127.9 QAR, the landscape is clearly divided between volume-driven giants and high-trust specialists. Brands like Medicube and Eucerin have secured their leadership through sheer product depth, creating a "one-stop-shop" presence that captures diverse consumer needs from anti-aging to hydration. Meanwhile, the data highlights a "Trust Tier" where brands such as Vichy and Esthederm win not just on price, but on superior average ratings, signaling to shoppers that their premium investment is backed by proven results. Shopper behavior in this category is heavily influenced by the "Value" price band (50–150 QAR), which houses over 58% of the total assortment. This segment acts as the primary engine for conversions, as it balances affordability with the high quality Qatari consumers demand. Beyond the numbers, emotional drivers play a pivotal role; our extraction of consumer sentiment shows that terms like "effective," "glow," and "lightweight texture" are the key hooks that transform a casual browser into a loyal buyer. Success on Noon is not merely about having a product; it is about positioning that product within the right price bracket while maintaining the "star power" of high ratings and positive emotional resonance. The Noon skincare category is a highly competitive ecosystem where scale, pricing, and trust signals collectively determine success. With an average category rating of 4.35 and a key price anchor around 127.9 QAR, the market clearly separates volume-led brands from high-trust specialists. Leaders such as Medicube and Eucerin dominate through extensive assortments that position them as “one-stop” skincare destinations, catering to a wide range of consumer needs. At the same time, premium brands like Vichy and Esthederm stand out within a clear “trust tier,” leveraging superior ratings to justify higher prices and signal consistent product performance. Consumer demand is heavily concentrated in the value-driven 50–150 QAR price band, which accounts for over half of the category and serves as the primary conversion engine. Within this range, emotional cues—such as perceptions of effectiveness, visible glow, and lightweight textures—play a decisive role in shaping purchase decisions. Ultimately, success on Noon depends on precise positioning: aligning products with the right price band, sustaining strong ratings, and reinforcing positive emotional associations. ### The Datahut Advantage In a fast-evolving marketplace like Noon, competitive advantage is built on accurate, timely intelligence. Brands that rely on clean, structured data can track pricing movements, monitor promotional intensity, identify assortment gaps, and respond to competitor actions in real time. Datahut enables this shift from intuition to precision. Through enterprise-grade data delivery, we transform complex marketplace signals into actionable insights—empowering brands to optimize visibility, defend market share, and confidently lead on the digital shelf. ## Turn These Insights Into Growth If you want to: - Track competitors’ pricing and discounts in near real time - Benchmark your brand against category leaders - Identify hidden revenue leaks, data gaps, or assortment blind spots - Build dashboards that your marketing, sales, and merchandising teams can act on - Power AI models, forecast demand, or evaluate category trends with clean data Datahut can deliver that data—fully structured, refreshed automatically, and tailored to your exact needs. 👉 Get a free 10-minute data audit and see how you stack up against competitors. Talk to a Datahut Data Strategist to explore category monitoring, pricing intelligence, or custom data pipelines. Start a pilot in days, not months. No infrastructure, no internal engineering dependency—just clean, reliable data delivered on schedule. Winning in e-commerce starts with knowing the market better than anyone else. Datahut gives you that edge. ## FAQs ### 1\. What factors most influence skincare purchase decisions on Noon? Purchase decisions are primarily driven by product efficacy, visible skin benefits, texture & feel, and safety concerns. Customers look for products that work effectively while being comfortable and safe for daily use. ### 2\. Do customers prefer budget, mid-range, or premium skincare products? The analysis shows a strong preference for budget and mid-range products, driven by value perception and affordability. However, premium products perform well when they strongly signal efficacy and visible results. ### 3\. Which skincare attributes are mentioned most frequently in customer reviews? The most frequently mentioned attributes include, hydration, glowing skin, lightweight texture, and smooth skin, indicating that both performance and experience matter equally to shoppers. ### 4\. How important is product texture and feel in influencing conversions? Texture and feel play a high-impact role. Attributes like lightweight, non-sticky, and fast absorption significantly affect repeat purchases, even when efficacy is strong. ### 5\. What skin concerns are shoppers most focused on solving? Customers most commonly mention acne and pigmentation, showing strong demand for problem-solving skincare rather than purely cosmetic benefits. ### 6\. How do negative sentiments affect buying behavior? Mentions of irritation, breakouts, and adverse reactions create hesitation and reduce conversion rates, especially among first-time buyers. Products positioned as gentle or suitable for sensitive skin perform better. ### 7\. How can brands use these insights to improve performance on Noon? Brands can improve performance by: - Highlighting proven effectiveness and visible results - Optimizing formula texture and comfort - Addressing specific skin concerns clearly - Reducing fear through safety-focused messaging and clean formulations ### Build vs Buy Web Scraping in 2026 | Data Teams URL: https://www.blog.datahut.co/post/build-vs-buy-web-scraping-in-2026/ Last updated: 2026-09-07T09:43:41.000Z ## Why I’m Writing This As a founder, I engage in conversations with product teams, data leaders, and day-to-day business managers. Almost every time I have one of these conversations, I hear the same question asked: ### “Should we build our own scraping stack, or should we buy?” This build vs buy web scraping debate is often framed as a tooling choice. In reality, it’s a strategic decision about where your most expensive and scarce resource—engineering time—should be spent. I have observed numerous teams spend a lot of time developing scraping tools that they later did not want to continue developing. Additionally, I have witnessed teams outsourcing their work without any knowledge, thereby losing control over portions that could have been leveraged. What I will do is demonstrate how I would assess that trade-off in very general and simple terms. If there’s one idea to keep in mind as you read: [Web scraping infrastructure](https://www.blog.datahut.co/post/web-scraping-tools/)[ ](https://www.blog.datahut.co/post/web-scraping-tools/)is no longer a differentiator. What you do with the data is. I recently spoke with a startup founder who spent nearly 9 months building an internal scraping stack. They had to use additional resources on the scraping stack and by the time their web scrapers were stable, their original product roadmap had slipped by a 5 months and could not raise the additional funds they needed from the investors. ## The Real Shift People Miss: Scraping Is No Longer “Just Code” Ten years ago, scraping was relatively straightforward: Back then, libraries like [Beautiful Soup](https://www.crummy.com/software/BeautifulSoup/?ref=blog.datahut.co) were often enough to get the job done. You could fetch a page, parse the HTML, and move on without worrying too much about how the site behaved. Scraping data often involves sending an HTTP request, checking the response for relevant data, and then moving on. You could reliably retrieve the data you wanted with a simple HTTP GET Request, avoiding issues with how pages were rendered, found, or behaved, since you didn't have to render the page in a browser to get your data. - Write a script - [Parse some HTML](https://techcrunch.com/2023/08/15/browse-ai-help-companies-build-bots-to-scrape-website-data-and-put-it-to-work/?utm%5Fsource) - Schedule a cron job Today, modern web scraping looks very different. What used to be simple scripts has evolved into long-running web scrapers that behave more like infrastructure than code. - Evolving and sophisticated[ ](https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/)[anti‑bot systems](https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/) - Client-rendered, JavaScript-heavy websites - Cookies and fingerprinting are seen as methods of detection. - Breakages that are experienced continuously create an ongoing need for adaptation At this point, scraping has turned into full‑time anti‑bot warfare. What teams are really dealing with now is continuous anti-bot evasion, not one-time scraper development. Modern web scrapers now require continuous adaptation just to maintain baseline data coverage. Without that adaptation, teams quickly run into IP bans that silently degrade coverage and data freshness. Self healing scrapers is the new normal - you building a similar tool does not make any difference. Building a distributed scraping infra is should not be your priority - building the product should be. Tools like Beautiful Soup (for scraping data from static webpages) are breaking down before our eyes as more and more sites rely on dynamic (rather than static) client-side rendering and behavioral detection. This problem has moved beyond a simple maintenance issue; it has become an infrastructure problem that requires sustained effort to continue operating. This is why many teams increasingly rely on [specialized ](https://www.datahut.co/solutions?ref=blog.datahut.co)[web scraping services](https://www.datahut.co/solutions?ref=blog.datahut.co) instead of maintaining fragile internal scrapers. ## The Question I Ask Founders and PMs When considering whether to buy or build web scraping capabilities, the question becomes even more relevant: "Is this where you want to win?" This is often when the project manager will pause before responding with: "Honestly? No. We thought we had to own it." At this point, that response will generally shift this PM's thought process from habitual to strategic. Your engineering resources are your most treasured and limited resource. But if 20-30% of the time spent by your engineers is working on broken web scraping tools, continuously rotating proxies, and/or reacting to changes on your target sites, then they are not: - Improving product insights - Building better models - Creating smarter analytics - Shipping features customers will pay for This is where the build vs buy decision stops being technical and becomes strategic. Most teams don’t set out to become experts at maintaining web scrapers—it happens accidentally. Instead of fixing scrapers, teams could use that time to analyze pricing data, assortment gaps, or competitive positioning. ## A Simple Mental Model: Commodity vs Differentiator This is the simplest way to explain the decision to non‑technical stakeholders. ### Commodity Layer (Not Where You Win) These are capabilities you need, but don’t get credit for: - Proxy rotation and IP management - Headless browsers and[ JavaScript rendering](https://web.dev/articles/rendering-on-the-web?ref=blog.datahut.co) Running [headless browser](https://pptr.dev/?ref=blog.datahut.co) farms inside a company is expensive and hard to manage. This is especially true when you need to scale beyond a few sites. - CAPTCHA handling Tasks like solving CAPTCHAs are necessary in modern scraping. They add extra work but do not make the product different. - General anti‑bot adaptation Everyone needs these. No one wins because of them. Keeping anti-bot evasion working well over time is operational work. It does not make the product different. Managing [residential proxies](https://www.blog.datahut.co/post/a-guide-to-using-proxies-for-web-scraping/) at scale is costly and rarely worth the internal overhead. This layer is typically handled by external web scraping services, not internal product teams. ### Differentiator Layer (Where You Actually Win) Data creation of actual value consists of: - The types of data collected (so there's no duplication), - How you [clean, enrich, and validate it](https://www.blog.datahut.co/post/7-metrics-to-measure-data-quality/) - Type of ETL processes used to transfer data for storage and analysis - Types of analytics, models, and business rules used to analyze, interpret, and create data, and - How tightly is this data integrated into your product? This is the layer your customers actually experience. Much of this work involves turning messy, unstructured data into formats your systems and decision-makers can actually use. Well-designed ETL pipelines change raw data into something reliable, comparable, and ready for decisions. Long ago, some teams discussed whether to develop their own logging or monitoring systems. Today, the discussion is not seen as a strategy. The same thing has happened with scraping: many teams have not yet updated their thinking about it. ## How do you evaluate your options? This is the code idea we use internally and recommend to most teams evaluating build vs buy web scraping. ![where to build vs buy in web scraping ](https://www.blog.datahut.co/content/images/2026/07/img-50.jpg.webp) The idea is simple: - You buy reliability, scale, and resilience at the infrastructure layer - You retain intelligence, logic, and competitive advantage at the value layer The boundary matters. You outsource the pain. You keep the brain. That’s what we call Bounded Buy. ## A Simple Way to Think About How Capabilities Mature You don’t need formal frameworks or jargon to understand this part. The idea is straightforward: Most capabilities develop in a predictable way over time. What starts as something rare and strategic eventually becomes something expected and operational. Here’s how that typically plays out: 1\. NovelAt this stage, very few teams can do the thing at all. It needs testing, deep knowledge, and creativity. Early on, building this in‑house can make sense because there are no reliable external options. 2\. Custom (Built In‑House)As more teams face the same problem, they start building their own internal versions. Each implementation looks slightly different, and a lot of engineering time goes into making it work for specific use cases. 3\. ProductizedOver time, vendors emerge. They standardize the problem, package it into tools or services, and make it easier to adopt. At this stage, buying often becomes cheaper and faster than building. 4\. CommodityEventually, the capability becomes expected. Everyone needs it. Best practices exist. Scale, reliability, and good operations matter more than clever design. [Web scraping infrastructure](https://www.blog.datahut.co/post/web-crawling-and-its-use-cases-for-2026/) has clearly reached this final stage. The web scraping industry has matured significantly over the last decade, with established vendors, best practices, and strong [economies of scale](https://www.investopedia.com/terms/e/economiesofscale.asp?ref=blog.datahut.co). What matters now is uptime, resilience, and consistency—not bespoke engineering. This is where many build vs buy web scraping decisions go wrong. Teams continue to treat mature, commodity infrastructure as if it were still a source of competitive advantage. Your analytics, models, data interpretation, and decision logic are not common or standard yet. They are shaped by your business context and customer needs—and that’s exactly where internal creativity and ownership belong. ## Option 1: Build Everything In-House (If You MUST Have Control Over Everything) If you cannot compromise at all on the build-out end of the web scraping build vs.-buy spectrum, then a completely in-house web scraping system is owned by teams. In reality, a legitimate in-house scraping system consists of: Behind the scenes, this often means operating headless browser farms just to render pages reliably at scale. - Multiple senior engineers - Dedicated DevOps support - Continuous proxy and infrastructure spending - Continuous maintenance for evolving sites Most teams are not budgeting enough for expenses. In reality, you should expect: - Between $150k and $400k for initial development - Several hundred thousand dollars each year for salary costs - Between 20-30% of engineering time will go toward maintenance activities. These numbers still understate the real issue: maintenance costs compound over time as sites change, detection evolves, and internal tooling ages. Much of this effort goes into keeping web scrapers functional rather than improving downstream insights. A large part of the spending often goes to buying and rotating residential proxies. It does not improve data quality or insights. ### Downsides of Pure In‑House Build - CAPTCHA resolution does not grow exponentially: Increasing the number of sites that use CAPTCHA results in high costs. - Anti-bot avoidance measures will never be complete: New techniques for detecting automated traffic are continually evolving, keeping the internal team in a cycle of reacting. - High opportunity costs: The time spent developing scrapers takes time away from developing product, models, and features that are customer-facing. - Compounding impact: Over time, the costs of maintaining scrapers, infrastructure sprawl, and the operational risks associated with poorly built scrapers continue to compound. ### This approach only makes sense when: - You operate in heavily regulated environments. - You need end‑to‑end auditability - You scrape highly proprietary or internal systems. For most teams, this is an expensive way to solve a non‑differentiating problem. ## Option 2: Bounded Buy (Hybrid Model) If you are looking for a way to maintain control of your data without reinventing the wheel, the Bounded Buy Model may be your best option. In the Bounded Buy Model, you will: - Utilize either commercial unblockers or scraping platforms as a reliable means of extracting data. - Maintain complete control over how you process, validate, and use your data. The Hybrid Model takes advantage of the best of both worlds by giving teams ownership of their internal data while also providing access to external web scraping services to increase their scale and resilience. Advantages - Reduced time to market from months to a matter of weeks. - Fixed costs become variable or usage-based - Engineering can focus on developing value for the customer. Disadvantages of the Hybrid Model - Anti-bot evasion techniques can still be detected: although many vendors provide reliable solutions for the commodity (unblocker/scraper), vendors operating upstream of the commodity will still run into operational issues when their techniques change. - Complex Integrations: Engineering teams will continue to have effort(s) integrating, monitoring, and adjusting to the vendor-provided information. - Partial Dependency: Your reliance on external vendors for the commodity layer results in the reliability of your use of that vendor’s service. - Zero-maintenance is a misnomer: Although operational oversight is considerably less, it is still necessary. Even in a hybrid environment, the teams are still responsible for monitoring any failure of the upstream web scraper. Most importantly, your competitive advantage stays in your codebase - not your vendor’s. ## Option 3: Fully Managed Service A[ ](https://www.datahut.co/solutions?ref=blog.datahut.co)[fully managed web scraping service](https://www.datahut.co/solutions?ref=blog.datahut.co)[ ](https://www.datahut.co/solutions?ref=blog.datahut.co)works best when speed and simplicity matter more than deep customization. ### This solution is a good fit in the following cases: - There is an immediate need for data - Requirements are well defined - Minimal custom logic is used ### There are also downsides to a Fully Managed Service: - Less flexible - you must operate within the vendor’s data model and delivery structure - Less control - the ability to modify scraping behavior is typically abstracted away - Vendor dependency - if you want to switch providers later, it may require remapping the workflow You are trading some flexibility for a focus on the analytics aspect of your solution. This trade-off makes a lot of sense and is often optimal for many analytics-driven and early-payment use cases. In this model, teams no longer need to think about how individual web scrapers are built or maintained. ## TCO Comparison: Three Web Scraping Sourcing Models This view shows how teams should evaluate build vs buy web scraping decisions over a realistic three‑year horizon. People often miss how maintenance costs add up over time. These costs include engineering time, extra infrastructure, and operational risks. Costs here are driven largely by proxy acquisition, especially residential proxies, and the effort required to keep them usable over time. ## Build vs Buy: A Practical Guide for Web Scraping Legal and Compliance: The Part You Can’t Ignore Scraping isn’t just a technical problem—[it’s a legal and operational one](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/). If you handle: - Personal data - Regulated markets - Strict retention or deletion requirements Then [compliance](https://gdpr.eu/?ref=blog.datahut.co) workflows matter as much as cost. This often becomes the hidden deciding factor in build vs buy web scraping decisions. Hybrid and managed models work well here because infrastructure risk is outsourced while compliance logic can remain tightly controlled. ## My Founder Takeaway (Where We Actually Land) We have seen hundreds of build vs buy web scraping decisions across companies. We have watched how the web scraping industry has changed. Our recommendation today is clear. For most teams, Fully Managed Service is the right default choice. Scraping infrastructure has crossed the point where owning it creates meaningful leverage. Continuous [anti-bot evasion](https://www.cloudflare.com/learning/bots/what-is-bot-management/?ref=blog.datahut.co) is now table stakes, not a competitive edge. Proxy management, browser orchestration, bot evasion, and scaling are now areas where teams quietly lose time and momentum. A fully managed approach allows teams to: - Eliminate months of setup - Avoid permanent headcount for non‑core problems - Transfer uptime and breakage risk to specialists - Pay only for the data they actually use Most importantly, it removes an entire class of operational distraction from the roadmap. ## Final Thoughts If you strip away pride, tooling bias, and sunk‑cost thinking, the decision becomes simple: If scraping does not directly differentiate your product, you should not be running scraping infrastructure. The best build-vs-buy decisions protect focus, accelerate learning, and keep teams working on what customers actually pay for—better insights, faster iteration, and cleaner data. That’s why, in practice, this is the model we see scale with the least friction over time. ## Frequently Asked Questions (FAQs) ### 1\. When does it actually make sense to build web scraping in-house? Only when the web scraping component of your site is fundamentally related to your main product or any required compliance issues is it worth doing your web scraping in-house. This generally occurs in regulated environments with private internal systems or if the way you gather data will provide a competitive advantage over others in the industry. For most analytics or research applications, building something in-house will likely create more headaches and disadvantages than building with a third party. ### 2\. Are web scraping services reliable for long-term use? Most mature, established web scraping services are designed to provide a dependable service over a long period of time. They continue to invest in proxy management, browser control systems, and effective anti-bot measures. In-house, these tasks tend to take a great deal of time and money. It is most beneficial to work with a vendor that can demonstrate a track record of scale, transparency in SLAs, and clarity about the quality of their data. ### 3\. How do I evaluate build vs buy web scraping from a cost perspective? Cost needs to be looked at as the total cost of ownership, not just the purchase price of tools. When you build an in-house solution, there are additional hidden costs (engineering time, maintenance time, infrastructure sprawl, and therefore opportunity cost) that are not included in the comparison of the initial cost vs. buying or using a fully managed service. ### 4\. Does using a fully managed web scraping service mean losing data control? Not necessarily true, when you use fully-managed services, you will not have any visibility to how the actual scraping is being done (e.g., are you using a scraper tool, and if so, what type), but you still have control over how your team stores the data, transforms it, validates it, and uses it internally. In many cases, control over the use of the data will be more important than control over the mechanics of data collection. ### 5\. What is the biggest mistake teams make in build vs buy web scraping decisions? The most common mistake is treating scraping infrastructure as a source of differentiation. Teams often think owning common infrastructure is more valuable than it is. They also forget about the long-term costs to run it. Teams mix up owning web scrapers with having a competitive edge. ### Amazon Menstrual Cup Category Insights & Trends URL: https://www.blog.datahut.co/post/amazon-menstrual-cup-category-analysis/ Last updated: 2026-09-07T09:43:43.000Z [Menstrual cups](https://www.treksandtrails.org/blog/the-menstrual-cup-changed-my-life-and-its-bound-to-change-yours-too/?ref=blog.datahut.co) are no longer a niche product. As more consumers seek sustainable and cost-effective menstrual hygiene solutions, [Amazon](https://www.amazon.in/?ref=blog.datahut.co) has become a key marketplace for discovering, comparing, and purchasing menstrual cups. With hundreds of brands, wide price variations, heavy discounts, and thousands of reviews, how do shoppers decide which menstrual cup to buy? To answer this, we scraped real [Amazon menstrual cup listings](https://www.amazon.in/s?k=menstrual+cup&ref=blog.datahut.co) and analyzed pricing, discounts, ratings, reviews, and brand positioning using [exploratory data analysis (EDA)](https://medium.com/@akshatsharma0610/a-tour-to-eda-exploratory-data-analysis-fafef76d38a7?ref=blog.datahut.co). In this blog, [we uncover menstrual cup pricing trends on Amazon](https://www.blog.datahut.co/post/scraping-amazon-s-menstrual-cup-data-using-playwright-and-curlcffi/), identify top-performing brands, analyze customer sentiment, and [reveal](https://www.blog.datahut.co/post/is-it-legal-to-scrape-amazon-unethical-uses-of-amazon-web-scraping/) what truly drives conversions in this fast-growing category. [GET THE FULL AMAZON MENSTRUAL CUP DATA HERE!](https://tally.so/r/aQ97o2?ref=blog.datahut.co) ## Key Findings from Amazon Menstrual Cup Category Analysis Based on Amazon marketplace pricing data, we find the following key themes: - A small set of brands dominate visibility, with the top 10 products capturing over 80% of all ratings. - Most menstrual cups compete below ₹1,000, creating intense mid-range price pressure. - Heavy discounting is widespread and used as a primary visibility lever. - Customer satisfaction remains consistent across price bands, indicating that higher prices do not guarantee better ratings. ## Amazon Menstrual Cup Market Overview ![Amazon category overview](https://www.blog.datahut.co/content/images/2026/07/img-77.jpg.webp) ### Category Leaders ### [Sirona](https://thesirona.com/?srsltid=AfmBOoqbRP-rnELSodXxEc-K4fX9SCrzTxBg3BrYBKDzZ5denBWvKW6M&ref=blog.datahut.co) Sirona is known for its strong brand presence, high review count, and wide product assortment. [Sirona](https://thesirona.com/?srsltid=AfmBOooSJGl7ECRdmrDU7IrNo3yhVfJeS008jOLq-d8JGkfWVQN3P-Iq&ref=blog.datahut.co) appears as a clear leader in shopper engagement. ### [Pee Safe](https://www.peesafe.com/collections/christmas%5Fnewyear%5Fsale?campaignid=PS%5FGoogle&adgroupid=186815805749&creative=787459958579&matchtype=e&network=g&device=c&keyword=pee%20safe&gad%5Fsource=1&gad%5Fcampaignid=22141549425&ref=blog.datahut.co) [Pee Safe](https://www.peesafe.com/?srsltid=AfmBOopZCnAZrNjM%5FbY50t5KRgjpy77PsREKSx5190GfeYb62nEIf5iH&ref=blog.datahut.co) is another widely recognized name in the menstrual hygiene space, holding its position through consistent ratings, pricing strategies, and marketing visibility. ## Which Brands Have the Highest Average Price? This plot provides a compact but powerful overview of the menstrual cup category on Amazon. It highlights the competitive landscape and pricing trends. ![Brands having the highest average price](https://www.blog.datahut.co/content/images/2026/07/img-78.jpg.webp) This bar chart compares the [menstrual cups](https://www.treksandtrails.org/blog/the-menstrual-cup-changed-my-life-and-its-bound-to-change-yours-too/?ref=blog.datahut.co) across brands, helping highlight how different brands position themselves in terms of pricing. Each bar represents a brand, and its height shows the average price customers actually pay. ### Key Insights - LENA clearly stands out as the most expensive brand, with a much higher average price than all others. - Cora and Asan follow, indicating a premium pricing strategy. - Most remaining brands cluster below ₹1,000, showing strong competition in the mid-price segment. The price compression suggests that real competition happens around affordability rather than premium differentiation. Strategic takeaway: New brands attempting premium positioning without established trust face a steep adoption barrier. ## How Shoppers Engage with Menstrual Cup Brands: Ratings Breakdown ![How Shoppers Engage with Menstrual Cup Brands: Ratings Breakdown](https://www.blog.datahut.co/content/images/2026/07/img-79.jpg.webp) From the chart, it’s clear that Sirona leads by a wide margin with an average rating count of 8.1k, making it the most reviewed and likely the most purchased menstrual cup brand in the dataset. Pee Safe follows with around 7.6k average ratings—showing strong brand presence but still behind Sirona. ### Key Insights - Sirona dominates the market, receiving almost 40% more ratings than the second-ranking brand. - Only two brands (Sirona and Pee Safe) have exceptionally high customer engagement. - The middle tier—brands like LENA (3.2k) and Carmesi (3.0k)—receives moderate but solid engagement. - The overall distribution is highly uneven, suggesting that customer trust and recognition are concentrated among a few leading brands. The average rating count per brand provides strong insight into market popularity and customer trust. A clear pattern emerges: a small number of brands dominate customer attention, while the rest compete in a much narrower visibility range. Understanding this helps us identify leaders in the category, gaps in competition, and opportunities for growth. ## Who Discounts the Most? ![Who Discounts the Most?](https://www.blog.datahut.co/content/images/2026/07/img-80.jpg.webp) From the chart, it’s clear that MW PARIS offers the highest average discount at 84%, making it the most aggressive discounter among the listed brands. This could be a strategy to capture more market share or increase product visibility. Nbhaag Enterprise (76.67%), Bildos (73%), and VSEI (70%) also provide steep discounts, indicating that they compete heavily on price. On the other end, brands like LADY GO (60.5%) and Mowell (59%) offer comparatively moderate discounts. These brands may rely more on product quality, brand reputation, or consistent pricing rather than heavy promotions. ### Key Insights 1. Strong promotional competition There is a clear trend of high discounting across most brands, showing that the menstrual cup category is price-competitive on Amazon. 2. A few brands rely heavily on discount-driven visibility Brands such as MW PARIS, Nbhaag Enterprise, and Bildos offer 70%+ discounts on average, suggesting a strategy geared toward increasing conversions through price drops. 3. Lower-discount brands may be positioned as premium or stable LENA and Ezcup maintain lower-than-average discount percentages, which could indicate: 4. A more premium positioning 5. Dependence on brand trust rather than promotion 6. Steady pricing with minimal fluctuations 7. Wide variation indicates fragmented pricing strategies Discount percentages range from 59% to 84%, showing that brands are experimenting with different approaches to attract buyers. 8. Discounting is a major driver in this category The overall high discount levels suggest that customers shopping for menstrual cups may be price-sensitive, and brands respond by offering deep discounts. The analysis of average discounts by brand reveals that the [menstrual cup category](https://www.pinkishe.org/blog-post/which-menstrual-cup-is-best-for-you?gad%5Fsource=1&gad%5Fcampaignid=21679242440&ref=blog.datahut.co) on Amazon is highly competitive, with many brands relying heavily on aggressive promotions. While a few brands maintain moderate discounts, most are using steep markdowns to stand out. This highlights the importance of pricing strategy in customer acquisition, especially in categories where multiple brands offer similar products. ## Price Band Analysis – How Menstrual Cup Products Are Distributed Across Budget, Mid-Range, and Premium ![Price bands overview](https://www.blog.datahut.co/content/images/2026/07/img-81.jpg.webp) This table explains how prices are distributed across the category. Understanding price bands helps to identify where most products are positioned, how brands target different customer groups, and whether higher-priced products deliver better customer satisfaction. The table summarizes how many SKUs (how many products belong to that price band) fall into each band, along with the average rating and average price for that group. ### Key Insights - The Budget segment (≤ ₹500) is the largest group, with 185 SKUs. This means most menstrual cup products on Amazon fall into the low-price category. Despite being the cheapest tier, the average rating is 4.24, showing that customers are largely satisfied even with budget-friendly options. - The Mid-range segment (₹501–₹1000) has 70 SKUs. This segment holds a balance between affordability and quality, and it actually has the highest average rating of 4.30 among all three tiers. This suggests that customers may perceive mid-range products as offering the best value for money. - The Premium segment (₹1001+) has only 36 SKUs, making it the smallest group. Even though these products are priced much higher (average price ₹1817), the average rating is 4.27, similar to the budget tier. This indicates that paying more does not necessarily translate to higher customer satisfaction. - Pricing does not strongly influence ratings: All segments have similar ratings (4.27–4.30), showing that customer satisfaction is not dependent on price. The price band analysis shows that the menstrual cup category on Amazon is heavily skewed toward budget products, with consistent customer satisfaction across all segments. Mid-range products emerge as the sweet spot for value, while premium offerings remain niche and do not necessarily deliver higher ratings. ## Rating Concentration Analysis – How Few Products Capture Most Attention ![How Few Products Capture Most Attention](https://www.blog.datahut.co/content/images/2026/07/img-82.jpg.webp) This table shows how customer ratings are distributed across top-performing products, helping us understand visibility and competition within the category. The data clearly indicates a highly concentrated market: the top product alone accounts for 18.55% of all ratings, while the top 5 products capture nearly 68%, and the top 10 account for over 84% of total engagement. This means customer attention is heavily focused on a small set of listings, making it harder for new or lesser-known products to gain visibility. In such a setup, success is driven by breaking into the top tier, as most demand and trust are concentrated around established bestsellers. Customer engagement is highly concentrated. Sirona and Pee Safe lead by a wide margin, benefiting from review momentum that reinforces Amazon’s search and recommendation algorithms. Once brands cross a critical review threshold, visibility compounds. Implication: New entrants must plan explicitly for review velocity and early trust-building to escape the long tail. ## Revenue Drivers of the Category: What Influences Customer Decisions ![Revenue Drivers of the Category: What Influences Customer Decisions](https://www.blog.datahut.co/content/images/2026/07/img-83.jpg.webp) This chart highlights the most frequently mentioned product attributes in customer reviews, showing what shoppers care about most when making purchase decisions. Quality leads by a clear margin, followed closely by comfort and ease of use, indicating that customers prioritize reliability and everyday usability over everything else. Value for money also plays a strong role, reinforcing price sensitivity, while softness, though mentioned less often, remains an important factor affecting user comfort. Overall, the chart shows that functional performance and user experience are the strongest drivers of customer trust and conversion in this category. ![Revenue Drivers of the Category: ](https://www.blog.datahut.co/content/images/2026/07/img-84.jpg.webp) ## Strategic Implications for Brands - The mid-range price band is the safest entry point for new launches. - Review accumulation is as critical as product quality. - Transparent pricing may build long-term trust in health-focused categories. - Messaging should emphasize comfort, leakage security, and ease of use. ## Conclusion: Turning Category Data Into Competitive Advantage ### Key Takeaways from Amazon Menstrual Cup Data - Budget products dominate listings but not satisfaction. - Mid-range products deliver the highest perceived value. - Heavy discounting is the norm, not the exception. - Comfort and leakage protection drive conversions more than price. - Brand trust concentrates around a few leaders like Sirona and Pee Safe. The [Amazon](https://www.blog.datahut.co/post/the-secret-weapon-of-successful-amazon-sellers/) menstrual cup category demonstrates a pattern common across many fast-growing [e‑commerce](https://www.shopify.com/blog/what-is-ecommerce?ref=blog.datahut.co) segments: intense price competition, heavy reliance on discounts, and extreme concentration of customer trust around a few dominant listings. In such environments, intuition and surface-level metrics are not enough. Winning brands are those that understand where competition is truly happening (price bands, review velocity, visibility thresholds) and why customers choose one product over another (comfort, leakage security, ease of use). Category-level data makes these dynamics visible—and actionable. For teams making decisions around pricing, positioning, product launches, or marketplace expansion, having access to clean, structured, and continuously refreshed data is no longer optional. It is a strategic requirement. ## About Datahut [Datahut](https://www.datahut.co/?ref=blog.datahut.co) provides end-to-end marketplace data extraction and category intelligence for brands, retailers, consultants, and data teams. We help organizations move from raw marketplace data to clear, defensible decisions. If you’re evaluating this category or any other Amazon or e‑commerce segment, Datahut can support your analysis with reliable, scalable data pipelines. ### FAQ Section ### 1\. What do shoppers look for most when buying menstrual cups on Amazon? Shoppers primarily focus on comfort, leakage protection, ease of use, and overall product quality. Customer reviews show that comfort and leakage security are the strongest emotional drivers influencing purchase decisions. ### 2\. Which price range of menstrual cups performs best on Amazon? Mid-range menstrual cups priced between ₹501 and ₹1000 perform best in terms of customer satisfaction. This segment balances affordability and quality, achieving the highest average ratings compared to budget and premium products. ### 3\. Do higher-priced menstrual cups receive better ratings on Amazon? Not necessarily. Data shows that premium menstrual cups do not receive significantly higher ratings than budget or mid-range products. Customer satisfaction remains consistent across price bands, indicating that higher prices do not always mean better performance. ### 4\. How important are discounts in the menstrual cup category on Amazon? Discounts play a major role in shopper decision-making. Most brands rely heavily on aggressive discounting to stay competitive, improve product visibility, and increase conversions in this crowded category. ### 5\. Which menstrual cup brands dominate customer engagement on Amazon? A small number of brands dominate customer engagement, with Sirona and Pee Safe leading in average rating counts. These brands benefit from strong brand trust, high visibility, and consistent customer feedback. ### 6\. What are the most common customer complaints about menstrual cups? The most common concerns raised in customer reviews relate to softness and leakage. While overall sentiment is positive, these issues significantly impact trust and repeat purchases. ### 7\. How competitive is the menstrual cup market on Amazon? The menstrual cup category on Amazon is highly competitive, with over 60 active brands and nearly 300 product listings. Pricing, discounts, reviews, and brand visibility all play critical roles in gaining shopper attention. ### Efficient Ethos Product Data Scraping with Python Guide URL: https://www.blog.datahut.co/post/scrape-ethos-product-data/ Last updated: 2026-07-23T07:48:30.000Z When we think about luxury watches, they are rarely just about telling time—they carry stories of craftsmanship, heritage, and personal style. In India, one name that consistently stands out in this space is [Ethos Watches](https://www.ethoswatches.com/?ref=blog.datahut.co), the country’s largest luxury and premium watch retailer. With a market share of 13% in the premium and luxury segment and a market cap of over ₹7,400 crore, Ethos has built a reputation for trust and authenticity. Every timepiece sold here goes through a series of checks to ensure it is 100% genuine, giving customers complete confidence in their purchase. From iconic names like Rolex, Omega, and TAG Heuer to niche Swiss watchmakers, Ethos curates a wide collection that caters to different tastes—whether it’s a minimal everyday piece, a sporty companion, or an intricate horological masterpiece. With more than 60 boutiques across 20 cities, along with a strong online presence through [ethoswatches.com](http://ethoswatches.com/?ref=blog.datahut.co), Ethos makes the world of fine watchmaking accessible to Indian buyers. Ever wondered what stories lie hidden behind luxury watch collections? By looking closely at data from Ethos Watches, we can uncover patterns that go beyond just brand names or price tags. From understanding which watch types are most popular to seeing how design choices influence buyer preferences, data helps us see the bigger picture of the luxury watch market in India. In the sections ahead, we’ll walk through the process of gathering and analyzing this data to bring those insights to light. ## Automated and Insightful Data Extraction We worked with Ethos Watches’ data in two stages: first, collecting product page links from their online store, and then preparing the information for deeper analysis. Our focus was on exploring different aspects of luxury timepieces—ranging from pricing patterns and brand positioning to design details like straps, glass types, and movement styles. ### Step 1: Gathering Product Page Links Before we can dive into analyzing luxury watches, the very first task is to gather the raw material: product page links. Think of it like preparing the foundation of a house—without strong building blocks, nothing else can stand firmly. In the same way, if our dataset doesn’t start with accurate product links, the later stages of analysis will fall apart. For Ethos Watches, this meant starting with their brands listing page at ethoswatches.com. This page acts like the front door to hundreds of premium timepieces, and our goal was to carefully step through each section, collect the product URLs, and save them for later. But since doing this manually would take days, we turned to automation. Using Playwright, a tool that lets us control a web browser with code, we wrote a script that could act like a diligent assistant: open the Ethos website, close any unexpected popups, and systematically record every watch link it found. Each product block on the page contains a clickable link, and by telling Playwright where to look (div.product\_sortDesc), we could extract these links one by one. Once the links were collected, we didn’t just leave them floating around in memory. To keep everything neat and reusable, the script saved the links in two places: 1. A database (SQLite) – This worked like a permanent storage cabinet, ensuring every product link was stored securely without duplication. 2. A JSON file – This provided an easy-to-read snapshot of the links from each page, which could be shared or checked later. Because Ethos Watches has multiple pages of products, the script also needed to handle pagination—that little “Next” button at the bottom of the page. Instead of us clicking it endlessly, the script kept moving forward until there were no more pages left, quietly recording every watch along the way. In short, Step 1 was all about building a strong dataset of product URLs. With over sixty boutiques and hundreds of timepieces online, Ethos offers a massive catalog, and this process gave us a structured way to capture it all. Having these links in hand is like having a detailed map before setting out on a journey—we now know exactly where each watch lives on the site, and we’re ready to explore deeper insights in the next steps. ### Step 2: Extracting Detailed Data from Each Product URL Once we had a solid collection of product links, the next logical step was to open each one and dig deeper. If Step 1 was about drawing the map, Step 2 was about walking into every boutique and carefully noting down the details of each watch. Every product page on Ethos Watches tells a story—brand, collection, price, movement type, water resistance, and more. But since no two pages are exactly alike, we needed to design our scraper to adapt. Here’s how the process worked: Using the product URLs we had stored earlier, our script revisited each page one by one and carefully extracted the details. Along the way, it was designed to handle unexpected interruptions like promotional or subscription popups by detecting and closing them automatically. Once inside, the scraper located key product information such as the title, brand, price, and technical specifications by targeting the right HTML elements to ensure accuracy. Finally, all captured details were stored securely in the same SQLite database alongside the product links, with a processed flag added to prevent re-scraping the same page. At the same time, the results were exported into a JSON file, making the dataset easy to inspect, share, or use for further analysis. The most important part of this step was reliability. Websites can be unpredictable—sometimes a page loads slower, sometimes a detail is missing. To tackle this, we included error handling in our code: if a page failed to load or a field wasn’t found, the script logged it instead of breaking down. This ensured that the scraping process could run smoothly across hundreds of pages without constant supervision. By the end of Step 2, we had transformed a simple list of product links into a rich dataset of Ethos Watches. Instead of just knowing where the watches were, we now had their full profiles: what they were, how much they cost, and what features they carried. ### Step 3: Cleaning the Extracted Ethos Watch Data When you first scrape data from a website like [ethoswatches.com](http://ethoswatches.com/?ref=blog.datahut.co), it feels exciting—you’ve just collected hundreds of rows about luxury watches, complete with brand names, models, prices, and technical details. But very quickly, you’ll notice that raw data is rarely “ready to use.” Instead, it often looks messy, inconsistent, and full of small details that can confuse your analysis. Think of it like bringing home fresh vegetables from the market. They look great at first glance, but before cooking, you need to wash, peel, and cut them. Data cleaning is exactly that process for datasets—it takes the raw, collected information and prepares it so that your analysis becomes smooth and reliable. Take the price column as an example. The raw data usually includes the ₹ sign and sometimes even commas, like ₹2,50,000\. While this looks fine to the human eye, computers prefer a simple number such as 250000\. So, one of the first steps in cleaning is removing those extra symbols to make prices easier to calculate and compare. Units also need attention. In Ethos data, details like 21,600 bph or Approx. 60 hours appear frequently. While these phrases are helpful for customers, they add unnecessary complexity to a dataset. By cleaning, we can strip away the word bph or remove Approx. from the power reserve, keeping only the clean numeric values like 21600 or 60 hours. This way, the data becomes uniform and ready for deeper analysis. The process may not sound as glamorous as scraping or visualizing trends, but it’s the bridge that connects raw collection to meaningful insights. Once the Ethos dataset is cleaned, we can confidently explore patterns—like how price relates to movement type, or whether certain watch features are linked to higher demand. Clean data doesn’t just make analysis possible; it makes the insights trustworthy. ## Essential Building Blocks Behind the Scenes: Python Libraries That Power the Workflow Behind every smooth and reliable data scraping project lies a carefully selected set of Python libraries working quietly in the background. Think of them as a team of skilled helpers—each with a specific role—making sure everything runs efficiently, stays organized, and doesn’t crash when things get tricky. In this project, we’ve brought together a handful of powerful tools that cover everything from web browsing to data saving, all while keeping the process fast and error-free. To start, we have asyncio, which acts like a traffic controller for the script. It allows different parts of the program to run at the same time without waiting in line. For example, while one product page is still loading, another can already begin processing—saving time and keeping the workflow smooth. Working closely with asyncio is playwright.async\_api, which opens and interacts with websites as if a person were browsing them. It clicks, scrolls, and fetches content—even from pages that load using JavaScript—making it perfect for modern websites. But fast scraping can also attract unwanted attention from websites with anti-bot systems. That’s where Playwright Stealth (not shown in this block but often used alongside) can be handy—it helps the browser behave more naturally, reducing the chances of being blocked. Meanwhile, TimeoutError from Playwright helps us gracefully handle situations where a page takes too long to load, so the program doesn’t crash but moves on smartly. Once data is captured, we need to store it in a reliable and organized way. That’s where sqlite3 comes in—a lightweight, file-based database that doesn’t need a server. It’s simple to use and perfect for saving structured information like URLs and product details. Think of it as a neat digital notebook that you can query anytime. Supporting all of this are a few unsung heroes. The random module helps introduce natural pauses between actions, mimicking how a real person might browse, which also helps avoid detection. The logging library quietly keeps track of what’s happening during the run—successes, failures, and everything in between—so you can go back and understand any issues. Finally, Path from the pathlib module makes it easier to manage folders and file locations in a clean and reliable way, no matter which operating system you’re using. Altogether, this thoughtful combination of Python libraries creates a strong foundation for scalable, efficient, and resilient data scraping. By letting each tool do what it does best, we ensure that the system remains easy to manage, beginner-friendly, and ready to handle real-world challenges. ## Step 1: Extracting Product URLs from the Dog Food Section on Ethos ### Imports and Initial Setup ``` # IMPORT REQUIRED LIBRARIES import asyncio import json import random import sqlite3 import logging from pathlib import Path from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError ``` Before diving into the actual scraping process, the very first thing we need to do is gather our tools—and in Python, that means importing the right libraries. Think of this like laying out everything you’ll need on your workbench before starting a DIY project. These imports are not just random names; each one plays a specific and important role in helping our scraper run smoothly and smartly. We begin with asyncio. This library is like a multitasking expert for our Python code. Normally, a program does one thing at a time—wait for a page to load, then process it, then move to the next one. But with asyncio, we can juggle several tasks at once. It’s like having multiple tabs open in your browser, each doing something useful in the background. This makes our scraper much faster and more efficient, especially when dealing with multiple pages. Next, we pull in json, which helps us handle structured data. If we ever want to save the information we collect in a format that other programs or people can read easily, JSON is the way to go. It’s like packing our data into neat little boxes, with labels on everything. Then comes random. This one might sound a bit odd in a scraper, but it actually serves a clever purpose. Websites are smart these days, and if they notice a bot clicking through their pages too quickly or in a predictable pattern, they might block it. So we use random to slow things down a little and add variation—maybe a 2-second pause here, a 3.1-second pause there—just to make our bot feel more human. Now let’s talk about sqlite3\. This is our way of storing the treasure we dig up. Think of it like a mini spreadsheet or a pocket-sized database that lives right on your computer. It doesn’t need any setup or internet access—just quietly saves everything in an organized file. This makes it perfect for projects where we’re collecting a lot of data, like hotel links, product info, or anything else. We also bring in logging. Imagine you’re keeping a diary while running your scraper. If something goes wrong—maybe a page didn’t load, or a piece of data was missing—you’ll want a record of that. That’s what logging does. It keeps a behind-the-scenes record of everything our code does, so we can look back and understand what worked, what didn’t, and why. Then we have Path from Python’s pathlib module. This is just a nice, readable way to handle file and folder paths on your computer. Instead of writing long, clunky strings to manage where files go, Path makes it all cleaner and more intuitive. And finally, we import async\_playwright from Playwright, along with TimeoutError. Playwright is the heart of our operation—it’s what allows our code to control a web browser. It can open a website, click buttons, scroll through listings, and wait for elements to appear—just like a human browsing manually. That’s especially important for modern sites that load content dynamically or hide it behind buttons. The TimeoutError part helps us deal with situations where a page takes too long to load—kind of like setting an alarm so we don’t sit around waiting forever. In short, this small block of imports is quietly doing a lot of heavy lifting. It gives our scraper the brain to think, the hands to interact, and the memory to store everything it finds. With these tools in place, we’re ready to start exploring the web—systematically, efficiently, and smartly. ### Logging Configuration ``` # SETUP LOGGING FOR DEBUGGING AND TRACKING log_dir = Path("logs") log_dir.mkdir(exist_ok=True) log_file = log_dir / f"scrape_log_{Path().cwd().name}.log" logging.basicConfig(filename=log_file, level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") """ This block sets up logging to track the script's behavior and any errors. It creates a `logs/` directory (if not already present) and stores log messages in a file with the current folder's name in it. """ ``` Once we’ve gathered our tools by importing all the necessary libraries, the next important step in building a scraper is to set up a way to keep track of what’s happening behind the scenes. Imagine you’re baking a cake for the first time and you decide to jot down everything as you go—what worked, what didn’t, and where you had to make changes. This is exactly what logging does for our code. It quietly records the journey of the script so we can look back later and understand how everything unfolded. In this part of the code, we’re creating a log system using Python’s built-in logging module. We start by making a folder named logs. This will be our storage room—it's where all the log files will live. The line log\_dir.mkdir(exist\_ok=True) ensures that if this folder doesn’t exist already, Python will create it for us. And if it’s already there, that’s fine too—it won’t throw any errors or complaints. Then, we build the name for our log file using the current working directory’s name. That’s what this line is doing: log\_file=log\_dirf"scrape\_log\_{Path().cwd().name}.log". Think of it like naming your notebook based on the kitchen you’re baking in. This helps keep our log files organized, especially if we’re running the scraper in different folders or for different projects. Next, we tell Python how to write in this log file. Using logging.basicConfig(...), we define a few things: where to save the log (filename=log\_file), how detailed the messages should be (level=logging.INFO), and the format of each log entry. This format includes the date and time something happened, the type of message (like INFO, WARNING, or ERROR), and a short message describing what occurred. Why is all of this important? Well, imagine your scraper is running for hours and suddenly stops. Without logging, you’d be left guessing what went wrong. But with logging, you can open the log file and see exactly what the script was doing before it crashed. Maybe a page didn’t load in time, or the internet disconnected briefly. Logging keeps a reliable diary of every action, which is incredibly helpful for both beginners and experienced developers alike. So before we even send our scraper out into the wild, we’re already building a system to help us understand and debug it later. It’s like packing a travel journal before a trip—you might not need it immediately, but when something unexpected happens, you’ll be glad it’s there. ### Defining Global Configuration and Constants ``` # DATABASE INITIALIZATION async def init_db(): """ Initializes the SQLite database by creating a table for storing product URLs. """ conn = sqlite3.connect(DB_PATH) c = conn.cursor() c.execute(""" CREATE TABLE IF NOT EXISTS product_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE ) """) conn.commit() conn.close() ``` Now that we’ve told our scraper where to go and how to behave, it’s time to prepare a place to store the information it collects. Think of this step as setting up a filing cabinet before you start sorting documents. In programming, that filing cabinet is often a database, and for small projects like ours, SQLite is a perfect choice. It’s lightweight, doesn’t require any server setup, and everything is saved in a simple file on your computer. The function init\_db() takes care of this setup. As the name suggests, it initializes the database—meaning it creates the structure we need to start storing data. But don’t worry, it’s not as complicated as it sounds. This function gently checks if our database already exists, and if it doesn’t, it builds the table we need. That table is called product\_urls, and it’s where all the product page links we scrape will be saved. Inside this table, we define two columns. The first is id, which is like a serial number that automatically increases each time we add a new row. You don’t need to think about this one too much—it’s just there to help keep everything uniquely identified. The second column is url, which will store the actual link to a product page. This column is marked as UNIQUE, which means no two entries can be the same. This is really helpful in web scraping because it prevents us from accidentally saving the same link more than once. The function works step-by-step, just like how you’d open a notebook, draw a table, and label the columns. First, it connects to the database file at the location we defined earlier (DB\_PATH). If that file doesn’t exist yet, SQLite will quietly create it. Then, we create a cursor, which is like a pen used to write commands inside the database. We use that cursor to write a CREATE TABLE command, which basically says, “If this table isn’t already here, make one now.” Once that’s done, we save the changes with commit() and close the connection—just like closing the notebook when you're done writing. This setup is a one-time task. You only need to run init\_db() once before scraping begins, to make sure everything is in place. After that, your scraper will know exactly where to put the URLs it collects, and it won’t bother saving the same one twice. It’s a quiet but crucial part of the process that ensures our data stays organized, tidy, and ready for analysis later. ### Saving Scraped URLs to the Database ``` # SAVE URL TO DATABASE async def save_to_db(url): """ Saves a single product URL to the SQLite database. """ conn = sqlite3.connect(DB_PATH) c = conn.cursor() try: c.execute("INSERT OR IGNORE INTO product_urls (url) VALUES (?)", (url,)) conn.commit() finally: conn.close() ``` Once our database is set up and ready to receive data, the next step is teaching our scraper how to actually store something in it. That’s where the save\_to\_db() function comes in. Imagine this function as a responsible assistant that carefully writes each new product link into a notebook—making sure it doesn’t write the same thing twice. This function is quite simple in its purpose: it takes in one URL—the link to a product page—and saves it to our database. That’s it. But it does it with care. The input to this function is a single string, the URL we’ve scraped from the website. We pass that URL into our database using a command called INSERT OR IGNORE. This phrase is key because it keeps things clean: if we accidentally try to insert the same URL more than once, SQLite will politely ignore it instead of throwing an error or cluttering our table with duplicates. Behind the scenes, the function starts by connecting to the database file we defined earlier with DB\_PATH. Once connected, we create a cursor—this is what allows us to send instructions to the database. Then, using that cursor, we run the SQL command to insert the URL. The (?,) part in the SQL line is a placeholder for the actual value, and it’s filled in with the url we passed. This method not only prevents errors but also helps protect against SQL injection—a common security issue. Notice how we wrap the database operations in a try block with a finally clause. Even though this function is small, we still want to be safe and responsible. The finally block ensures that the database connection is always closed, no matter what. Think of it like turning off the lights and locking the door when you leave a room—even if something unexpected happens. In short, save\_to\_db() is a quiet worker. It doesn’t try to do too much. It takes one piece of data, checks if it’s new, and files it away if it hasn’t seen it before. This kind of function becomes incredibly useful when scraping hundreds or thousands of links, helping us avoid duplication and making sure every piece of data has its proper place. ### Store URLs Safely with JSON Backup ``` # SAVE URL TO JSON async def save_to_json(url_list, page_num): """ Saves a list of URLs from a single page into a JSON file. """ json_path = Path(f"Data/data1/ethos_page_{page_num}.json") with json_path.open("w") as f: json.dump(url_list, f, indent=2) ``` After collecting product links from a webpage, we often want to keep a backup—not just for safety, but also for reviewing, sharing, or using the data outside of the database. That’s where our save\_to\_json() function comes into play. Think of this step as saving your progress in a game, or keeping a soft copy of your handwritten notes. While our database stores everything in a structured, query-friendly format, the JSON file gives us a more portable and human-readable version of the same information. This function takes two inputs: a list of URLs (url\_list) and the page number (page\_num) they were scraped from. The page number is especially useful because it helps us organize the files. Instead of dumping everything into one huge document, we save each page’s data separately. This keeps things neat and makes it easier to debug or resume scraping later if something goes wrong. Inside the function, we first define the path where this JSON file should be saved. Using Python’s Path from the pathlib module, we construct a clean, consistent file location—something like Data/data1/ethos\_page\_3.json, where “3” is the current page number. It’s a simple but powerful way to make our files self-explanatory just by their names. Then we use Python’s built-in json module to handle the actual saving. The with open() block ensures the file is opened safely and closed properly when we’re done writing. The json.dump() function takes our list of URLs and writes them into the file in JSON format. The indent=2 part just makes the file prettier and easier to read if we open it later—sort of like adding clean line breaks and indentation in your notebook. This small function may not seem flashy, but it’s incredibly practical. By saving each batch of URLs page by page, we’re giving ourselves a safety net. If the scraper stops midway or we need to inspect specific data later, we don’t have to rerun everything from scratch. Each JSON file becomes a snapshot of that moment in the scraping process—organized, readable, and ready to use. ### Smart Pop-up Handler ``` # HANDLE UNEXPECTED POPUPS async def close_popups(page): """ Attempts to close any pop-up elements that may appear and block access to the main content of the web page. """ selectors = [ 'div[role="button"][aria-label="Close"]', 'a.mdl-cls-btn.ctClickNew' ] for selector in selectors: try: await page.locator(selector).first.click(timeout=2000) logging.info(f"Closed popup: {selector}") except: pass ``` While scraping a website, things don’t always go exactly as planned. Sometimes, just as your script is trying to read or click something on the page, an unexpected pop-up appears—like a newsletter sign-up, a promotional offer, or a cookie consent banner. If you’ve ever visited a shopping website and been greeted by a sudden overlay blocking the content, you already know how frustrating these pop-ups can be. For a human, it’s easy to click the close button and move on. But for a scraper, unless we teach it what to do, it gets stuck—unable to move forward. That’s exactly why we have the close\_popups() function. This function is a simple but thoughtful piece of our scraper that plays the role of a quiet troubleshooter. It receives the current webpage (or "tab") being controlled by Playwright and begins scanning for pop-ups that might be hiding parts of the site we want to scrape. To do that, it goes through a small list of CSS selectors—these are patterns that help it find specific elements on the page, such as the close button on a pop-up window. For each selector in the list, the function tries to locate the first matching element. If it finds it, it clicks the button to close the pop-up. If the element isn’t found—maybe the pop-up didn’t show up this time, or it was already dismissed—the function simply moves on without raising an error or stopping the script. This approach keeps the scraper flexible and resilient, ready to handle unpredictable behaviors on the site without breaking. What makes this function so valuable is its subtlety. It doesn’t scrape data or save anything, but it quietly clears the way so that everything else can work properly. Pop-ups often sit right on top of product listings or navigation buttons. If we don’t close them, our script might fail to click on the “Next” button or miss a set of links. By handling these interruptions in advance, close\_popups() helps ensure that the scraping flow remains smooth and uninterrupted, no matter what the website throws at us. ### The Core Scraping Function ``` # MAIN ASYNCHRONOUS SCRAPING FUNCTION async def scrape_ethos(): """ Main asynchronous scraping function for Ethos Watches website. """ await init_db() async with async_playwright() as p: browser = await p.firefox.launch(headless=False) context = await browser.new_context(extra_http_headers=HEADERS) page = await context.new_page() current_page = 1 total_scraped = 0 while True: url = f"{START_URL}?p={current_page}" if current_page > 1 else START_URL logging.info(f"Navigating to: {url}") retry = 0 success = False while retry < 3 and not success: try: await page.goto(url, timeout=60000) await close_popups(page) success = True except PlaywrightTimeoutError: retry += 1 wait = 2 ** retry + random.uniform(0, 2) logging.warning(f"Timeout on page {current_page}. Retrying in {wait:.1f}s...") await asyncio.sleep(wait) if not success: logging.error(f"Failed to load page {url} after retries.") break product_desc_blocks = await page.query_selector_all("div.product_sortDesc") product_urls = [] for block in product_desc_blocks: try: a_tag = await block.query_selector("a") href = await a_tag.get_attribute("href") if a_tag else None if href and href.startswith("https://www.ethoswatches.com/"): await save_to_db(href) product_urls.append(href) except Exception as e: logging.warning(f"Error parsing product block: {e}") await save_to_json(product_urls, current_page) logging.info(f"Page {current_page}: Scraped {len(product_urls)} product URLs.") total_scraped += len(product_urls) next_button = await page.query_selector('a.next.page-link') if next_button: current_page += 1 delay = random.uniform(3, 6) logging.info(f"Waiting {delay:.2f}s before next page...") await asyncio.sleep(delay) else: logging.info("No more pages.") break await browser.close() logging.info(f"Total products scraped: {total_scraped}") ``` After setting up all the building blocks—defining global variables, preparing the database, handling pop-ups, and saving data to files—we’re finally ready to bring everything together in one place. That’s the job of our main function, scrape\_ethos(). This function is the heart of the scraping process. You can think of it as the conductor of an orchestra, calling on each instrument—our earlier functions—to play its part at just the right time. And because the web is dynamic and involves waiting for pages to load, we use Python’s asynchronous tools to manage everything efficiently and smoothly. We begin by calling init\_db(), which makes sure our database is set up and ready to store product URLs. Then we launch a web browser using Playwright—a tool that lets us control a browser as if a human were using it. Instead of hiding the browser (as scrapers often do to stay fast), we use it in visible mode (headless=False). This is useful while testing, so we can see what’s happening on each page. Next, the function starts visiting each page of the website one by one. Ethos Watches uses pagination, meaning that product listings are spread across multiple pages. To move through them, we update the page number in the URL using a loop. For each page, we use a while loop to handle retries. Websites don’t always load smoothly—sometimes they’re slow or have temporary hiccups. So if the page doesn’t load the first time, we wait a bit and try again. This retry logic uses exponential backoff, a fancy way of saying, “wait a little longer each time you fail, but don’t give up too quickly.” Once the page loads, we immediately call our close\_popups() function. This checks if any unwanted pop-ups are covering the content and closes them so we can continue scraping without interruptions. Then we search for all the blocks that contain product information. These are usually HTML elements that follow a certain pattern, in this case, div.product\_sortDesc. From each block, we try to pull the link ( tag) that leads to the individual product’s page. We check if the link is valid and starts with the correct base URL to ensure we’re only saving proper product URLs. Each valid link is saved in two places—first into the SQLite database using save\_to\_db(), and then into a JSON file using save\_to\_json(). This two-pronged saving approach gives us both a structured database for analysis and a human-readable file for quick inspection or backups. After collecting all the links from a page, the scraper checks if there is a “Next” button. If so, it waits a few seconds (randomly between 3 and 6) to mimic human browsing, then moves to the next page. If the “Next” button isn’t there, the scraper understands that it has reached the end and stops gracefully. Finally, when all pages have been processed, we close the browser and print the total number of products scraped. This final message is a small but satisfying checkpoint—it lets us know everything has completed as expected. In the end, this function not only automates the entire scraping process from start to finish, but it also does so in a way that’s thoughtful, reliable, and respectful of the website being visited. ### Script Entry Point ``` # SCRIPT ENTRY POINT if __name__ == "__main__": """ Entry point of the script. """ asyncio.run(scrape_ethos()) ``` When you're writing a Python script, it's important to have a clear starting point—something that tells the computer, "Begin here." In this script, the block that starts with if name == "\_\_main\_\_": is exactly that. Think of it like the front door to your program. If someone runs this script directly, this block is what gets executed. But if someone just wants to borrow part of your script (maybe they want to use your functions in their own code), then this section will stay quiet and not run automatically. That’s what makes this line so useful—it helps control when your code should actually do something. Inside this block, there's a single line: [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)(scrape\_ethos()). This is where the real action begins. The scrape\_ethos() function is the core of your scraping process—it’s where the browser launches, pages are visited, and product links are collected. But because it’s written as an asynchronous function (which is useful when you're doing tasks like waiting for websites to load), you can’t just call it like a regular function. That’s where [asyncio.run](http://asyncio.run/?ref=blog.datahut.co)() comes in. It sets up the necessary environment and tells Python, “Okay, this is an async function, so let’s run it properly.” It’s a bit like turning on the engine before you can drive a car. Using this entry-point pattern might feel like a small detail, but it’s actually a good habit for anyone writing Python scripts—especially as your code grows or you work in teams. It keeps things clean, flexible, and prevents surprises when your code is used in new ways. So, even though it's just a few lines, it plays a key role in making sure your scraping project starts at the right time, in the right way. ## Step 2: Turning URLs into Complete Product Profiles ### Imports and Initial Setup ``` # IMPORT REQUIRED LIBRARIES import asyncio import json import logging import sqlite3 from pathlib import Path from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError from playwright_stealth import stealth_async import re import random ``` The script begins by importing a set of essential Python libraries—both built-in and third-party—that provide the backbone for everything that follows. Modules like sqlite3, json, and logging help manage data storage, structure, and track the script’s progress. asyncio supports asynchronous operations, which allow the script to perform tasks like web browsing without freezing or waiting unnecessarily. The powerful playwright and playwright\_stealth libraries handle browser automation and help avoid detection while scraping. Additional helpers like Path, re, and random make tasks like handling file paths, working with patterns, and introducing randomness smooth and reliable. These imports set the stage for a well-coordinated, efficient web scraping process. ### Essential Paths for Database, JSON, and Logs ``` # CONSTANTS FOR FILE AND DATABASE PATH DB_PATH = "ethos_products.db" JSON_PATH = "ethos_product_data.json" LOG_PATH = Path("logs/scrape_details.log") LOG_PATH.parent.mkdir(exist_ok=True) """ This section defines key file paths used throughout the scraper: """ ``` Before diving into the actual scraping work, it's important for our script to set up a few things behind the scenes—much like laying out your tools before starting a project. In this case, we begin by defining a few constants that tell the script where to save different types of information as it runs. First, we specify where the scraped data will be stored using DB\_PATH. This points to a local SQLite database file named ethos\_products.db. Think of this database like a digital notebook with two main sections—one for just keeping track of product page links (product\_urls), and another for saving the full details scraped from each of those pages (product\_data). Having this separation helps the script stay organized and know what has already been done versus what’s still pending. Next is JSON\_PATH, which refers to a file named ethos\_product\_data.json. This file acts like a backup copy of the product data, but in a format that’s easy to read and share. JSON files are especially useful if you want to later open the data in tools like Excel or convert it into other formats. Then we have LOG\_PATH, which leads to a log file stored inside a folder called logs/. Logs are like a running diary for the script—they record everything from successful actions to unexpected problems. This helps us trace issues later or just understand how the script performed over time. Lastly, we make sure that the logs/ folder actually exists before the script tries to save anything there. The line LOG\_PATH.parent.mkdir(exist\_ok=True) takes care of this by quietly creating the folder if it’s missing. And if it’s already there, the script simply moves on without complaint. Altogether, this section is about setting up smart, reusable paths so that our scraper knows exactly where to place its output, how to track its progress, and where to look if something goes wrong. It’s a small but important step in making the script stable, organized, and beginner-friendly. ### Logging Configuration for Debugging and Monitoring ``` # SETUP LOGGING FOR DEBUGGING AND TRACKING logging.basicConfig( filename=LOG_PATH, level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s" ) """ This section sets up logging to help monitor the scraping process. """ ``` To keep track of what the script is doing at each step, we set up something called logging. Think of logging like keeping a diary for the scraper—it notes down what’s happening, when it happens, and if anything goes wrong. This is especially useful when scraping large websites or running long scripts, where things might fail quietly if we’re not paying attention. In this setup, every log message is saved to a file named scrape\_details.log, stored in a folder called logs. ### Organizing Storage for Product Data ``` # DATABASE SETUP def setup_database(): """ Create and update the required database tables. """ conn = sqlite3.connect(DB_PATH) cur = conn.cursor() # Add processed column to product_urls table cur.execute("PRAGMA table_info(product_urls)") columns = [row[1] for row in cur.fetchall()] # Add 'processed' column if missing (default = 0 → not scraped) if "processed" not in columns: cur.execute("ALTER TABLE product_urls ADD COLUMN processed INTEGER DEFAULT 0") # Create structured table for scraped product data cur.execute(""" CREATE TABLE IF NOT EXISTS product_data1 ( url TEXT PRIMARY KEY, name TEXT, brand TEXT, mrp TEXT ) """) conn.commit() conn.close() ``` Before we start collecting any product data, we need to prepare a place to store everything—and that’s where the database setup comes in. Think of it like setting up a filing cabinet before you begin sorting and storing documents. If the cabinet isn’t ready or doesn’t have the right folders, things can easily get lost or mixed up. This part of the script handles that preparation for us. In this function, we connect to an SQLite database—basically, a small file-based system where we’ll keep track of the URLs we want to scrape and the product details we extract. First, we check if a column called processed exists in the product\_urls table. This column acts like a checkbox that helps us remember whether we’ve already scraped a product URL. If it’s not there, we add it. Any URL with processed = 0 means we still need to visit it, while processed = 1 means we’ve already handled it. Next, we set up a new table called product\_data. This is where the detailed product information—like name, brand, price, and even things like case shape or power reserve—will be stored. Each row in this table represents one product, and the URL is used as a unique identifier to avoid saving duplicates. By running this setup function before we begin scraping, we make sure everything is neatly organized. The database will be ready to record and track every product we scrape without confusion or duplication. It also helps us safely pause and resume scraping without starting over, since we already know which URLs have been processed. This makes the whole process much smoother and far less error-prone. ### Mark URLs as Processed ``` # UPDATE PROCESSED STATUS IN DATABASE def update_processed_status(url):    """    Marks a product URL as "processed" in the database after it has been successfully scraped.    """       conn = sqlite3.connect(DB_PATH)    cur = conn.cursor()    cur.execute("UPDATE product_urls SET processed = 1 WHERE url = ?", (url,))    conn.commit()    conn.close() ``` After scraping data from a product page, it’s important for our script to remember that the job is done for that specific page. Otherwise, the next time the script runs, it might go back and repeat the same work, which wastes time and can lead to duplicate entries in our database. That’s where this small but essential function, update\_processed\_status, comes into play. Think of it like checking off a task on a to-do list. Once we’ve visited a product page and saved its details, we want to mark it as “complete.” This function connects to our SQLite database and updates the corresponding product URL’s status to show that it has already been handled. It does this by setting the processed column to 1, which simply means “done.” The function works quietly in the background. When you pass in a URL, it opens the database, finds the matching entry in the product\_urls table, and marks it as processed. It then saves the change and closes the connection, making sure everything is neat and tidy before moving on. Why is this so important? Imagine scraping thousands of product URLs—without a reliable way to track which ones have already been scraped, the script would either reprocess everything from scratch or get confused. By using this approach, we give our scraper a memory. It knows exactly where it left off and can pick up from there, especially helpful if the script is interrupted or needs to be restarted later. In short, this function keeps our data collection process organized and efficient, helping us avoid repetitive work and maintain clean, accurate results. ### Store Scraped Entries in JSON ``` # SAVE SCRAPED DATA TO JSON FILE def save_to_json(json_path, entry):    """    This function appends a single scraped product entry (in dictionary format) to a JSON file.    """    if Path(json_path).exists():        with open(json_path, "r", encoding="utf-8") as f:            data = json.load(f)    else:        data = []    data.append(entry)    with open(json_path, "w", encoding="utf-8") as f:        json.dump(data, f, indent=4) ``` After collecting product information from a webpage, we need a safe and accessible place to store that data. While databases are excellent for structured storage, sometimes it's also helpful to have a simple file you can open and read directly—something you can easily share, move around, or check with your own eyes. That’s exactly what this save\_to\_json function is designed to do. Think of a JSON file like a notebook where each page holds details about one product. This function takes a single product entry (structured as a Python dictionary), and adds it to that notebook. If the file doesn’t already exist, it creates one from scratch. If it does exist, the function opens it up, reads all the previous entries, adds the new one to the list, and then saves everything back neatly in the same file. Here’s how it works step by step. First, the function checks if the file (given by json\_path) already exists on your computer. If it does, it loads all the existing entries into a list. If not, it just starts fresh with an empty list. Then, it appends the new product information to that list—like adding one more page to our notebook. Finally, it writes the entire updated list back to the file, formatting it with indentations to keep things clean and easy to read. This approach is especially helpful if you want a backup of your scraped data or if you're not ready to work with a database just yet. JSON files are flexible, human-readable, and compatible with many tools. Later, you can even convert them to CSV or Excel if needed. So in a nutshell, this function helps keep your data safe, organized, and always within reach—even outside your code. ### Popup Management for Smooth Scraping ``` # HANDLE UNEXPECTED POPUPS async def close_popups(page):    """    Attempts to close any pop-up windows that may appear on Ethos product pages.    """    # First popup    try:        await page.locator("div#close[role='button']").click(timeout=2000)        logging.info("First popup closed")    except:        pass    # Second popup    try:        await page.locator("a[onclick='cnscbEthClose()']").click(timeout=2000)        logging.info("Second popup closed")    except:        pass ``` When you're building a scraper to collect product details from a website, it's not always a smooth ride. Sometimes, pages throw up unexpected pop-ups—like cookie notices, promotional banners, or welcome messages—that cover the content you're trying to grab. These popups can block important product details and interrupt your scraper's flow, causing it to either miss key information or crash entirely. That’s where this function, close\_popups, steps in to help. Imagine you’re trying to read a book, but every few pages, someone waves an ad in front of your face. You’d have to gently push it aside before continuing. Similarly, this function quietly checks for two known types of popups that appear on the Ethos product pages. It does this by looking for specific HTML patterns—kind of like scanning a page for a "close" button and clicking it automatically if it's there. It works asynchronously, meaning it runs alongside other tasks without blocking them. That makes it suitable for fast-paced web scraping, where every second counts. The function tries to click each popup’s close button using a short timeout of two seconds. If the popup isn’t found or doesn’t close in time, it doesn’t throw a tantrum—it just moves on silently. This helps keep your scraping process steady and uninterrupted. Why is this important? Because when a popup covers the screen, even if your code is perfect, it might still fail to fetch the data. For example, if a cookie banner is sitting on top of the "Add to Cart" button or product name, your scraper might think the element doesn't exist or isn't clickable. By handling these interruptions upfront, close\_popups ensures your scraper has a clear view of the page—just like cleaning your glasses before reading. You typically call this function right before you start extracting data from a page. Just write await close\_popups(page), and it will do its job quietly in the background. It’s one of those small touches that makes your scraping setup much more reliable and professional, especially when dealing with unpredictable websites. ### Retrieve Label-Based Data from Pages ``` # EXTRACT A SPECIFIC PRODUCT SPECIFICATION VALUE async def extract_spec_value(page, label):    """    Extracts the value of a specific product specification from a product detail page.      """    try:        # Find by specRow or li.calibre_sepcColumn        element = await page.query_selector(f"xpath=//div[@class='specRow'][span[@class='specName' and contains(text(), '{label}')]]/span[@class='specValue']")        if element:            return (await element.inner_text()).strip()        # Try li fallback        element = await page.query_selector(f"xpath=//li[@class='calibre_sepcColumn specRow'][span[@class='specName' and contains(text(), '{label}')]]/span[@class='specValue']")        if element:            return (await element.inner_text()).strip()    except Exception as e:        logging.warning(f"Failed to extract {label}: {e}")    return None ``` In web scraping, especially when dealing with product pages, it’s often necessary to pull out specific details—like the size of a watch, its strap color, or the type of movement inside. These pieces of information usually sit in a structured part of the webpage called a specification section. Think of it like a table where each row shows a label and its value—for example, “Case Size” on the left and “42 mm” on the right. Now, the challenge here is that not all web pages structure these rows the same way. Some use one kind of HTML layout, while others might use a different one altogether. So, to handle this smartly, we use a flexible function called extract\_spec\_value. This function is designed to take in a page (which is the product page we’re looking at) and a label (like “Strap Color”), then go and find the value that matches that label. It does this by first looking for a block of HTML with a class name called specRow. Inside that, it checks if there’s a span with the class specName that includes the label we’re searching for. If it finds that, it pulls the value from the corresponding specValue span. But the page might be using a different structure, so the function is smart enough to try a second layout as a fallback—specifically, a list item (li) with a slightly different class name. This two-step search ensures we don’t miss the data just because the structure changes a little. The beauty of this function is that it hides all the technical details behind a clean, readable interface. You simply ask, “What’s the Case Size?” and it brings you back the answer, if it exists. If it can’t find the label or something goes wrong during the search, it quietly logs a warning and returns None. This way, the scraping process doesn’t crash—it just skips that bit and keeps moving. As a result, the rest of the data collection can continue uninterrupted, and we still have a note in the logs about what went wrong, in case we want to fix or revisit it later. Functions like this one help keep our scraping code clean, reusable, and resilient—even when websites throw us curve balls with changing layouts. ### Fetching and Organizing Product Info ``` #  EXTRACT ALL PRODUCT DETAILS FROM THE PAGE async def extract_data(page): """ Extracts detailed product information from an Ethos Watches product detail page. """ data = {} try: # Name title = await page.query_selector("h1.ethos_title span.fWeight_regular") data["name"] = await title.inner_text() if title else None # Brand brand = await page.query_selector("div.specCol span.specName:text('Brand') + span.specValue a") if not brand: brand = await page.query_selector("h1.ethos_title a") data["brand"] = await brand.inner_text() if brand else None # MRP price = await page.query_selector("span.price") data["mrp"] = await price.inner_text() if price else None except Exception as e: logging.error(f"Error extracting data: {e}") return data ``` At the heart of our scraping project, there’s a function called extract\_data(). This is the part of the script responsible for visiting a single product page on the Ethos Watches website and pulling out all the important details about the watch listed there. Think of it as the person who walks into a store, reads every tag on a product, notes down everything neatly, and leaves. That’s what this function does—but on the web. Now, since websites are built using HTML, we can target specific parts of the page using tools like Playwright, which lets us control a browser with code. And since websites sometimes load information slowly or use a lot of background scripts, we use asynchronous programming here. That just means we let our code “wait patiently” for things to load, without freezing the whole program. Inside the function, we start with a blank dictionary called data. This is like an empty notepad where we’ll write down each detail we find about the product. We begin by trying to get the name of the product. Usually, it’s found in the main heading at the top of the page. So, we use Playwright to find that element using something called a CSS selector—basically, a way to point to specific parts of a webpage. If we find the name, we grab the text and store it in our dictionary. Next, we look for the brand of the watch. Sometimes it's right there under the "Brand" label; other times, it might be part of the title. So, we try both possibilities to make sure we don’t miss it. This “try multiple ways” approach is common in scraping because websites aren’t always perfectly consistent. Then we move on to the model number. This could be sitting inside a tag with the ID [#psku](https://www.blog.datahut.co/blog/hashtags/psku), or if it's not there, we ask our helper function extract\_spec\_value() to look it up for us in the specifications section. That helper is smart—it knows how to look for labels and get their values, which is great for structured data. We do something similar for the price, or MRP. We look for a span with the class “price”, grab the number if it exists, and store it. Now comes the bulk of the work: all the technical specifications of the watch. This is where the function shines. We’ve prepared a list of fields we care about—things like “Strap Color”, “Dial Colour”, “Movement”, “Power Reserve”, and many others. For each of these, we loop through and call our extract\_spec\_value() helper. This helper searches the page for that label and grabs whatever value is next to it. To keep our keys in the dictionary consistent and clean, we lowercase each label and replace spaces with underscores. So, for example, “Strap Material” becomes strap\_material—easier to work with in code later. Throughout this whole process, we wrap everything in a try block. That’s a safety net. If something goes wrong—maybe the website structure changes, or an element is missing—we don’t want the whole program to crash. Instead, we log the error, skip that one detail, and continue collecting the rest. That way, we don’t lose everything just because of one hiccup. Finally, once all the available information is collected, the function returns the data dictionary. At this point, it's like a neatly filled-out form with all the details of the watch, ready to be saved into a database or analyzed further. In simple terms, this function behaves like a reliable assistant: it opens a product page, reads every label carefully, grabs the matching information, and organizes it all into a clean, consistent format. It handles surprises gracefully and keeps working even when some data is missing. This kind of organized, flexible approach is what makes a scraper both powerful and dependable—even when the website it's working on isn't perfect. ### Main Asynchronous Scraper ``` #  MAIN ASYNCHRONOUS SCRAPING FUNCTION async def scrape(): """ Main asynchronous function that controls the entire scraping workflow. """ setup_database() conn = sqlite3.connect(DB_PATH) cur = conn.cursor() cur.execute("SELECT url FROM product_urls WHERE processed = 0") urls = [row[0] for row in cur.fetchall()] conn.close() if not urls: logging.info("No unprocessed URLs found.") return logging.info(f"Total URLs to process: {len(urls)}") async with async_playwright() as p: browser = await p.chromium.launch(headless=False) context = await browser.new_context() page = await context.new_page() await stealth_async(page) for idx, url in enumerate(urls, 1): try: logging.info(f"[{idx}] Navigating to {url}") await page.goto(url, timeout=30000) await asyncio.sleep(random.uniform(1, 3)) await close_popups(page) data = await extract_data(page) data["url"] = url # Save to DB conn = sqlite3.connect(DB_PATH) cur = conn.cursor() columns = [ "url", "name", "brand", "mrp" ] values = [data.get(col, None) for col in columns] cur.execute(f""" INSERT OR REPLACE INTO product_data1 ({', '.join(columns)}) VALUES ({','.join('?' * len(columns))}) """, values) conn.commit() conn.close() # Save to JSON save_to_json(JSON_PATH, data) # Mark as processed update_processed_status(url) logging.info(f"[{idx}] Scraped and saved: {url}") except PlaywrightTimeoutError: logging.error(f"[{idx}] Timeout while loading {url}") except Exception as e: logging.error(f"[{idx}] Unexpected error: {e}") await browser.close() ``` Let’s imagine you’re tasked with collecting detailed product information from a website—like a digital assistant that visits a page, reads every detail carefully, writes it down, and moves on to the next page without repeating itself. That’s exactly what the scrape() function is designed to do. It acts like the brain behind our scraping operation, coordinating every step from start to finish. To begin with, this function is asynchronous, which simply means it can perform multiple tasks at once without waiting for one to finish before starting another. This is especially helpful when you're dealing with the internet, where delays are common. Think of it like reading a book while waiting for water to boil—you're making good use of your time. This kind of efficiency is what async def scrape() brings to the table. Now, before diving into any scraping, the function prepares the environment by calling another function named setup\_database(). This ensures all necessary tables in the database are ready. Then it connects to our SQLite database and fetches all product URLs that haven’t been visited yet—these are marked with processed = 0\. If there are no URLs left, it logs that there’s nothing to do and gently exits. This is a thoughtful checkpoint to avoid unnecessary effort. When there are URLs to process, the real journey begins. We launch a Chromium browser using Playwright, which is a tool that can control web browsers just like a human would—with clicks, scrolls, and even typing. But here’s the twist—we also apply a stealth mode using stealth\_async(page). This makes our scraper look and behave like a real user, helping us avoid getting blocked by the website’s defenses. Once the browser is ready, we loop through each URL one by one. For every product page, we navigate to it and wait for it to load. Since many websites now have pop-ups (like newsletter signups or cookie warnings), we call close\_popups(page) to get them out of the way, just like you would click the 'X' before browsing a page. Next comes the heart of the scraping—extract\_data(page). This function collects everything we care about, such as the product’s name, brand, price, model number, strap material, and much more. After gathering this information, we also tag it with the URL it came from, so we always know the source. Once the data is in hand, we save it in two places. First, we insert it into a structured SQLite database, which acts like a local spreadsheet where each row is a product and each column holds specific details. Then, we also save the same data into a JSON file. This acts like a backup or an easy-to-share export that other programs or people can use later. After saving the data, we make a note that the URL has been processed by calling update\_processed\_status(url). This is like checking off a task on your to-do list—it ensures we don’t visit the same page again in future runs. Throughout the entire process, we also maintain detailed logs. Whether the scraping is successful, a timeout happens, or an unexpected error pops up, we log it all. This is extremely helpful when debugging or resuming a scrape that was interrupted. Finally, after all URLs are processed, we close the browser gracefully. Just like shutting down your computer at the end of the day, this step ensures all resources are freed up and everything ends cleanly. In summary, the scrape() function ties together many moving parts into one well-orchestrated routine. It ensures we only visit unprocessed URLs, collect rich product data carefully, save it reliably, and avoid repeating our work. For anyone starting out in web scraping, this function offers a complete, real-world example of how to build a smart, efficient, and polite scraper—one that’s both effective and respectful to the website it visits. ### Starting the Scraper (Main Entry Point) ``` # SCRIPT ENTRY POINT if name == "__main__":    """    Program entry point. Runs the async scrape function and logs fatal errors if any.      """    try:        asyncio.run(scrape())    except Exception as e:        logging.critical(f"Fatal error: {e}") ``` At the very end of the script, we have a simple block that acts like a "start button." When you run this script directly, it kicks off the main scraping process. In our case, this is handled by the if name == "\_\_main\_\_": block. If anything goes seriously wrong, it logs the error so you know what happened. ## Conclusion In the world of online retail, understanding detailed product information can make a real difference—whether you're tracking trends, comparing brands, or building your own catalog. This blog walked through how we can extract rich product details from Ethos Watches using Playwright and asynchronous Python code. By carefully navigating each product page, reading labels, and collecting specifications like price, strap material, or movement type, we’ve created a reliable system to turn unstructured website content into clean, organized data. For beginners or interns just starting out, this blog shows that with the right tools and a step-by-step approach, even a dynamic website can be broken down into meaningful insights that power smarter analysis and applications. ## Libraries and Versions Used Name: asyncio Version: Built-in Python module (no separate installation required) Name: json Version: Built-in Python module (no separate installation required) Name: random Version: Built-in Python module (no separate installation required) Name: sqlite3 Version: Built-in Python module (no separate installation required) Name: logging Version: Built-in Python module (no separate installation required) Name: pathlib Version: Built-in Python module (no separate installation required) Name: playwright Version: 1.48.0 ### AUTHOR I’m Anusha P O, Data Science Intern at Datahut. I specialize in building smart scraping systems that automate large-scale data collection from complex e-commerce websites. In this blog, I walk you through how we extracted and structured detailed product information from Ethos Watches using Playwright, SQLite, JSON, and asynchronous Python workflows—turning intricate product pages into clean, analysis-ready datasets. At Datahut, we help businesses unlock the full potential of web data by designing robust, scalable scraping solutions for competitive intelligence, product research, and market analysis. If you’re exploring data-driven strategies for e-commerce or want to organize large product datasets efficiently, reach out via the chat widget on the right. Let’s transform your raw web data into actionable insights. FAQ SECTION ### 1\. Is it legal to scrape Ethos product data using Python? Scraping Ethos product data can be legal if it is done responsibly and in compliance with Ethos’ terms of service, robots.txt rules, and applicable data protection laws. Publicly available product information such as prices, brands, and availability is generally safer to scrape for research or analysis purposes. ### 2\. Why should async tools be used for scraping Ethos product data? Async tools allow multiple requests to be processed simultaneously, making scraping faster and more efficient. When scraping Ethos, which may have multiple product pages and categories, async frameworks significantly reduce scraping time while maintaining performance. ### 3\. Which Python libraries are best for async web scraping? Popular Python libraries for async web scraping include aiohttp, asyncio, Playwright, and httpx. These tools help manage concurrent requests efficiently and handle dynamic content commonly found on eCommerce websites like Ethos. ### 4\. How can I avoid getting blocked while scraping Ethos? To avoid blocks, use techniques such as rotating user agents, setting proper request delays, limiting request frequency, and handling headers correctly. Async scraping should be carefully throttled to mimic human-like browsing behavior. ### 5\. What type of product data can be scraped from Ethos? You can scrape product-related data such as product names, brand details, prices, discounts, ratings, availability, product descriptions, and category information. This data can be used for competitive analysis, price monitoring, and market research. ### Why Open Source Web Scraping Tools Fail at Scale URL: https://www.blog.datahut.co/post/open-source-web-scraping-tools/ Last updated: 2026-09-07T09:43:45.000Z Picture this: Your team builds a beautiful internal scraping platform using Open Source libraries. It scrapes 20 e-commerce sites, powers dashboards, feeds pricing models… and becomes part of your company’s heartbeat. You scale from 10K → 100K → 1M pages per day. Suddenly: - your prices stop updating - your stock signals lag - your competitor feeds look “too perfect” - your alerts never fire - your data scientists complain about anomalies - and your engineering team starts firefighting daily You didn’t “break” anything. You simply pushed Open Source tools past their natural shelf life — a limit most teams only discover when it’s too late. If you're using Open Source tools for large-scale web scraping — stop and read this first. Open Source web scraping tools are brilliant, and we’ve written about their strengths in our deep-dive on[ ](https://www.blog.datahut.co/post/web-scraping-vs-api/)[Web Scraping vs API](https://www.blog.datahut.co/post/web-scraping-vs-api/) for teams comparing extraction strategies. They democratized scraping, taught millions, and helped founders ship prototypes fast. But here’s the truth almost everyone discovers too late: Open Source scraping tools break fast — and the moment you scale to millions of records, that short shelf life becomes a serious business risk. That’s why even experienced engineering teams – and many web scraping companies themselves – eventually discover that a quick Open Source stack is very different from a battle-tested, production-grade scraping platform. Not because open source is bad.Not because maintainers don’t care. Simply because the web evolves aggressively,[ ](https://developers.cloudflare.com/bots/?spm=a2ty%5Fo01.29997173.0.0.547c51712sJynJ&ref=blog.datahut.co)[anti-bot systems evolve even faster](https://developers.cloudflare.com/bots/?spm=a2ty%5Fo01.29997173.0.0.547c51712sJynJ&ref=blog.datahut.co), and scale amplifies every tiny weakness. To make this engaging (and honest), here is the story in the right order — starting from the pain companies feel first, then revealing why it happens. You can also checkout the list of[ ](https://www.blog.datahut.co/post/web-scraping-tools/)[open source web scraping tools](https://www.blog.datahut.co/post/web-scraping-tools/) : ## 1\. The Biggest Danger: Silent Failures (Where Companies Lose Real Money) Most companies don’t complain when scrapers crash.They complain when scrapers pretend to work. Silent failures return: - empty HTML - incomplete product data - soft 404 pages - CAPTCHA HTML masked as real pages - JavaScript-heavy websites returning unhydrated or partially-rendered DOM snapshots - sanitized versions of content - headless browsers returning incomplete or pre-hydration HTML The dashboard shows “green.”Your datasets look “valid.”But behind the scenes: - competitor price drops go unnoticed - stockouts aren’t detected - new variants don’t appear - discounts go untracked - attributes break silently [At scale](https://docs.scrapy.org/en/latest/topics/deploy.html?ref=blog.datahut.co) — when scraping millions of records per day — a 2% silent failure rate becomes a massive business lOpen Source. Silent failures are the [#1](https://www.blog.datahut.co/blog/hashtags/1) reason Open Source scraping “fails.”[ ](https://www.blog.datahut.co/post/competitive-price-intelligence-can-determine-profitability/)[You can see real examples of these failure modes in our guide on E‑commerce Pricing Intelligence](https://www.blog.datahut.co/post/competitive-price-intelligence-can-determine-profitability/), where even small data gaps lead to major pricing mistakes. For brands that rely on web scraping services to power pricing, assortment, and availability decisions, these silent failures don’t just break dashboards — they translate directly into missed revenue and margin erosion. ## 2\. Why These Failures Happen: The Anti-Bot Arms Race Anti-bot companies iterate faster than Open Source projects pOpen Sourceibly can.They constantly update: - [TLS/JA3 fingerprint detection](https://developer.mozilla.org/en-US/docs/Glossary/TLS?ref=blog.datahut.co) - Canvas/WebGL fingerprinting - Timing + behavioral scoring - IP reputation models - JavaScript challenge flows - Hidden trap endpoints These updates happen daily — sometimes hourly. Cloudflare highlights this in their own documentation on evolving[ bot challenges](https://developers.cloudflare.com/bots/?ref=blog.datahut.co) And here’s the part few mention: Anti-bot companies download Open Source scraping tools the moment they’re released.They study them. They fingerprint them. They train ML detectors on them.They block them. Most in-house teams only realize this after days of unexplained failures; seasoned web scraping companies know that the real game is staying just unpredictable enough that you don’t become an easy signature in somebody’s bot-detection model. Open source is public. Anti-bot teams reverse-engineer faster.With AI-assisted patching, defenses update in hours, not weeks. This alone gives Open Source scrapers a very short shelf life. ## 3\. Websites Change Even Faster — And Scale Magnifies Every Break Modern websites shift constantly. These JavaScript-heavy pages rely on dynamic hydration, client-side rendering, and API-driven blocks that break frequently and often silently . [ ](https://www.nngroup.com/articles/ecommerce-product-pages/?ref=blog.datahut.co)[Nielsen Norman Group’s UX research](https://www.nngroup.com/articles/ecommerce-product-pages/?ref=blog.datahut.co) shows how often e-commerce teams run layout experiments and product-page redesigns. These continuous shifts are the exact drift patterns we highlight in our Datahut article on why bad product data quietly destroys revenue. — something we break down extensively in our article on How[ ](https://www.blog.datahut.co/post/amazon-data-hygiene-webscraping/)[Retailers Lose Money to Bad Product Data](https://www.blog.datahut.co/post/amazon-data-hygiene-webscraping/). - template layouts - CSS classes - data blocks - variant cards - API endpoints - JS hydration flows For small prototypes, these are tiny issues.But at enterprise scale, each break becomes a disaster: - 1 selector break → 50,000+ failed pages - 1 layout change → 200,000+ unusable rows - 1 JS tweak → entire datasets wiped Open Source scrapers aren’t weak.They simply cannot adapt to fast, continuous change. ## 4\. The Concurrency Collapse (The Hidden Breaker at Scale) This is where most internal systems fall apart. Teams try to solve delays by “just increasing parallelism.” But Open Source scraping frameworks aren’t built to handle: - thousands of parallel sessions - context isolation - distributed job orchestration - global rate limits - dynamic region rotation The result is[ ](https://link.springer.com/chapter/10.1007/978-3-030-29400-7%5F26?ref=blog.datahut.co)[concurrency collapse](https://link.springer.com/chapter/10.1007/978-3-030-29400-7%5F26?ref=blog.datahut.co): - queues stall - threads hang - browser contexts leak memory - sessions freeze - proxies burn out in batches - backpressure cascades acrOpen Source the pipeline On dashboards it looks like “slow scrapers.”In reality, the system is choking under its own concurrency load. This is one of the biggest reasons internal pipelines degrade over time. ## 5\. Data Freshness Degradation (Your Pipeline Slowly Falls Behind) Even if nothing “breaks,” Open Source scrapers degrade gradually: - captcha loops slow batches - retries pile up - backlogs delay next cycles - failed crawls add recrawl load - retry storms overwhelm proxies Your once “hourly” pipeline becomes: - 3 hours behind - then 6 - then 12 - eventually 24+ hours delayed In pricing intelligence, retail, travel, or real-time availability monitoring —a delayed pipeline is a broken pipeline. [Freshness degradation is one of the biggest hidden costs of Open Source-based setups.](https://thenewstack.io/the-hidden-cost-of-open-source-waste/?ref=blog.datahut.co) ## 6\. Technical Limitations That Hit Hard at Scale These issues are invisible at 500 pages.They become catastrophic at 5 million. ### A. Static selectors + rigid flows One popup, cookie banner change, or DOM shift → mass failure. ### B. Browser-based tools degrade over long runs Playwright/Puppeteer/Selenium suffer from: - memory leaks - zombie processes - context bloat - slowdown drift - massive RAM usage ### C. Bypass techniques lag behind anti-bot vendors (Cloudflare, Imperva, Akamai, PerimeterX) Stealth libraries have a lifespan of days to weeks.[ Akamai’s bot management insights](https://www.akamai.com/products/bot-manager?ref=blog.datahut.co) explain how modern anti‑automation systems fingerprint browser behavior at a granular level, making any static spoofing method. ## 7\. Organizational Reality: Scraping Is a Full-Time Engineering Discipline This is where most teams underestimate complexity. Scraping at scale requires: - continuous selector maintenance - centralized logic management - observability pipelines - drift detection - proxy pool management - concurrency tuning - multi-region failovers - infrastructure orchestration Internal engineers end up spending 60–70% of their time on: - fixing selectors - debugging page states - chasing layout issues - patching scripts - managing proxy burnouts — something Distil Networks (now Imperva Bot Management) has analyzed in detail in their reports on automated traffic patterns. - handling backlogs - babysitting browser sessions Burnout becomes inevitable.Velocity drops.Your roadmap slows down.And scraping becomes a black hole of engineering hours. This is why even strong tech teams eventually abandon internal scrapers for managed solutions. The most effective teams treat data extraction as a dedicated function. They either build an internal capability that thinks like a specialist web scraping company, or partner with managed web scraping services that live and breathe reliability, anti-bot evasion, and data quality. ## 8\. Open Source Maintainers Can’t Patch as Fast as the Web Breaks Maintainers are volunteers, students, weekend contributors. They cannot patch: - new anti-bot techniques - browser rendering changes - network fingerprint updates - protocol shifts - at enterprise speed. This isn’t criticism — it’s simply not their job. ## 9\. The Real Bottleneck: Millions of Requests Amplify Every Weakness ### Open Source tools are great for: - 10 sites - 50 categories - 100k pages/day But at millions of pages/day, tiny cracks become: - retry storms - proxy exhaustion - infrastructure overload - cascading timeouts - huge recrawl storms - delayed data Open source wasn’t designed for: - 24×7 uptime - multi-region scraping - enterprise logging - compliance audits - unpredictable anti-bot escalation Not a flaw.Just not the purpose. ## 10\. Open Source Is Amazing — It’s Just Being Used for the Wrong Job Open Source tools are perfect for many use cases, especially when combined with structured approaches like those outlined[ ](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch/)[in our Definitive Guide to Building Web Crawlers](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch/): - prototyping - proofs of concept - enrichment tasks - academic research - lightweight crawls - one-off data pulls - early-stage products The problem is not that open source is weak. The problem is when companies take a prototype and try to scale it to a multi-million-page production system. That’s where the shelf life ends. ## 11\. Extending the Shelf Life of Open Source Scrapers To extend the shelf life of open‑source scrapers, teams must go beyond scripts and adopt production‑grade engineering patterns. A robust setup includes: - monitoring (success rates, DOM drift, CAPTCHA events, anomaly spikes) - drift detection (schema changes, attribute movement, JS hydration differences) - indirection layers (centralized selector logic, one‑patch‑fix‑all architecture) - hardened infrastructure (proxy pools, geo-routing, autoscaling, retries) - concurrency control (dynamic throttling, region-aware rate limits) - failover scrapers (redundant flows, backup extractors, hybrid browser/API strategies) - ML-based soft‑404 detection (classifying fake pages, honeypots, and trap responses) When implemented together, these turn Open Source setups from fragile → operationally reliable. This is also the baseline you should expect from serious web scraping services: not just scripts that run, but an ecosystem of monitoring, drift detection, proxy intelligence, and failover strategies that keep data flowing even as the web fights back. ## Final Thought The short shelf life of Open Source scraping tools isn’t because they’re [weak.It](http://weak.it/?ref=blog.datahut.co)’s because the web — and anti-bot defenses — evolve faster than volunteer-maintained tools Open Source tools possibly can. Open source is a foundation.But at scale, you need an ecosystem — architecture, observability, resilience. ## Why Enterprise Teams Prefer Datahut If you’ve ever wondered why Fortune-500 teams, large retailers, financial platforms, and marketplaces trust [Datahut](https://www.datahut.co/solutions?ref=blog.datahut.co) over Open Source tools. Here’s the simple truth: Not all web scraping companies are built the same. Datahut operates as a deeply specialized, compliance-first partner rather than a generic vendor. This is why our customers treat us as critical infrastructure instead of a disposable tool. Our shelf life is longer because our technology never becomes predictable. Unlike Open Source tools or web scraping apis that are available to public, fingerprintable, and quickly patched against, Datahut’s scraping stack is: - fully private - continuously adaptive - region-sharded - anti-bot-aware - proxy-intelligent - shielded from public eyes - never exposed to customers This means anti-bot vendors cannot: - study our behavior - fingerprint our flows - model our traffic - patch against our techniques Our stealth remains effective far longer. That’s why enterprises rely on Datahut when accuracy, uptime, and scale directly impact revenue. If your brand depends on reliable, large-scale data extraction — talk to[ ](https://datahut.co/?ref=blog.datahut.co)[Datahut](https://datahut.co/?ref=blog.datahut.co). We’ll show you what stable, enterprise-grade scraping really looks like. ## FAQ ### Question 1: What are some good open source web scraping and crawling tools Answer: Some of the most widely used open source web scraping and crawling tools are: - Scrapy – A Python-based crawling framework that’s great for large, structured spiders and pipelines. - Playwright / Puppeteer – Headless browser automation tools that work well for JavaScript-heavy websites. - Selenium – A mature browser automation framework originally built for testing, often reused for scraping. - BeautifulSoup / lxml – Lightweight HTML/XML parsing libraries, usually combined with requests or httpx. These tools are excellent for prototypes, research, internal tooling, and low-to-medium scale crawls—especially when you have in-house engineering capacity. ### Question 2:Why should I choose open source tools for web scraping over paid alternatives? Open source tools are a good choice when: - You want full control over the code, infrastructure, and data flow. - You have engineers who enjoy building and maintaining scraping logic. - Your use case is limited in scope (fewer sites, lower volume, non–time-critical). - You’re running experiments, POCs, or academic projects where budget is tight and risk is low. They let you move fast early, learn how the target websites behave, and avoid vendor lock-in. The trade-off is that, as volume, complexity, and anti-bot pressure increase, you’ll need to invest heavily in monitoring, maintenance, and infrastructure to keep those open source scrapers healthy. ### Question 3: How do open source web scraping tools differ from commercial web scraping software? #### Open source tools are usually: - Do-it-yourself: you assemble the pieces (fetching, parsing, storage, monitoring). - Public and fingerprintable: anti-bot vendors can download, study, and detect common patterns. - Community-maintained: updates and fixes depend on volunteer time and priorities. #### Commercial / managed web scraping solutions typically offer: - End-to-end pipelines: collection, cleaning, normalization, delivery, and monitoring in one place. - Private, non-public stacks: harder for anti-bot systems to fingerprint and block. - SLAs, support, and compliance: uptime commitments, legal review, and dedicated teams. - Operational maturity: proxy management, drift detection, alerting, and failover already built in. I n short: open source is great for building; commercial platforms are built for running at scale without constant firefighting. ### Question 4: How can I avoid getting blocked or detected when using open source web scraping tools? There’s no magic switch to “never get blocked,” but you can reduce issues by: - Throttling and scheduling: slow down request rates, add jitter, and spread crawls over time instead of spiking traffic. - Using quality proxies: distribute traffic across regions and IPs instead of hammering from a single address. - Rotating headers and sessions: send realistic user agents, cookies, and session data rather than obvious default fingerprints. - Handling JavaScript-heavy pages carefully: use headless browsers when needed, and make sure you wait for content to render before scraping. - Monitoring for drift and failures: track error rates, HTML changes, CAPTCHA frequency, and soft 404s so you notice problems early. Even with all of this, open source stacks will still hit limits at large scale. At that point, many teams either build a dedicated in-house scraping platform or move to a managed web scraping service that’s designed to handle anti-bot defenses and constant website changes for them. ### Amazon Menstrual Cup Data Analysis Using Web Scraping URL: https://www.blog.datahut.co/post/scraping-amazon-s-menstrual-cup-data-using-playwright-and-curlcffi/ Last updated: 2026-09-07T09:43:47.000Z When thinking about [menstrual cups](https://www.treksandtrails.org/blog/the-menstrual-cup-changed-my-life-and-its-bound-to-change-yours-too/?ref=blog.datahut.co), they are more than just a reusable alternative to pads or tampons—they represent convenience, sustainability, and personal health. On Amazon, one of the largest online marketplaces in the world, a wide range of menstrual cups is available, catering to different sizes, materials, and preferences. By scraping menstrual cup data from [Amazon’s website](https://www.amazon.in/?ref=blog.datahut.co), including product titles, brands, prices, and reviews, it is possible to uncover insights about which products are most popular, how buyers respond to different designs, and which features matter most to users. Similar to exploring a curated store, this data shows patterns in consumer choices and helps identify trends in the growing market for menstrual hygiene products. In the sections ahead, the blog will walk through how this [Amazon data was collected](https://www.blog.datahut.co/post/the-secret-weapon-of-successful-amazon-sellers/), what key details were extracted, and what the numbers reveal about menstrual cup usage and preferences across different buyers, giving a clear picture of this essential segment of women’s wellness products. ## From URLs to Insights: The Two-Stage Scraping Process for Amazon Data Before diving into analysis, the first step in working with Amazon’s menstrual cup dataset was to build a system that could reliably collect product links and then extract meaningful information from each page. Instead of treating scraping as a single step, I approached it as a small two-stage journey—first gathering every valid URL from the search results, and then visiting those pages to extract clean, structured data. This helped create a steady, well-organized workflow where each part supported the next, much like laying a strong foundation before building the rest of the structure. 1. URL Collection From Amazon’s Menstrual Cup Section When working with [Amazon’s menstrual cup listings](https://www.amazon.in/s?k=menstrual+cup&ref=blog.datahut.co) , the first challenge is not collecting product details—it’s simply finding all the product URLs hidden across the pages. Although Amazon doesn’t use endless scrolling in the same way some sites do, it still loads content dynamically, displays pagination that changes based on filters, and sometimes even reorders results during repeated visits. To handle this variability, an automated URL-scraping system using Python tools like Playwright, asyncio, and Playwright Stealth was built . These tools work together to open the [Amazon results page](https://www.blog.datahut.co/post/is-it-legal-to-scrape-amazon-unethical-uses-of-amazon-web-scraping/), behave like a real visitor, and move through the listings one page at a time without raising suspicion. Since Playwright allows browser-level interactions, the script could scroll naturally, wait for elements to load, and extract the product link from each listing in a clean, structured manner. Because real-world websites often introduce delays or network hiccups, the script was designed with gentle pauses using the random module to imitate human browsing. Meanwhile, logging quietly kept track of each step so I could revisit the process later if needed. Every link collected was stored inside a SQLite database, rather than a simple text file, making retrieval more reliable and reducing duplication. If a product appeared under multiple URLs—something that happens frequently on Amazon—the script checked the database first before saving anything new. By the end of this phase, I had a clear, well-organized set of product URLs ready for deeper analysis, almost like collecting the raw ingredients before beginning the actual recipe. 1. Structured Data Extraction From Each Product Page With a complete list of URLs ready, the next stage was to visit each product link individually and extract meaningful information from the page. This part required a slightly different approach because product pages on Amazon often load parts of their content asynchronously, and some sections appear only after certain scripts finish running. To ensure nothing important was missed, I used a combination of Playwright and curl\_cffi, allowing the scraper to switch between a full browser and a lightweight request method depending on the situation. Playwright handled pages that needed dynamic loading, while curl\_cffi provided speed on simpler pages—making the workflow flexible and efficient. Once the HTML was fetched, BeautifulSoup took charge of parsing the page, reading through titles, brand names, prices, ratings, reviews, descriptions, and the small but meaningful aspects such as quality, comfort, and ease of use. To keep the process clean, each product’s extracted data was stored as structured JSON, and also saved into SQLite so that no information was lost even if the program stopped unexpectedly. Tools like urllib.parse and urlparse helped decode and clean Amazon’s often complex URLs, while modules like re and json ensured that the scraped text could be shaped into clean, readable data ready for analysis. This second phase felt like opening each product page one by one and carefully noting the details in a notebook, but done with the precision and consistency of automation. By the end, every menstrual cup product—filtered, cleaned, and structured—was ready for insights, comparisons, and further storytelling through data. Together, these steps form a complete cycle of intelligent scraping: discovering product URLs first, then harvesting the information they hold in a way that is reliable, human-like, and well-organized. 1. Turning Messy Amazon Data into Meaningful Information with OpenRefine When you first scrape data from a website like Amazon India , like the menstrual cup search results , it feels rewarding—you’ve just gathered a large collection of product details, complete with titles, brands, prices, ratings, and customer impressions. But just like most real-world datasets, the raw output is rarely clean. It often arrives mixed with unrelated products, repeated entries, symbols that computers can’t interpret easily, and columns that need more structure before analysis becomes meaningful. A good way to think about this is to imagine returning from a grocery store with a big basket of fruits. They look bright and colorful, but before you can actually eat them, you still need to wash, sort, and cut them. [Data cleaning in OpenRefine](https://southampton-rsg.github.io/openrefine-data-cleaning/aio.html?ref=blog.datahut.co) works exactly like that—slowly transforming raw, uneven information into something neat, organized, and ready to work with. OpenRefine’s interface makes this process almost conversational, helping you clean step by step. While cleaning the Amazon menstrual cup dataset, the first thing I noticed was that not every product in the results was actually a menstrual cup. Some listings were only for washers, some for sterilizers, and some even for pouches. These items may be related, but they would introduce noise into the analysis, so I filtered out those URLs right at the start. Another issue was duplication—the same product often appeared under different URLs, which is common on large marketplaces. [OpenRefine’s](https://openrefine.org/?ref=blog.datahut.co) clustering options helped identify and remove these duplicates so that only unique products remained. Price cleaning was another important step. Amazon usually displays prices with the Indian Rupee symbol (₹) and commas, which look fine to us but make calculations complicated for machines. By removing symbols and formatting the values into plain numbers, the price column became clean, consistent, and ready for comparison. A similar approach was used for the brand data. Many product titles included brand names mixed with extra words, so I used OpenRefine to separate these into a structured brand column with clear, meaningful values. One of the more interesting parts of the cleaning process involved the aspects column, which contains phrases like “quality: positive” or “ease of use: negative.” In their raw form, these look like scattered text, but by flattening them into a tidy structure—where each aspect becomes a clear field—it becomes much easier to analyze customer sentiment later. This step almost felt like unfolding a crumpled sheet of paper: the information was always there, just waiting to be organized properly. ## Essential Tools Behind the Scraper: The Python Libraries That Make Everything Work When you look at a working scraper from the outside, it often seems simple—just run a script and data appears. But behind that small command lies a group of tools quietly doing their part, almost like a backstage team that keeps a theater production running smoothly. In this project, the scraper relied on a set of Python libraries that worked together in harmony, each one handling a specific responsibility. At the heart of the workflow was [asyncio](https://realpython.com/async-io-python/?ref=blog.datahut.co). It allowed the program to perform multiple tasks at once, so while one Amazon product page was loading, another could already be parsed. This created a natural sense of flow, especially when working with a large list of URLs. Pairing with asyncio was [Playwright](https://playwright.dev/?ref=blog.datahut.co) , which acted as the browser window of our script. It opened pages, scrolled through dynamic content, and captured HTML—even on sections that would normally load only when a user interacts with the site. To make Playwright behave more like a human visitor, Playwright Stealth blended in small browser adjustments, reducing the chances of triggering Amazon’s anti-bot systems. Once the HTML was fetched, BeautifulSoup stepped in as the gentle cleaner, reading through the raw markup and helping extract only the pieces we needed—titles, brands, prices, and aspects. But sometimes even Playwright can be too heavy for smaller requests, which is where [curl\_cffi](https://pypi.org/project/curl-cffi/?ref=blog.datahut.co) offered speed, sending fast, lightweight HTTP calls whenever possible. Behind the scenes, tools like urllib.parse, urlparse, and unquote quietly ensured that every URL was decoded and formatted correctly. Data also needs a safe place to live, and sqlite3 provided exactly that—a simple, file-based database that works without any additional setup. It became our storage shelf where URLs and product records were saved in an organized, query-friendly manner. Throughout the process, logging acted like a diary, noting each success and failure, while random introduced small, natural-looking pauses to make the scraper feel less robotic. Even modules like re, Path, and contextmanager played their part, helping with text cleaning, file handling, and resource management in a clean and predictable way. ## STEP 1: Gathering All Product Links from Amazon’s Menstrual Cup Category ### Importing Libraries ``` # IMPORTS import asyncio import sqlite3 import random import logging import re import urllib.parse from pathlib import Path from playwright.async_api import async_playwright from playwright_stealth import stealth_async ``` When you’re just starting out with web scraping, the first steps can feel a little overwhelming. There are tools, libraries, environments, and a whole lot of new words. But once you slow it down and look at each part one piece at a time, the picture becomes much clearer. At first glance, this block might look like a list of random ingredients. But just like a recipe, each one has a purpose. And the magic happens when they work together. We begin by bringing in modules like asyncio, which helps Python handle tasks without waiting around, almost like letting your program multitask gracefully, while [sqlite3](https://www.sqlite.org/docs.html?ref=blog.datahut.co) gives you a built-in way to store scraped data locally, similar to keeping your notes organized in a single notebook rather than scattered files. You’ll also notice we import random and logging, which may sound simple, but they quietly play big roles: randomness helps you rotate actions like user agents to appear more natural online, and [logging](https://docs.python.org/3/library/logging.html?ref=blog.datahut.co) keeps a transparent record of what your script is doing behind the scenes—much like a diary you can check later when something breaks. The re module steps in to help with pattern matching when you need to clean or extract text, and urllib.parse helps safely manage and format URLs so your scraper doesn’t stumble on messy links. Then comes Path from pathlib, which offers an easy way to handle file paths without worrying about operating-system differences, making your project feel tidier and more predictable. Finally, we have imports from Playwright—async\_playwright and the [stealth](https://pypi.org/project/playwright-stealth/?ref=blog.datahut.co) module—both of which allow the script to open and browse websites the way a human would, loading pages smoothly while trying to avoid detection from strict websites . Together, these imports set the foundation for a scraper that is simple to understand, friendly for beginners, and strong enough to grow into a more advanced project as your skills improve. ### Setting Up the Scraper Configuration ``` # CONFIGURATION START_URL = "https://www.amazon.in/s?k=menstrual+cup" """str: The main Amazon search page URL from where the scraper starts collecting product links.""" DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon_cup/DATA/amazon_menstrual_cups.db" """str: File path of the SQLite database where all scraped product URLs and product details will be stored.""" USER_AGENT_FILE = "/home/anusha/Desktop/DATAHUT/chewy-and-petco/user_agents.txt" """str: Path to the text file containing a list of user agents. The scraper randomly picks one user agent to avoid detection.""" LOG_FILE = "/home/anusha/Desktop/DATAHUT/Amazon_cup/LOG/scraper_log.log" """str: File path for saving all log messages (errors, success messages, warnings). Helpful for debugging and tracking scraper activity.""" ``` Setting up the configuration for a web-scraping script often feels like arranging the starting points of a small adventure, and the variables in this section give the scraper clear instructions on where to begin, where to save information, and how to disguise itself while exploring pages. The [START\_URL](https://developer.mozilla.org/en-US/docs/Learn%5Fweb%5Fdevelopment/Howto/Web%5Fmechanics/What%5Fis%5Fa%5FURL?ref=blog.datahut.co) points directly to an [Amazon](https://www.amazon.in/?ref=blog.datahut.co) search page for menstrual cups, acting as the entry gate from which product links are discovered. The DB\_PATH then guides the script toward a specific SQLite database stored on the local system, allowing all collected product details to be organized safely in one place, much like labeling a box so it can be found later . The USER\_AGENT\_FILE plays an equally important role because it stores different browser identities, and the scraper quietly picks one at random each time to appear more natural while visiting Amazon—similar to how different people have different browsing patterns. Meanwhile, the [LOG\_FILE](https://docs.python.org/3/howto/logging.html?ref=blog.datahut.co) creates a dedicated home for all log messages, collecting errors, warnings, and activity notes so the script’s progress stays transparent and easy to troubleshoot . Together, these simple configuration lines act like a roadmap, storage room, disguise kit, and diary for the scraper, giving the entire process a sense of order and continuity before the actual crawling work begins. ### Adding Logging to Monitor the Scraper’s Performance ``` # LOGGING logging.basicConfig( filename=LOG_FILE, level=logging.INFO, format="%(asctime)s | %(levelname)s | %(message)s", ) logger = logging.getLogger(__name__) """Sets up the logging system for the scraper.This helps track what the script is doing, including errors, warnings, and progress.All logs will be saved to the file defined in LOG_FILE.Logger object used throughout the script to record messages in the log file.""" ``` Setting up logging in a scraping script often feels like placing a quiet observer in the background, someone who carefully notes what happens at every stage, and the configuration here does exactly that by directing [Python’s logging](https://realpython.com/python-logging/?ref=blog.datahut.co) system to write all activity into the file defined by LOG\_FILE. The logging.basicConfig call shapes this observer’s behavior by telling it where to save messages, what level of detail to record, and how each line should look once written; beginners often find it helpful to think of this as keeping a diary for the script, one that records important moments like errors, warnings, and general progress updates in a neatly formatted style. The line logger = logging.getLogger(\_\_name\_\_) simply retrieves a logger that the rest of the script can use, allowing every function to add notes to the same diary without needing any extra setup. Once the logger is in place, the script gains a sense of transparency, making it easier to revisit what happened during long scraping sessions and offering clear clues when something behaves unexpectedly, much like a trail of footprints left behind after a long walk through a forest of web pages. ### Setting Up the SQLite Database for URL Storage ``` # DB SETUP def init_db(): """ Create (or connect to) the SQLite database and set up the required table. """ conn = sqlite3.connect(DB_PATH) cur = conn.cursor() cur.execute(""" CREATE TABLE IF NOT EXISTS product_urls( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE ); """) conn.commit() return conn ``` Creating the database setup function in a scraping project often feels like preparing a dedicated storage shelf before collecting anything, and the init\_db function serves that purpose by opening a connection to the SQLite file defined in DB\_PATH and ensuring the necessary table is ready to hold product URLs. When the function calls sqlite3.connect(DB\_PATH), it either opens the existing database or quietly creates a new one if it is not already present, and this simple behavior is one reason beginners often find [SQLite](https://docs.python.org/3/library/sqlite3.html?ref=blog.datahut.co) approachable. Once the connection is established, the script prepares a cursor and executes a CREATE TABLE IF NOT EXISTS command, which guarantees that a table named product\_urls is available without worrying about errors if it already exists. This table holds two fields—an automatically generated ID and a unique URL—forming a clean structure for storing links gathered from Amazon as the scraper moves through different pages. After committing the change, the function returns the connection so the rest of the script can continue using it, allowing later parts of the code to insert, update, or read data without setting up the database again. In many ways, this small function lays the foundation for the entire scraping process, similar to setting up a notebook before beginning research, ensuring that everything collected has a proper and organized place to stay. ### Preparing Clean Amazon URLs Before Scraping ``` # CLEAN AMAZON URL def clean_url(raw): """ Clean and normalize Amazon product URLs. """ if not raw: return raw # Sponsored redirect: # /sspa/click?...&url=%2Fbrand-product%2Fdp%2FASIN%2Fref... if "sspa/click" in raw and "url=" in raw: parsed = urllib.parse.urlparse(raw) qs = urllib.parse.parse_qs(parsed.query) if "url" in qs: true_url = qs["url"][0] decoded = urllib.parse.unquote(true_url) return decoded # Normal product link → keep EXACTLY return raw ``` Cleaning Amazon URLs can feel a bit like removing extra stickers from a package before storing it, and the clean\_url function is designed to do exactly that by taking the raw link collected from the page and checking whether it is a straightforward product link or a sponsored redirect that Amazon often uses for advertisements. When the function receives a URL, it first makes sure the value is not empty and then looks for signs of a redirect, especially the pattern often seen in links beginning with /sspa/click, which usually contains the “real” product link hidden inside the url= parameter. Using tools from [Python’s urllib.parse module](https://docs.python.org/3/library/urllib.parse.html?ref=blog.datahut.co)[,](https://docs.python.org/3/library/urllib.parse.html?ref=blog.datahut.co) the function breaks the redirected link into its parts, extracts the original product URL, and decodes it so it becomes readable again, much like opening a folded note to reveal the actual message inside. If the URL doesn’t contain any sponsored indicators, it is returned exactly as it is, since many Amazon product pages follow a predictable structure that works without any adjustments. By the time the function finishes, each link becomes clean and trustworthy, making the scraping process smoother and the stored URLs easier to handle later, creating a natural flow in the pipeline where messy inputs are quietly converted into neat and usable product links. ### Saving Product URLs into the Database ``` # SAVE def save_url(conn, url): """ Save a product URL into the database. """ try: cur = conn.cursor() cur.execute("INSERT OR IGNORE INTO product_urls(url) VALUES (?)", (url,)) conn.commit() if cur.rowcount > 0: logger.info(f"Saved URL: {url}") else: logger.info(f"Duplicate skipped: {url}") except Exception as e: logger.exception(f"Error saving URL: {url} | {e}") ``` Saving each product link into a database may seem like a small step in the scraping process, yet the save\_url function shows how important it is to store information carefully and avoid unnecessary duplicates while collecting data from a large site like Amazon. When this function receives a database connection and a URL, it prepares a simple SQL command—[INSERT OR IGNORE](https://www.sqlite.org/lang%5Finsert.html?ref=blog.datahut.co)—which gently attempts to add the link into the product\_urls table without causing an error if the same link was already stored earlier, much like placing an item on a shelf only if it is not already there. After running the command and committing the change, the function checks whether a new row was actually added, and based on that, a message is written through the logger to indicate whether the link was successfully saved or skipped because it already existed; this makes it easy to trace how the scraper progresses over time. In case something unexpected happens—perhaps because of a malformed link or a temporary database issue—the except block catches the error and records a detailed message for debugging, helping identify and solve problems without interrupting the entire scraping workflow. Through this small function, the script gains a reliable and organized way to store every valuable URL it discovers, creating a continuous sense of structure as the scraper moves from collecting links to processing them later. ### Loading User Agents to Keep the Scraper Undetected ``` # USER AGENT LOADER def load_user_agents(): with open(USER_AGENT_FILE, "r") as f: agents = [ua.strip() for ua in f.readlines() if ua.strip()] return agents """ Load user agents from the text file. """ ``` Loading user agents might sound like a technical detail, but the load\_user\_agents function gives the scraper an important ability: the chance to appear like different browsers each time it visits Amazon, reducing the chances of getting blocked and making the entire scraping process run more smoothly. The function simply opens the file specified by [USER\_AGENT\_FILE](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/User-Agent?ref=blog.datahut.co), reads each line, removes any extra spaces, and returns a clean list of user agent strings; this list acts almost like a collection of digital disguises, allowing the scraper to rotate between identities the way a person might try different doorways to avoid drawing attention. By returning a neatly prepared list, the function creates a smooth hand-off to the part of the script responsible for selecting a random user agent during requests, forming a natural connection between this early setup step and the later stages where pages are actually fetched. This simple routine, though small in appearance, becomes an essential part of the scraper’s overall strategy, giving it the flexibility to blend in and continue gathering data without interruptions. ### How the Scraper Collects Product URLs from Amazon Search Results ``` # SCRAPER async def scrape_page(page, conn): """ Scrape all product URLs from the current Amazon search results page. """ logger.info("Scraping current page...") product_links = page.locator('a.a-link-normal.s-line-clamp-3.s-link-style.a-text-normal') count = await product_links.count() logger.info(f"Found {count} product links") for i in range(count): try: href = await product_links.nth(i).get_attribute("href") if not href: continue cleaned = clean_url(href) # Ensure full URL full_url = urllib.parse.urljoin("https://www.amazon.in", cleaned) save_url(conn, full_url) await asyncio.sleep(random.uniform(0.5, 0.8)) except Exception as e: logger.exception(f"Error processing link {i}: {e}") ``` The scrape\_page function works like a careful collector moving through an Amazon search results page, gathering product links one by one and preparing each of them so they can be stored properly for later use, and its flow becomes easier to understand once the steps are seen as part of a single smooth process rather than separate technical tasks. When the scraper arrives on a page, it starts by locating all elements that match Amazon’s product link pattern, using a CSS selector/html content that points specifically to the titles shown in the search results; this selector may look complex at first glance, but it simply tells Playwright which pieces of the page represent actual product entries. After counting how many such links exist, the function loops through each one, extracting the href attribute, which is the raw link Amazon provides. Some of these links may be redirected or cluttered with tracking information, so the clean\_url function is called to tidy them up. Once cleaned, the link is turned into a full Amazon URL using urllib.parse.urljoin, ensuring it always starts with the proper domain instead of being left in a relative form. The function then hands this prepared link to save\_url, which stores it safely in the database. A short, random pause is added between each processed link to mimic natural browsing speed, which reduces suspicion when scraping large websites. If anything unexpected happens, the error is logged through the existing logging setup, allowing issues to be understood later without breaking the overall scraping session. Each step here flows into the next, creating a quiet rhythm where the script observes the page, gathers links, cleans them, stores them, and moves on, much like someone working steadily through a list without losing focus. ### Handling Pagination Across Amazon Search Results ``` # PAGINATION async def pagination_loop(page, conn): """ Loop through all Amazon search result pages and scrape each one. """ page_number = 1 while True: logger.info(f" Processing page {page_number}") await scrape_page(page, conn) next_btn = page.locator("a.s-pagination-next") if await next_btn.count() == 0: logger.info("❌ No next button found — stopping pagination") break disabled = await next_btn.get_attribute("aria-disabled") if disabled == "true": logger.info("❌ Next button disabled — scraping finished") break logger.info("➡ Clicking next page...") await next_btn.click() await asyncio.sleep(random.uniform(1, 2)) await page.wait_for_load_state("load") page_number += 1 ``` The [pagination\_loop](https://www.smashingmagazine.com/2016/03/pagination-infinite-scrolling-load-more-buttons/?ref=blog.datahut.co) function guides the scraper through Amazon’s multi-page search results in a steady, predictable manner, almost like turning the pages of a long catalog one by one, making sure nothing is missed along the way. It begins on the first page and immediately calls scrape\_page, which collects all the product links from that section, and once that page has been fully processed, the function looks for Amazon’s familiar “Next” button to determine whether there is another set of results waiting to be explored. If the button is not present or marked as disabled, it becomes clear that the last page has been reached, and the loop stops naturally without forcing the script to continue. But if the button is active, the function clicks it, waits briefly for the new page to load—much like a user pausing while a site refreshes—and then proceeds to scrape the next page in the exact same way. Each cycle includes a small, intentional delay so the scraper behaves more like a real person browsing through search results, which reduces the chances of triggering Amazon’s automated checks. Over time, this loop builds a quiet rhythm: scrape, check for the next page, move forward, and repeat until the entire trail of results has been followed from start to finish, forming a continuous flow that keeps the scraping process both organized and reliable. ### Controlling the Scraping Workflow with the Main Function ``` # MAIN async def main(): """ Main entry point of the Amazon scraper. """ conn = init_db() user_agents = load_user_agents() logger.info(" Scraper started") async with async_playwright() as p: browser = await p.chromium.launch( headless=False, args=["--disable-blink-features=AutomationControlled"] ) ua = random.choice(user_agents) context = await browser.new_context( user_agent=ua, viewport={"width": 1280, "height": 800} ) logger.info(f"Using User-Agent: {ua}") page = await context.new_page() await stealth_async(page) logger.info(f"Opening: {START_URL}") await page.goto(START_URL, wait_until="load", timeout=60000) await asyncio.sleep(random.uniform(2, 4)) await pagination_loop(page, conn) await browser.close() conn.close() logger.info(" Scraping completed!") ``` The main function acts as the central controller of the entire scraping process, guiding each step in a smooth sequence so everything works together without confusion, much like how a well-organized workflow begins with preparation, moves through the task itself, and ends by cleaning up properly. It first sets up the database that will store all collected URLs and then reads the list of user agents from the local file, which helps the scraper behave more like a regular browser. Once the preparation is complete, the function launches Playwright and assigns a random user agent so the scraper blends in more naturally while visiting the Amazon search results page defined by the START\_URL, and a short pause gives the browser time to load fully, similar to how a person waits for a page to settle before scrolling. After that, the scraper enters the pagination loop, which patiently moves through the full set of search result pages, gathering URLs from each one through the earlier functions; this creates a connected chain of steps rather than isolated pieces of work. When the loop finishes, the browser and database connection are both closed to avoid leaving loose ends, and a final log message marks the completion of the task. By moving through each phase in a steady and predictable rhythm—setup, browsing, scraping, and closure—the main function keeps the workflow simple and approachable, even for someone just beginning to understand how automated scraping scripts operate. ### Entry Point ``` # ENTRY POINT """ Runs the scraper when this file is executed directly. This triggers the main() function and starts the entire scraping process. """ if __name__ == "__main__": asyncio.run(main()) ``` When a Python script reaches the if name == "\_\_main\_\_": line, it simply means the file has been opened intentionally to run the program, and this tiny section quietly becomes the gateway that starts everything; calling [asyncio.run](https://docs.python.org/3/library/asyncio.html?ref=blog.datahut.co)(main()) is like turning the key in an engine, allowing the main scraper function to take over, prepare the browser, and begin collecting data step by step. This entry point keeps the project organized by making sure the scraping workflow runs only when the file itself is executed, not when it is imported somewhere else, similar to how a door opens only when the right key is used. This small entry block may look simple, but it plays an important role in keeping the scraper predictable, readable, and easy to extend as the project grows. ## Step 2: From Links to Comprehensive Product Information ### Importing Libraries ``` # IMPORTS import asyncio import random import sqlite3 import json import logging import re from contextlib import contextmanager from typing import Optional, List, Tuple, Dict from bs4 import BeautifulSoup from urllib.parse import urlparse, unquote from playwright.async_api import async_playwright, Page from playwright_stealth import stealth_async from curl_cffi import requests ``` In many scraping projects, different libraries handle different parts of the job, and seeing names like [curl\_cffi](https://curl-cffi.readthedocs.io/en/latest/?ref=blog.datahut.co) for the first time can feel a little unfamiliar, but it helps to think of it as a faster, more flexible version of the usual request tools, designed to mimic real browser traffic more closely so that websites treat the scraper like a normal visitor. Alongside it, modules such as urllib.parse and BeautifulSoup quietly take care of tasks like cleaning messy URLs and interpreting raw HTML, while type-hints like Optional or Dict bring clarity to how data flows through the script. The moment these imports appear at the top of a file, the code silently prepares itself with all the tools needed for the steps ahead, similar to how a workspace is set up before starting a task. Even though these imports may look like a simple list, they lay the foundation for the entire scraper, allowing the later functions to focus on the actual logic without repeatedly rewriting these essential features. ### Defining Paths, User Agents, and Log Settings ``` # CONFIG / CONSTANTS """Holds all file paths and important settings used by the scraper. These constants make the script easier to configure and reuse. """ DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon_cup/DATA/amazon_menstrual_cups.db" """str: Path to the SQLite database where scraped product URLs will be stored.""" USER_AGENT_FILE = "/home/anusha/Desktop/DATAHUT/chewy-and-petco/user_agents.txt" """str: Path to a text file containing multiple user-agent strings. The scraper picks one randomly to reduce blocking by Amazon.""" LOG_FILE = "/home/anusha/Desktop/DATAHUT/Amazon_cup/LOG/amazon_scraper_merged.log" """str: Path to the log file where the scraper will record all errors and progress.""" ``` Setting up a scraper often begins with gathering a few core details in one place, and that is what these configuration constants aim to do by keeping paths and settings neatly organized so the rest of the script stays clean and easy to read. The database path simply tells the program where to store collected product information, much like pointing someone to the correct drawer before filing documents. The user-agent file plays an equally helpful role by holding different browser identities that the scraper can rotate through, a small trick that makes requests appear more natural. The log file path takes care of another practical need by giving the scraper a place to record errors and progress, making it easier to revisit what happened during long runs. Although these lines might seem simple at first glance, they quietly set the foundation for the entire workflow, allowing the rest of the code to focus on the actual scraping rather than repeatedly hunting for paths or settings. ### How Simple Request Headers Help a Scraper Communicate Better ``` # HEADERS """Default HTTP headers used when sending requests. These help make the scraper look more like a normal browser request. """ HEADERS = { "content-type": "text/html", "Connection": "keep-alive", "cache-control": "no-transform", } ``` When a scraper sends a request to a website, the server often expects it to behave like a typical browser, and that’s where these default headers come in, acting like a small introduction that tells the site what kind of content is being requested and how the connection should be handled. The content-type header simply explains the format being asked for, while Connection: keep-alive helps maintain a stable link so the scraper doesn’t reconnect repeatedly, which can slow things down. The cache-control setting adds another layer of clarity by asking for fresh content rather than something stored from a previous visit. Even though these lines may look tiny, they quietly shape how the scraper communicates with the server, making the interaction smoother and more reliable as the rest of the script runs. ### Understanding the Core Settings That Guide the Scraper’s Behavior ``` # RETRIES & SCRAPER CONSTANTS """General settings that control scraper behavior. """ RETRIES = 3 PLAYWRIGHT_TIMEOUT = 120_000 # ms MIN_HTML_LEN = 5_000 SLEEP_RANGE = (3.0, 4.0) TABLE_NAME = "product_cffi_3" URLS_TABLE = "product_urls" ``` A scraper often needs a few guiding rules to handle the unpredictable nature of websites, and these constants help set that foundation by defining how many times a request should be attempted again when something goes wrong, how long Playwright should wait for a page to load, and even the minimum amount of HTML needed to consider a response useful. The retry count keeps the script from giving up too quickly when a site responds slowly, while the timeout prevents it from waiting forever on a frozen page. A small random pause, controlled by the sleep range, adds a touch of natural behavior so the scraper doesn’t look too mechanical, and the table names simply point the script to where product details and collected URLs should be stored inside the database. ### Keeping Track of the Scraper’s Journey with Simple Logging ``` # LOGGING SETUP """Configures logging for the scraper. """ logging.basicConfig( filename=LOG_FILE, filemode="a", level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s" ) logging.info("🚀 Merged Scraper Started") ``` A scraper benefits greatly from a clear record of what happens during its run, and this logging setup creates that trail by writing every important message into a dedicated log file without deleting previous entries. Each line is stored with a timestamp and a label that describes the type of message, making it easier to understand when something succeeded or when an error needs attention. This simple structure turns the log file into a quiet companion that keeps track of each step, helping beginners trace issues without feeling overwhelmed. ### Why a Scraper Needs User Agents and How This Utility Helps ``` # UTILITIES def load_user_agents(path: str = USER_AGENT_FILE) -> List[str]: """ Load user-agent strings from a text file. """ try: with open(path, "r") as f: ualist = [ua.strip() for ua in f if ua.strip()] if not ualist: raise ValueError("User agent file is empty") return ualist except Exception as e: logging.warning(f"Could not load user agents ({e}), using fallback UA.") return [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 " "(KHTML, like Gecko) Chrome/120.0 Safari/537.36" ] USER_AGENTS = load_user_agents() """List[str]: Loaded user agents used for rotating requests and avoiding detection.""" ``` This small utility focuses on gathering user-agent strings from a simple text file, almost like collecting different “browser masks” that help a scraper blend in while requesting pages online. The function reads each line carefully, ignores empty entries, and returns a clean list that can be rotated later during scraping. If something goes wrong—maybe the file is missing or contains no usable data—the code gently switches to a safe fallback user agent and records the issue through logging, keeping the flow of the script steady instead of stopping unexpectedly. ### Cleaning and Standardizing Amazon Product URLs ``` # URL CLEANING def clean_amazon_url(raw_url: str, asin: Optional[str]) -> str: """ Generate clean canonical URL: https://amazon.in//dp/ASIN """ parsed = urlparse(raw_url) domain = f"{parsed.scheme}://{parsed.netloc}" # If ASIN not found → return raw if not asin: return raw_url parts = raw_url.split("/") slug = None # Try to extract for p in parts: if p and p not in ["dp", "gp", "product", asin] and len(p) > 3: slug = p break if slug: return f"{domain}/{slug}/dp/{asin}/" return f"{domain}/dp/{asin}/" ``` This URL-cleaning function helps transform long, cluttered Amazon links into simple, consistent versions by breaking the original address into parts and keeping only the essential pieces, such as the domain and the product’s ASIN. The logic looks through the URL for a readable slug—usually a hint of the product name—and if one is found, the function rebuilds the link in a tidy format that is easier to store or reuse later; if no slug is available, the code still creates a clean fallback using only the ASIN. This concepts such as parsing and path segments are explained in detail, making it easier to understand how this function turns a messy link into a predictable, standardized one without interrupting the natural flow of the scraping process. ### A Simple and Safe Way to Handle Database Connections ``` # DATABASE CONTEXT MANAGER @contextmanager def db_conn(path: str = DB_PATH): """ A simple context manager for opening and closing the database connection. """ conn = sqlite3.connect(path) try: yield conn conn.commit() finally: conn.close() ``` This small database helper creates a safe workspace for interacting with a SQLite file, allowing the code to open a connection, run queries, and close everything cleanly without extra effort. The idea is similar to borrowing a tool for a moment and returning it once the job is done, ensuring nothing is left half-open or forgotten. By placing database actions inside a with block, the function quietly handles tasks like committing changes or closing the connection, even if an unexpected error appears along the way. ### Setting Up the SQLite Database for Storing Product Data ``` # DATABASE: init + helpers def init_db(): """ Initialize the SQLite database used for scraping. """ with db_conn() as conn: cur = conn.cursor() # product_url cur.execute(f""" CREATE TABLE IF NOT EXISTS {URLS_TABLE}( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE ) """) # ensure scraped column exists cur.execute(f"PRAGMA table_info({URLS_TABLE})") existing = {r[1] for r in cur.fetchall()} if "scraped" not in existing: try: cur.execute(f"ALTER TABLE {URLS_TABLE} ADD COLUMN scraped INTEGER DEFAULT 0") logging.info("Added 'scraped' column to product_url") except Exception as e: logging.warning(f"Could not add scraped column: {e}") # combined product table - ONLY KEEPING SPECIFIED FIELDS cur.execute(f""" CREATE TABLE IF NOT EXISTS {TABLE_NAME} ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, title TEXT, brand TEXT, price TEXT, mrp TEXT ) """) logging.info(f"DB initialized, tables: {URLS_TABLE}, {TABLE_NAME}") ``` Setting up the database begins with a small helper that quietly creates the tables needed for storing product links and the final scraped details, making sure everything is ready before the main scraper starts collecting information. The function opens a connection through the context manager defined earlier, creates the URL table if it is missing, and adds a “scraped” column when the database does not already have it, which helps the script remember what has been processed. It then prepares another table to hold the cleaned product data—only essential fields like the title, brand, and pricing—so the final dataset stays tidy and easy to handle. Each action is logged for clarity, giving a clear trail in case something needs to be checked later. ### Fetching Only the URLs That Are Still Waiting to Be Scraped ``` # DATABASE: fetchers & savers def fetch_pending_urls(limit: Optional[int] = None) -> List[Tuple[int, str]]: """ Fetch product URLs that are not yet scraped. """ with db_conn() as conn: cur = conn.cursor() q = f"SELECT id, url FROM {URLS_TABLE} WHERE scraped = 0" if isinstance(limit, int) and limit > 0: q += f" LIMIT {limit}" cur.execute(q) return cur.fetchall() ``` A beginner often wonders how a scraper knows which links still need attention, and this small function offers a simple way to manage that flow by quietly checking the database and pulling out only the URLs that are yet to be processed, much like picking unread messages from an inbox. The idea is straightforward: the database keeps a table named product\_url, and inside it sits a column called scraped, which marks each link as either done or pending. Whenever the function runs, it opens a safe connection using db\_conn()—a helper often explained earlier in the project—and gently reads only the rows where the scraped value is 0, meaning they are untouched and ready for processing. If a limit is provided, such as when testing or handling a small batch, the query simply adds a cap; otherwise, it gathers everything remaining. The result is returned as a clean list of pairs containing each URL and its corresponding ID, making the next parts of the pipeline easier to handle. This kind of step forms the backbone of many scraping workflows, because managing progress reliably is just as important as fetching the data itself. ### How Product Details Are Stored Using a Smart Upsert Method ``` # SAVE PRODUCT DATA WITH UPSERT LOGIC def save_product(data: Dict[str, Optional[str]]): """ Save product information into the database with an UPSERT (insert-or-update) logic. """ url = data.get("url") if not url: logging.error("save_product called without url") return # ONLY KEEPING SPECIFIED FIELDS cols = ["title", "brand", "price", "mrp"] with db_conn() as conn: cur = conn.cursor() # ensure row exists cur.execute(f"INSERT OR IGNORE INTO {TABLE_NAME} (url) VALUES (?)", (url,)) # build update statement with COALESCE set_parts = [] params = [] for c in cols: set_parts.append(f"{c} = COALESCE(?, {c})") params.append(data.get(c)) params.append(url) sql = f"UPDATE {TABLE_NAME} SET {', '.join(set_parts)} WHERE url = ?" cur.execute(sql, params) logging.info(f"Saved/updated product: {url}") ``` Storing scraped product details becomes much easier when a function gently handles both new entries and updates, and this save\_product function does exactly that by treating the product’s URL as its identity and building the rest of the process around it. The moment the function receives a dictionary of product information, it first checks whether a URL is present, because that single value decides where the data belongs in the database; without it, the function simply logs an error and steps back. Once a URL is confirmed, the database connection created through db\_conn() helps ensure that a row exists for that URL by using an INSERT OR IGNORE statement, and this approach quietly avoids duplicate-row errors. After the row is secured, the function prepares an update that relies on SQLite’s COALESCE() feature, allowing each new piece of data—title, brand, price, or mrp—to be added only when it is not missing; if a value is None, the database keeps the older non-empty value instead of replacing it. This small detail becomes very helpful when some pages fail to load fully or when partial data is all that is available during a scrape. Each update runs smoothly without disturbing existing information, and the process fits naturally with earlier steps like fetching pending URLs, forming a clear workflow that moves from discovering product links to storing their details safely. ### Marking URLs as Completed in the Scraping Process ``` # MARK URL AS SCRAPED def mark_scraped(row_id: int): """ Mark a URL record as scraped. """ with db_conn() as conn: cur = conn.cursor() cur.execute(f"UPDATE {URLS_TABLE} SET scraped = 1 WHERE id = ?", (row_id,)) logging.info(f"Marked scraped → ID {row_id}") ``` Tracking progress becomes much easier when each URL clearly shows whether it has already been processed, and the mark\_scraped function handles this step in a simple, predictable way that even beginners can follow without confusion. The moment this function receives an ID, it opens a database connection through the same db\_conn() context manager that earlier parts of the system rely on, creating a smooth flow from fetching pending URLs to saving product details and finally marking them as completed. Inside the connection, a small SQL update statement sets the scraped column to 1 for the matching row, which acts like flipping a switch that tells the scraper not to revisit that link again. This idea is similar to maintaining a checklist during a long task: once an item is marked as done, it never needs attention again unless something goes wrong. By fitting naturally with earlier helper functions—like the one that fetches pending URLs and the one that stores scraped product details—the mark\_scraped function becomes part of a continuous story: URLs are discovered, processed, safely stored, and finally checked off, creating a clear and predictable cycle that keeps the scraping process organized from start to finish. ### Handling Reliable Fetching with curl-cffi ``` # FETCHERS def fetch_using_curl(url: str, ua: str, retries: int = RETRIES) -> Optional[str]: """ Fetch HTML content using curl-cffi (requests with browser impersonation). """ headers = {"User-Agent": ua, **HEADERS} for attempt in range(1, retries + 1): try: r = requests.get(url, impersonate="chrome", headers=headers, timeout=30) status = getattr(r, "status_code", None) if status == 200 and getattr(r, "text", None) and len(r.text) >= MIN_HTML_LEN: logging.info(f"cURL success [{len(r.text)} bytes]: {url}") return r.text logging.debug(f"cURL short/non-200 (status={status}) attempt {attempt}: {url}") except Exception as e: logging.warning(f"cURL attempt {attempt} error for {url}: {e}") return None ``` Fetching web pages reliably can be tricky, especially when sites try to block automated access, and the fetch\_using\_curl function provides a beginner-friendly way to handle this by using curl-cffi, a Python library that acts like a real browser. This function starts by taking a URL and a user-agent string, which tells the website which browser is being simulated—this is important because sites often serve different content or block requests that don’t look like they come from a real user. Inside the function, custom headers are built to resemble a standard browser request, and a loop tries multiple times to fetch the page in case of temporary network issues or bot protection that can return very short HTML pages. Each attempt sends a GET request using curl-cffi’s impersonation mode, which mimics Chrome, and the response is checked to ensure the status code is 200 and the HTML length is above a minimum threshold, which avoids saving incomplete or error pages. If the page is successfully fetched, it logs the success with the size of the response, and if it fails, warnings or debug messages record the status and attempt number for easier troubleshooting. This approach fits naturally with other parts of a scraper, such as database helpers like fetch\_pending\_urls and save\_product, creating a smooth workflow from discovering URLs to safely storing product data . ### How Playwright Helps Load Dynamic Amazon Pages ``` # PLAYWRIGHT FETCHER async def playwright_fetch(page: Page, url: str, retries: int = RETRIES) -> Optional[str]: """ Fetch full HTML content using Playwright with retries. """ for attempt in range(1, retries + 1): try: resp = await page.goto( url, timeout=PLAYWRIGHT_TIMEOUT, wait_until="networkidle" # Wait until no network requests ) # Additional waits for Amazon dynamic elements await page.wait_for_load_state("networkidle") await page.wait_for_timeout(3000) # Wait 3 seconds extra # Wait for essential selectors selectors = [ "#productTitle", "#feature-bullets", "#bylineInfo", "#detailBulletsWrapper_feature_div", "#prodDetails", ] for sel in selectors: try: await page.wait_for_selector(sel, timeout=5000) except: pass # ignore if not found (not mandatory fields) html = await page.content() if html and len(html) >= MIN_HTML_LEN: logging.info(f"Playwright success [{len(html)} bytes]: {url}") return html logging.warning(f"Playwright short HTML attempt {attempt} for {url}") await asyncio.sleep(2.0) except Exception as e: logging.warning(f"Playwright attempt {attempt} error for {url}: {e}") await asyncio.sleep(2.0) return None ``` Using playwright\_fetch provides a modern way to fetch web pages that load content dynamically with JavaScript, which is common on sites like Amazon, and this function is designed to handle those challenges in a beginner-friendly way. It starts by taking a Playwright Page object and a URL, then tries multiple times to load the page while waiting until network activity has settled, ensuring that all dynamic elements are loaded. After the main page load, the function pauses a few extra seconds and checks for essential elements like the product title, features, and details, but it doesn’t fail if some of these selectors are missing, which makes it flexible across different product pages. Once the page is fully loaded and the HTML content meets a minimum length requirement, it logs the success with the size of the page, which helps track scraping progress. If an attempt fails due to network issues or incomplete content, it waits a short period and retries, providing a robust way to handle transient errors. This approach works smoothly alongside other tools in the scraper, like fetch\_using\_curl for simpler requests and database helpers such as save\_product to store clean data. ### A Simple Helper for Safely Extracting Text ``` # PARSING HELPERS def safe_text(soup: BeautifulSoup, selector: str) -> Optional[str]: """ Safely extract text from a BeautifulSoup selector """ el = soup.select_one(selector) return el.get_text(strip=True) if el else None ``` The safe\_text function is a simple yet powerful helper for extracting text from HTML using BeautifulSoup, which is especially useful when scraping pages that may or may not have certain elements, like product descriptions or titles. It takes a [BeautifulSoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co) object representing the parsed HTML and a CSS selector string, then looks for the first matching element. If the element exists, it returns the cleaned text without extra spaces; if it’s missing, it safely returns None instead of causing an error, which helps prevent the scraper from crashing on inconsistent pages. This approach fits seamlessly with other parts of the scraper, like playwright\_fetch or fetch\_using\_curl, because it allows the dynamic or static HTML content to be parsed reliably, and the results can then be stored in the database using functions like save\_product. ### How Product Information Is Pulled from Rows ``` # PARSE KEY-VALUE ROWS FROM def parse_key_value_rows(soup: BeautifulSoup) -> dict: """ Extract key–value information from table rows (). """ out = {} for tr in soup.select("tr"): th = tr.find("th") td = tr.find("td") if not th or not td: continue key = th.get_text(separator=" ", strip=True) val = td.get_text(separator=" ", strip=True) if not key: continue k = re.sub(r"\s+:\s*$", "", key.strip().lower()) out[k] = val return out ``` The parse\_key\_value\_rows function offers a gentle way to turn messy product detail tables into clean, structured data by scanning through each element in the HTML and picking out the and pairs that usually hold important product facts like brand names, item weights, or manufacturer details. It checks every row safely, skips incomplete ones, and carefully cleans both the key and value so the final result feels consistent and easy to use. The key text is converted to lowercase, extra spaces are trimmed, and trailing colons are removed, which means “Brand:” and “Brand” are treated the same. Each cleaned key and value is then stored in a simple dictionary, allowing the rest of the scraper—whether it uses playwright\_fetch, fetch\_using\_curl, or helper functions like safe\_text—to rely on predictable data instead of dealing with different table formats. This becomes particularly helpful on pages like Amazon’s, where the layout can shift but the basic idea remains the same. The overall flow of this function makes the data extraction step feel like a natural continuation of the scraping process, turning raw table rows into readable information without adding unnecessary complexity. ### How List Items Are Parsed Into Structured Product Details ``` # PARSE LIST ITEMS FOR KEY-VALUE PAIRS def parse_list_items(soup: BeautifulSoup) -> dict: """ Extract key–value data from
  • list items. """ out = {} for li in soup.select("li"): # find bold span commonly used on Amazon: Key : bold = li.select_one(".a-text-bold") if not bold: # try finding a span with class containing 'a-text-bold' bold = li.find(lambda t: t.name == "span" and t.get("class") and any("a-text-bold" in cls for cls in t.get("class"))) if not bold: continue key = bold.get_text(" ", strip=True) key = re.sub(r"\s*:\s*$", "", key).strip().lower() try: bold.extract() except Exception: pass val = li.get_text(" ", strip=True) if val: out[key] = val return out ``` The parse\_list\_items function plays a helpful role when product details appear inside bullet lists, especially on pages like Amazon where a bold label is often followed by the actual information, and the goal of this function is to gently separate these two parts and turn them into clean key–value pairs that a scraper can easily understand. It moves through each
  • element one by one, looks for the bold span that usually marks the key, cleans it by trimming extra spaces and removing trailing colons, and then removes that bold span so the remaining text becomes the value. This approach works well with Amazon-style layouts where bold text such as “Manufacturer:” or “Item Weight:” is placed at the start, followed by the details. The outcome is stored in a simple dictionary, making it easy for other parts of the scraper—whether they rely on playwright\_fetch, fetch\_using\_curl, or helpers like safe\_text and parse\_key\_value\_rows—to work with consistent data. This creates a smooth flow where the raw HTML first gets fetched, then cleaned, and finally organized into structured information, helping beginners see how every step naturally leads into the next without adding unnecessary technical difficulty. ### Finding the ASIN: How the Code Locates Amazon’s Product ID ``` # EXTRACT ASIN def extract_asin(soup: BeautifulSoup, html: str) -> Optional[str]: """ Extract the ASIN (Amazon Standard Identification Number) from a product page. """ # 1) from product details kv = parse_key_value_rows(soup) for kname, v in kv.items(): if "asin" in kname: m = re.search(r"([A-Z0-9]{10})", v) if m: return m.group(1) return v # 2) input fields el = soup.select_one("#ASIN") or soup.select_one("input[name='ASIN']") if el: return el.get("value") or el.get_text(strip=True) # 3) regex m = re.search(r"ASIN[^A-Z0-9]*([A-Z0-9]{10})", html, re.I) if m: return m.group(1) return None ``` Understanding how the extract\_asin function works becomes much easier when thinking of an Amazon product page as a large room full of scattered clues, and the ASIN is the small but important label that identifies the product, similar to a unique ID on a warehouse box; the function patiently walks through different corners of this room to locate that label, starting with the product-detail section where Amazon often places key information inside neat table rows, and this part is decoded using a helper like parse\_key\_value\_rows, which gathers structured data from HTML just like explained earlier in internal sections of the project; if the ASIN hides somewhere else, the function then checks input fields inside the page—tags such as —because Amazon sometimes stores important data in hidden fields meant for internal site use; and when both of these methods come up empty, the function uses a regex search directly on the raw HTML string, which acts like scanning the entire page for any text that matches the ASIN pattern, ensuring nothing is missed; this multi-step approach works well because Amazon frequently changes layouts, and relying on one fixed selector is rarely enough, so the function quietly adapts by checking several likely spots before giving up, the heart of this function is simply about being thorough—reading the page carefully, searching step by step, and returning the ASIN only when a confident match is found, making the overall scraper more reliable even when Amazon updates page designs. ### Step-by-Step Logic Behind Amazon Product Parsing ``` # GATHER DETAILS FROM SOUP def gather_details_from_soup(soup: BeautifulSoup, html: str) -> dict: """ Extract key product details from an Amazon product page. """ kv = parse_key_value_rows(soup) li_map = parse_list_items(soup) def get_field(possible: List[str]) -> Optional[str]: for name in possible: ln = name.lower() if ln in kv: return kv[ln] if ln in li_map: return li_map[ln] return None # Only extracting specified fields title = safe_text(soup, "#productTitle") brand = safe_text(soup, "#bylineInfo") or get_field(["brand", "manufacturer"]) price = safe_text(soup, "span.a-price-whole") or safe_text(soup, ".a-price .a-offscreen") mrp = safe_text(soup, "span.a-text-price span[aria-hidden='true']") or get_field(["mrp", "list price", "price"]) # Extract ASIN only for URL cleaning asin = extract_asin(soup, html) or get_field(["asin"]) return { "title": title, "brand": brand, "price": price, "mrp": mrp, "asin": asin # Only used for URL cleaning } ``` The gather\_details\_from\_soup function works like a careful reader that goes through an Amazon product page and slowly picks up the details that matter, using both structured sections and scattered elements in the HTML to build a clear set of values; it begins by calling parse\_key\_value\_rows and parse\_list\_items, which act like small helpers that organize information found inside table rows and list items, making it easier to look up fields later without searching the whole page again, and then defines a tiny method named get\_field that patiently checks different possible names for the same detail—because one product may label the brand as “Brand” while another uses “Manufacturer”—so the function remains flexible even when layouts vary; once these foundations are ready, the code turns its attention to visible elements using safe\_text, reading the product title from [#productTitle](https://www.blog.datahut.co/blog/hashtags/productTitle), the brand from [#bylineInfo](https://www.blog.datahut.co/blog/hashtags/bylineInfo), and the price from selectors such as .a-price .a-offscreen, with the function gently switching to fallback options when something is missing; the MRP is gathered the same way, sometimes through a selector and sometimes from the earlier dictionaries, depending on the page structure; finally, the ASIN is extracted with extract\_asin, which checks multiple places inside the HTML before giving up, making it useful for tasks like URL cleaning where this identifier is needed, by the time the function finishes, it returns a simple dictionary holding the title, brand, price, MRP, and ASIN, bringing together all the pieces it collected in a way that feels like watching a puzzle come together naturally, without abrupt jumps or isolated steps. ### Scraping a Single Product Page with Dual-Fetch Strategy ``` # Scrape single product (dual fetch) async def scrape_product(page: Page, url: str, ua: str) -> Optional[dict]: """ Scrape a single Amazon product page using a dual-fetch strategy. """ html = await playwright_fetch(page, url) if html is None: html = fetch_using_curl(url, ua) if html is None: logging.error(f"Both curl & Playwright failed for {url}") return None soup = BeautifulSoup(html, "html.parser") details = gather_details_from_soup(soup, html) asin = details.get("asin") details["url"] = clean_amazon_url(url, asin) # Remove asin from final data since we don't store it for sample data if "asin" in details: del details["asin"] return details ``` The scrape\_product function works like a careful two-step safety net for loading an Amazon product page, beginning with Playwright to handle pages that rely on dynamic content and then quietly switching to a curl-based request if the first attempt does not return any HTML, giving the script a dependable way to keep moving without stopping at the first failure; once the HTML is available, the function hands the page to BeautifulSoup, which works like a gentle reader that turns the raw markup into something easier to navigate, allowing gather\_details\_from\_soup to pull out the title, brand, price, MRP, and the ASIN with the help of several parsing helpers placed throughout the codebase; after collecting the data, the function creates a clean Amazon URL through clean\_amazon\_url, using the ASIN to remove long tracking parameters and give a neater link, similar to tidying a long string into something short and readable, and as the final touch, the ASIN is removed from the output because it is needed only for URL cleaning and not for storing in the final dataset, keeping the returned dictionary focused on the essential fields; with this flow, the entire process feels like watching someone methodically fetch the page, fall back when necessary, parse the content, extract meaningful information, and hand back an organized result without abrupt jumps or technical clutter, making the logic easy to follow even for someone new to scraping. ### Coordinating the Scraping Workflow with the Main Function ``` # MAIN RUNNER async def main(limit: Optional[int] = None): """ Main controller function for scraping Amazon product pages. """ init_db() pending = fetch_pending_urls(limit) if not pending: print("No URLs left.") logging.info("No URLs left to process.") return async with async_playwright() as pw: browser = await pw.chromium.launch(headless=False) try: for row_id, url in pending: # Pick user agent once for this URL ua = random.choice(USER_AGENTS) # Create context/page per URL context = await browser.new_context(user_agent=ua) page = await context.new_page() # Apply stealth BEFORE scraping try: await stealth_async(page) except Exception as e: logging.debug(f"stealth_async warning: {e}") # Scrape product using dual method (Playwright → curl) try: data = await scrape_product(page, url, ua) except Exception as e: logging.exception(f"Unexpected scrape error for {url}: {e}") data = None # Save only if successful if data: save_product(data) mark_scraped(row_id) logging.info(f"Saved & marked scraped: {url}") else: logging.error(f"Failed to scrape (kept pending): {url}") await context.close() await asyncio.sleep(random.uniform(*SLEEP_RANGE)) finally: await browser.close() ``` The main function acts like the central coordinator that quietly guides the entire scraping process from start to finish, beginning with a simple step of pulling the list of product links stored in a SQLite database and then preparing a Playwright browser session so each page can be visited in a controlled environment; once the pending URLs are ready, the function moves through them one by one, choosing a random User-Agent for each request so the scraper behaves more like a regular visitor, and resources such as new browser contexts are created fresh for every link to avoid letting any previous page leave behind clues that automation is being used, which is especially important for websites with detection systems; before opening each page, stealth mode is enabled to reduce signs of automation, and after that, the scraper tries to collect details through scrape\_product, where Playwright handles most pages and curl-cffi becomes a fallback if dynamic loading fails; successful results are saved in the database using helper functions like save\_product, and the URL is marked as scraped so the system knows not to repeat the work the next time the script runs, allowing the whole workflow to feel smooth and dependable even if some pages fail and remain pending for later attempts; the function closes each context carefully, adds a small random delay to mimic normal browsing habits, and eventually shuts down the browser once every link has been processed, bringing the scraping cycle to a clean finish without returning data directly because everything is already written into SQLite and captured through logs for review. ### Endpoint ``` # ENDPOINT """ Runs the main function when the script is executed directly.""" if __name__ == "__main__": # optional: run with a limit asyncio.run(main(limit=None)) ``` This small endpoint acts like the front door of the script, making sure the main scraping routine starts only when the file is run directly rather than being imported somewhere else. Think of it as a simple trigger that tells Python to launch the main function, which then carries the entire workflow forward. Adding the optional limit parameter gives extra control during testing, helping the script run safely without processing more pages than needed while still keeping the flow of execution clean and predictable. ## Conclusion Choosing a menstrual cup becomes much simpler once the basic ideas behind its design, material safety, and long-term benefits are clearly understood, and going through the details on trusted product pages—whether from official brand websites or well-structured informational resources like health blogs and verified guides—helps build confidence for anyone trying it for the first time. As the properties of the cup start to make sense, such as how medical-grade silicone keeps it flexible yet safe, or how its reusable nature reduces monthly waste, the product stops feeling like a complicated new tool and more like a practical alternative worth considering. Exploring FAQs, user experiences, and comparison sections on the same site brings a clearer picture of how sizing works, how insertion becomes easier with practice, and how proper cleaning keeps everything hygienic. External references from reliable health platforms also reassure beginners that menstrual cups are globally recommended for comfort, sustainability, and cost-effectiveness. After understanding these points step-by-step, the idea of switching no longer feels overwhelming; instead, the cup comes across as a thoughtful, modern solution that supports comfort, confidence, and a more eco-friendly period experience—something that grows easier and more empowering with each cycle. ## Libraries and Versions Used Name: asyncio Version: Built-in Python module Name: random Version: Built-in Python module Name: sqlite3 Version: Built-in Python module Name: json Version: Built-in Python module Name: logging Version: Built-in Python module Name: re Version: Built-in Python module Name: contextlib Version: Built-in Python module Name: urllib.parse Version: Built-in Python module Name: pathlib Version: Built-in Python module Name: BeautifulSoup (bs4) Version: 4.12.3 Name: playwright Version: 1.48.0 Name: playwright-stealth Version: 1.0.6 Name: curl-cffi (requests module) Version: 0.6.2 ## AUTHOR I’m Anusha P O, a Data Science Intern at Datahut, with hands-on experience in building intelligent, scalable web-scraping systems. In this blog, I break down how we extracted structured product information for Menstrual Cup listings from Amazon using a hybrid scraping pipeline powered by Playwright, [curl-cffi](https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/), SQLite, JSON, and an async dual-fetch workflow.From handling anti-bot challenges to cleaning and normalizing product attributes, this project shows how messy Amazon product pages can be transformed into clean, reliable, analysis-ready datasets. At [Datahut](https://www.datahut.co/?ref=blog.datahut.co), we help businesses unlock the power of web data by designing robust scraping architectures for price monitoring, product research, competitive intelligence, and large-scale e-commerce analytics.If you’re working on data-driven strategies for online retail or want to organize large product datasets efficiently, feel free to reach out through the chat widget on the right.Let’s turn raw web data into clear, actionable insights. FAQ SECTION ### 1\. What is the purpose of scraping Amazon’s menstrual cup data? Scraping menstrual cup data helps analyze pricing, reviews, ratings, product features, and competitor positioning. It enables ecommerce sellers and analysts to make informed decisions based on real-time market trends. ### 2\. Why use Playwright for scraping Amazon product pages? Playwright is ideal for handling dynamic content, JavaScript rendering, and tough anti-bot measures. It mimics human browsing behavior, making it more reliable for scraping Amazon product detail and listing pages. ### 3\. How does curl-cffi help in scraping Amazon? curl-cffi bypasses strict bot detection by using TLS fingerprinting identical to real browsers. This makes it effective for fetching Amazon HTML pages without being blocked. ### 4\. Is it legal to scrape Amazon? Web scraping Amazon for personal research, price monitoring, or competitive insights is generally allowed when done ethically—without breaching login walls, violating robots.txt, or harming servers. Always follow Amazon’s terms and respect data privacy laws like GDPR. ### 5\. What insights can be extracted from menstrual cup product data? You can extract price variations, best-selling brands, rating distribution, keyword-rich descriptions, feature comparisons, and customer sentiment insights to support ecommerce product analysis. ### Web Crawling: Use Cases & Business Benefits for Companies URL: https://www.blog.datahut.co/post/web-crawling-and-its-use-cases-for-2026/ Last updated: 2026-09-07T09:43:49.000Z In the data-driven landscape of 2026, access to external web data isn't just an advantage, it's a baseline requirement. However, acquiring this data efficiently remains a major hurdle. Many businesses find themselves navigating high operational costs and complex technical barriers just to keep their data pipelines flowing. And that’s exactly why [web crawling has become one of the most valuable capabilities for businesses in 2026](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/)[.](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/) Nearly every company today relies on external data competitor prices, market trends, customer reviews, job postings, product catalogs, regulatory updates, and more. The surprising truth? Most of this data is already public. The challenge isn’t access, it’s scale, accuracy, compliance, and freshness. Web crawling solves all of that. In this blog, we break down: - What web crawling actually is (2026 version) - Why it’s more important now than ever - Real-world, high-impact use cases across industries - How businesses save time, money, and engineering effort - Why companies are shifting from buying third-party data to owning their crawling pipelines ## What Is Web Crawling? Web crawling (often referred to as a web spider or bot) is the automated process of discovering and visiting web pages by following links and sitemaps, usually to build a collection of URLs that can then be scraped or indexed like a search engine. In simple terms: A crawler is a smart bot that moves through the World Wide Web on its own discovering pages, following links, obeying site rules, and tracking what’s new or updated. It’s closely related to web scraping, though they serve different functions. For a deeper technical comparison, you can explore the nuances of[ ](https://www.blog.datahut.co/post/web-scraping-vs-api/)[web scraping vs APIs and crawling](https://www.blog.datahut.co/post/web-scraping-vs-api/). - Web crawling → discovers and organizes URLs using an XML sitemap or link following. - Web scraping → extracts specific data (e.g., price, title, rating) from those URLs, converting raw HTML content into structured data. In real systems, they work together: Crawler: “Find me all product pages in this category across these sites.” Scraper: “From each page, extract the product name, price, stock, rating, image, etc.” Modern crawlers are far from the simple bots of 2010\. At scale they need to handle: - JavaScript-heavy frontends and headless browser rendering. - Geo-targeted content and distributed crawling. - Managing user agents to appear as legitimate traffic. - Rate limits and CAPTCHAs. - Bot protection systems like[ ](https://www.cloudflare.com/en-gb/application-services/products/bot-management/?ref=blog.datahut.co)[Cloudflare Bot Management](https://www.cloudflare.com/en-gb/application-services/products/bot-management/?ref=blog.datahut.co) and similar tools. ![the modern crawler workflow ](https://www.blog.datahut.co/content/images/2026/07/img-69.jpg.webp) ## Why Web Crawling Is a Bigger Deal in 2026? Two big shifts changed the game: ### 1\. Anti-bot & AI controls got serious Cloudflare and similar providers protect millions of active websites and now offer one-click blocks for AI scrapers and crawlers, with default blocking for many new domains. On top of that,[ recent disputes](https://techcrunch.com/2025/03/27/open-source-devs-are-fighting-ai-crawlers-with-cleverness-and-vengeance/?utm%5Fsource=chatgpt.com) made ethical, compliant crawling a board-level topic. To navigate these barriers effectively, businesses must understand[ ](https://www.blog.datahut.co/post/how-to-maintain-anonymity-when-web-scraping-at-scale-expert-tips/)[how to maintain anonymity when web scraping at scale](https://www.blog.datahut.co/post/how-to-maintain-anonymity-when-web-scraping-at-scale-expert-tips/). ### 2\. Bot traffic exploded [Reports suggest](https://www.imperva.com/resources/resource-library/reports/bad-bot-report/?ref=blog.datahut.co) that over half of internet traffic is now bots, and a big portion of that is malicious or non-compliant. Organizations are struggling to distinguish “good” bots (like search engines or compliant crawlers) from abusive ones. Result: Businesses that want to use web data now need serious, compliant crawling infrastructure, not hobby Python libraries or scripts. That’s exactly where a well-designed crawling strategy or a managed vendor starts saving huge money. ## How Businesses Benefit: Web Crawling Use Cases (2026) ### Web Crawling Use Cases Table ![how businesses benefit from web crawling](https://www.blog.datahut.co/content/images/2026/07/img-70.jpg.webp) ## Deep Dive into Key Use Cases ### 1\. Price Intelligence & Dynamic Pricing Price is still one of the most powerful growth levers.[ ](https://www.blog.datahut.co/post/competitive-price-intelligence-can-determine-profitability/)[Competitive price intelligence can determine profitability](https://www.blog.datahut.co/post/competitive-price-intelligence-can-determine-profitability/) by allowing you to see exactly where you stand in the market. Crawlers constantly visit competitor sites, marketplaces, and even regional country sites for price monitoring to collect: - Current product prices - Discounts & promotions - Stock / availability - Shipping fees - Bundle offers This feeds: - [Dynamic pricing systems](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/how-we-help-clients/dynamic-pricing?ref=blog.datahut.co) that update prices in near real time. In fact, effective[ ](https://www.blog.datahut.co/post/dynamic-pricing-can-boost-online-retail-profits/)[dynamic pricing can boost online retail profits](https://www.blog.datahut.co/post/dynamic-pricing-can-boost-online-retail-profits/) significantly when backed by accurate real-time data. - MAP monitoring to ensure resellers don’t undercut agreed minima. - Margin optimization based on competitors’ moves and demand. Multiple providers highlight price intelligence as one of the dominant web data use cases in 2026. ### 2\. Product Data & Catalog Enrichment [If you run a marketplace, comparison site, or aggregator](https://www.blog.datahut.co/post/ecommerce-marketplace-data-for-business/), you can’t manually copy product details from thousands of suppliers. Web crawling + scraping help you: - Discover all product URLs in a category - Extract specs, features, images, and descriptions - Normalize attributes (size, color, material, etc.) - Keep data fresh when suppliers change content Analyses of data extraction applications show that [product data extraction](https://www.shopify.com/blog/product-description?ref=blog.datahut.co) and catalog building remain core use cases across industries. Impact: - Faster time-to-market for new products - Consistent, rich catalog without manual data entry - Better search, filters, and SEO performance ### 3\. SEO, SERP & Content Intelligence Search teams crawl: - Google/Bing SERPs for target keywords - Their own sites for broken links, redirects, and metadata - Competitor blogs / docs / landing pages [This helps them detect ranking losses early, identify content gaps, and fix technical SEO issues at scale.](https://www.blog.datahut.co/post/search-rank-tracking/) Instead of paying agencies for static [“SEO audits”](https://ahrefs.com/blog/technical-seo/?ref=blog.datahut.co) every quarter, teams run continuous crawling to keep a live picture of their search visibility. ### 4\. Lead Generation & B2B Intelligence For B2B companies, crawling can turn the open web into a structured lead engine. Leveraging[ ](https://www.blog.datahut.co/post/web-scraping-for-lead-generation/)[web scraping for lead generation](https://www.blog.datahut.co/post/web-scraping-for-lead-generation/) transforms static directories into live intent data: - Job boards → hiring patterns (e.g., “Hiring 5 data engineers” = strong growth signal) - Startup & company directories → basic firmographic data - Public company websites → product lines, locations, tech stack hints The difference in 2026 is quality and context: instead of just collecting emails, teams look at signals like rapid headcount growth, opening new locations, or heavier hiring in artificial intelligence. Those signals can feed scoring models and outbound sequences. ### 5\. Brand Protection & Fraud Monitoring If your brand has value, someone will try to abuse it. Crawlers patrol: - Marketplaces for counterfeit products - Gray/black-market sites for stolen or discounted items - Unofficial “support” sites using your logo - Social or classifieds listings using your brand assets Research on web scraping use cases highlights brand protection and compliance monitoring as growing applications, especially as fraud moves online. Benefits: - Faster takedowns - Reduced IP leakage and counterfeit impact - Stronger control over channel pricing and representation ### 6\. Real Estate & Travel Aggregation [Travel and property platforms](https://www.blog.datahut.co/post/web-scraping-for-travel-industry/) rely heavily on crawling to stay competitive: - Hotels, vacation rentals, and airlines → prices, availability, policies - Real-estate portals → listings, photos, amenities, neighborhood information This enables Metasearch engines, price-comparison widgets, and market trend dashboards. The value here is consistency and freshness—if your pricing or inventory is outdated, users bounce. ### 7\. Financial & Alternative Data Investors and analysts are hungry for “alternative datasets” that give an edge before official earnings. Crawling gives them a structured, time-series view of signals like job postings, product reviews, pricing changes, and public announcements. Web data is now a core part of market research, revenue forecasting, competitive intelligence, and risk analysis. ### 8\. AI & LLM Training Data Large language models and AI agents need clean, domain-specific training data: - Documentation & API references - Knowledge bases and FAQs - Public blogs and help centers - Regulatory and standards documents Many platforms now position themselves as web data providers for AI, emphasizing ethical, public data collection and robust compliance. Web crawling is the discovery backbone of those AI training pipelines. ## Build vs Buy: Why Many Teams Choose Managed Crawling in 2026 On paper, building your own crawler looks simple. In reality, the moment you operate at scale, you face a long list of moving parts. Businesses often underestimate the difficulties involved in[ scaling web scraping from prototype to production](https://www.blog.datahut.co/post/scaling-web-scraping-from-prototype-to-production-challenges-explained/): - IP rotation & global proxy networks - JavaScript rendering using headless browsers - CAPTCHAs and enterprise-grade anti-bot systems - Continuous HTML and DOM structure changes - Monitoring, retries, logging, and crawl budgets - Data quality validation and schema enforcement Most engineering teams don’t struggle with the first version of a crawler — they struggle with keeping it alive. That’s why leading data vendors highlight “zero-maintenance, fully managed crawling” as the primary reason enterprises outsource. Internal teams are often crushed not by building the crawler, but by the ongoing maintenance and compliance burden. ![managed web crawling](https://www.blog.datahut.co/content/images/2026/07/img-71.jpg.webp) This is exactly where Datahut comes in. Datahut provides a fully managed, compliance-first web crawling and data delivery service designed for modern, JS-heavy, anti-bot-protected websites. Instead of maintaining fragile scripts, teams use Datahut to: - Scale effortlessly across millions of pages without touching infrastructure - Avoid IP blocks with intelligent rotation, fingerprinting, and region-specific access - Handle complex rendering with production-grade headless browser pipelines - Stay compliant with regional regulations, and ethical crawling rules - Receive clean, analysis-ready data instead of raw HTML - Eliminate maintenance issues ,Datahut handles breakages, retries, and updates In short: You focus on insights. Datahut handles the crawling, compliance, and engineering complexity. If your teams need accurate, fresh, and scalable web data without the engineering complexity, Datahut can help. Our fully managed crawling infrastructure delivers clean, analysis-ready datasets tailored to your business needs — with zero maintenance, zero downtime, and complete compliance. Talk to us today:[ ](https://www.datahut.co/solutions?ref=blog.datahut.co)[https://www.datahut.co/solutions](https://www.datahut.co/solutions?ref=blog.datahut.co) ## Compliance, Ethics & 2026 Reality Check With all the hype around AI and data, it’s easy to forget the basics. To operate safely, you need a clear[ ](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/)[guide to legal and transparent data practices in web scraping](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/): Ensure GDPR compliance at every stage: Data Privacy Laws like [GDPR](https://gdpr.eu/?ref=blog.datahut.co) require organizations to follow principles like data minimization, purpose limitation, lawful basis, and transparency. Even public-facing data can be subject to privacy rules if it contains personal identifiers. A responsible crawling pipeline must include governance controls, validation mechanisms, and audit trails to prevent accidental collection or misuse of personal data. Any serious web crawling strategy in 2026 must integrate this into the design from day one. ## The Hidden Costs of Web Scraping When evaluating web crawling strategies, the "free" option of building internally often carries the highest long-term price tag. 1. Building Internally Diverting engineering resources to build crawlers pulls them away from your core product. Your team ends up managing a scraping infrastructure company inside your actual company, reducing focus on your primary business goals. 2. Investing in Building Own Infra The cost of maintaining a robust crawling infrastructure is deceptive. Beyond server costs, you face escalating expenses for residential proxies, CAPTCHA solving services, and headless browser clusters required to bypass modern anti-bot systems. 3. Talent Shortage Specialized web scraping engineers are difficult to find and expensive to hire. This is a niche skill set involving deep knowledge of reverse engineering, network protocols, and browser fingerprinting—skills that generalist full-stack developers often lack. 4. Complexities at Scale A crawler that works for 1,000 pages often fails at 10 million. Handling scale requires sophisticated logic for concurrency, error handling, and data validation. As websites update their layouts and security measures, internal teams are trapped in a constant cycle of "fix and patch" maintenance. 5. [Legal](http://5.legal/?ref=blog.datahut.co) & Compliance Liability: Navigating the minefield of global data privacy laws (GDPR, CCPA) and emerging AI regulations is a full-time job. Internal teams often lack the specialized governance frameworks to handle IP rights, cross-border data transfers, and PII protection, exposing the organization to significant regulatory risk and potential lawsuits. 6. The Data Quality Trap: Scraping the web is only half the battle; making the data usable is the other. Raw web data is notoriously messy, full of duplicates, broken HTML, and inconsistent formatting. You will likely spend as much time building and maintaining normalization, deduplication, and QA pipelines as you do on the crawler itself. ## Final thought Web crawling in 2026 has become a critical capability for any organization that relies on external data to stay competitive. As industries shift toward real-time decision-making, companies increasingly recognize that depending on third-party data providers creates limitations in cost, accuracy, and flexibility. Public web data is abundant, but accessing it consistently and responsibly requires the right infrastructure. A modern crawling strategy offers several advantages: - Greater control over data sources and update frequency - Improved accuracy and freshness for operational decisions - Reduced long-term cost compared to high-priced external datasets - Stronger visibility into markets, competitors, and customer behavior - A scalable foundation for analytics, automation, and machine learning initiatives - Better compliance and governance through transparent data lineage However, building this capability internally is challenging. Teams must manage rendering, proxies, anti-bot systems, schema changes, monitoring, retries, and legal considerations. The complexity grows every month as websites evolve. ## Frequently Asked Questions (FAQ) 1.What is web crawling? Web crawling is the automated process of discovering and navigating web pages through links and sitemaps. It helps map a website so data can be collected efficiently. 2\. How is web crawling different from web scraping? Crawling finds and organizes URLs, while scraping extracts specific information such as product prices, reviews (for sentiment analysis), or product details. They usually work together in real-world data pipelines. 3\. Is web crawling legal in 2026? Yes, crawling public pages is legal when you follow site terms and Data Privacy Laws like GDPR and CCPA. The legality depends on responsible behavior and avoiding personal data collection without a lawful basis. 4\. Why do companies need web crawling in 2026? Businesses rely on crawling for competitive pricing, product intelligence, market research, and AI training data. It gives them fresher, more accurate datasets than buying static third-party reports. 5\. Why do businesses outsource web crawling instead of building it? Maintaining crawlers at scale requires handling proxies, anti-bot systems, user agents, JavaScript rendering, and constant site changes. Outsourcing removes this burden so teams can focus on using reliable data insights instead of fixing broken crawlers. ### How Small Typos Hurt Conversions on Amazon Marketplace URL: https://www.blog.datahut.co/post/amazon-data-hygiene-webscraping/ Last updated: 2026-09-07T09:43:51.000Z Have you ever wondered how much money a single typo could be [silently draining from your Amazon sales](https://www.blog.datahut.co/post/discover-the-hidden-value-in-amazon-data-a-comprehensive-guide/) ? Not a bad review. Not a pricing mistake. Not a logistics issue. Just one incorrect letter—quietly wrecking your discoverability, relevance, and conversions. It sounds absurd… until you see it happen, and a bit of web scraping and large language models can help you fix it. Last week, while analyzing product data in the femcare category, I stumbled upon something that looked trivial… but turned out to be a silent revenue leak. One brand had consistently misspelled “Small” as “Smal” — not in the bullet points, not in the description, but right in the title across dozens of listings. ![Amazon data hygiene](https://www.blog.datahut.co/content/images/2026/07/img-55.jpg.webp) Anyone in femcare knows: Size is not a minor attribute. It’s a core purchase driver. Yet because of this one tiny mistake: - Customers searching for “small” pads couldn’t find the products - Browsers who did land on it became confused or suspicious - Amazon’s search algorithm reduced the listing’s relevance for “small” queries - Competitors ranking correctly for the keyword overtook it effortlessly This wasn’t just an embarrassing oversight. It was algorithmic self-sabotage. - A missing letter was costing them visibility. - Lost visibility cost them clicks. - Fewer clicks cost them conversions. - Lower conversions triggered further ranking decline. - And that downward spiral was happening quietly, every single day. They weren’t losing because of competition. They were selling female hygiene products — but their data hygiene was poor. They were losing because of a typo. ## The Compounding Effect of Bad Catalog Data on Amazon Rankings On Amazon — and on most marketplaces — your product data is the signal the algorithm relies on. - Every misspelling → lower relevance - Every inconsistent attribute → lower ranking - Every mismatch → lower trust - Every missing detail → lower conversions These small inconsistencies accumulate into major performance drops. Research across metadata-heavy industries (digital libraries, academic catalogs, commercial search systems) consistently shows that poor metadata harms searchability and discoverability: No matter the platform, one rule is universal:Bad data erodes visibility. Good data amplifies it. ## Why Amazon Dashboards Hide Critical Data Hygiene Problems Amazon Seller Central and Brand Analytics show a slice of your listing — not the full reality.(For reference:[ ](https://sell.amazon.com/blog/amazon-seo?utm%5Fsource=chatgpt.com)[How Amazon SEO actually works](https://sell.amazon.com/blog/amazon-seo?utm%5Fsource=chatgpt.com)) They miss critical issues because: 1. They show what you uploaded — not what Amazon renders. 2. They don’t reveal silent drift or overwritten fields. 3. They don’t show cross-catalog inconsistencies. 4. They don’t store historical snapshots. 5. They can’t detect unstructured errors (typos, odd phrasing, weak SEO signals). Humans catch these manually. Dashboards cannot. ## Why Teams Miss Critical Catalog Problems Most brands assume their catalog is clean because: - “We used templates.” - “The agency uploaded it correctly.” - “We optimized the listings at launch.” - “Everything looked fine six months ago.” But marketplaces silently: - merge content - suppress content - split listings - change category rules This creates catalog decay, where your metadata loses coherence over time. Quarterly audits can’t catch this but Continuous monitoring can. ## Scrape Amazon Product Data: The Missing Dimension of Amazon Catalog Optimization Every brand scrapes competitors. Almost none scrape their own listings. To understand why scraping matters, see:[ ](https://www.blog.datahut.co/post/how-scraping-amazon-data-can-help-you-price-your-products-right/)[How scraping Amazon data helps with pricing](https://www.blog.datahut.co/post/how-scraping-amazon-data-can-help-you-price-your-products-right/) Scraping your own catalog reveals: 1. The live version of your listing 2. Cross-listing consistency problems 3. Metadata drift 4. Image-level issues 5. Variant misalignment Scraping creates a mirror Amazon doesn’t provide. ## The Simple Amazon Data Hygiene Fix: Scrape → LLM Audit → Fix → Monitor This is the emerging 2025 standard for catalog quality. ### 1\. Scrape your entire catalog regularly If you know how to code:[Here’s a 20-line Python scraping guide](https://www.blog.datahut.co/post/web-scraping-in-python/) ### 2\. Feed this structured data into an LLM LLMs are exceptional at spotting: - typos - missing attributes - keyword dilution - duplicate bullets - tone inconsistencies - mismatched variants - broken image references - metadata drift They act as a semantic quality-control engine. ### 3\. Produce a Catalog Health Report The LLM's can produce a structured output like this which is action driven and anyone can fix. ``` SKU: FCM-SMALL-47 Issue 1: Title misspelling ("Smal" → "Small") Severity: HIGH Impact: Lost relevance for size-driven searches Fix: Correct title and reinforce keywordIssue Issue 2: Missing material attribute Severity: MEDIUM Impact: Lower trust, higher returns Fix: Add material to bullet Issue 3: Keyword drift detected Severity: HIGH Details: "Ultra-thin" was removed vs last month Impact: Ranking decline expected Fix: Reintroduce keyword naturally ``` ## What Happens When Brands Fix Catalog Hygiene Brands that adopt continuous QA typically see: 1. Higher organic ranking 2. Improved conversion rate 3. Lower ACOS 4. More stable Buy Box performance 5. Fewer returns 6. Stronger variant ecosystems For reference:[Amazon listing optimization best practices](https://gotrellis.com/resources/blog/amazon-listing-optimization/?utm%5Fsource=chatgpt.com) Catalog hygiene is not a content task. It is a profit function. ## Data Hygiene Is a Profit Lever — Not an Editorial Task Bad data quietly erodes: - visibility - conversion - relevance - buy box share - ad efficiency Good data amplifies all of the above. ## The Shift Every Amazon Seller Must Make in 2025 Old workflow: Upload → Forget → Notice when sales drop → Scramble to fixNew workflow: Scrape → Audit → Detect → Fix → Monitor → Repeat This is continuous catalog governance. ## Compare Your Catalog With Top Sellers — And Discover Hidden Patterns One of the most powerful ways to improve catalog hygiene is to compare your product metadata with the top-selling products. See example competitor scraping insights:[Price comparison via Amazon scraping](https://www.blog.datahut.co/post/price-comparison-on-amazon-how-web-scraping-helps-companies-win-the-e-commerce-game/) ## Look for Correlations — They Will Surprise You Patterns we’ve seen across categories include: - 5–7 images outperform 3–4 - Titles with size early convert better - Consistent variant images reduce bounce - Products with high attribute completeness rank better ## Going Beyond Fixing — Building a Culture of Catalog Quality Advanced teams use: - Weekly metadata snapshots - Variant governance playbooks - SEO drift alerts - Structured image audits - Competitor metadata monitoring - Pre-launch QA - Monthly governance reports For inspiration:[Listing quality dashboards explained](https://www.data4amazon.com/blog/chart-a-profitable-future-with-amazon-listing-quality-dashboard/?utm%5Fsource=chatgpt.com) ## The Bigger Picture: Metadata as a Strategic Asset When your catalog is clean: - search relevance strengthens - conversion rates lift - ads convert efficiently - organic ranking climbs - Buy Box share stabilizes Metadata is no longer a back-office chore. It’s a profitability lever. Related reading:[How metadata improves discoverability](https://blog.invgate.com/how-does-metadata-improve-data-quality?utm%5Fsource=chatgpt.com) ## Conclusion: The Smallest Details Decide the Biggest Outcomes A typo, a missing size, a broken image, a keyword that quietly dropped out, and a variant that fell out of sync may look like tiny, isolated issues. But on Amazon, these small cracks in your catalog create disproportionately large ripple effects—reducing relevance, confusing shoppers, weakening trust, and triggering algorithmic penalties that push your products further down the search results. Each flaw compounds the next, turning what appears to be minor oversights into silent, long-term revenue leaks. These don’t look like major problems. But on a marketplace where millions of products compete for the same shoppers, they create algorithmic consequences far larger than their appearance. ## The brands that win on Amazon in 2026 will be the ones that: - Scrape their own catalog regularly, - Use LLMs to audit metadata deeply, - Fix issues proactively, not reactively, - Monitor their listings continuously, - and make catalog hygiene a strategic discipline. Because marketplace success isn’t just about supply chains or ad budgets. It’s about the consistency, accuracy, and clarity of the data that represents your products. And in a world governed by algorithms, clean data is compounding leverage. ## Start Improving Your Catalog Today If this resonates, here are next steps you can take immediately: 1. Set up a weekly scrape of your own catalog - Even a simple script can surface issues your team has never seen. 2. Run your scraped data through an LLM - Ask it to detect inconsistencies, missing attributes, and signs of drift. 3. Create a basic “Catalog Quality Checklist” - Define how every title, bullet, variant, and attribute should look. 4. Start monitoring metadata like you monitor PPC - Treat catalog health as a measurable growth input. The smallest details move the biggest numbers. And the brands that understand this early will win the next decade of marketplace competition. ### Need an audit of your products or an entire category? Need an audit of your products or an entire category? Get in touch with Datahut. ## Amazon Data Hygiene FAQs ### 1\. What is Amazon data hygiene? Amazon data hygiene refers to the accuracy, consistency, and completeness of all product information in your Amazon catalog—including titles, bullets, images, attributes, keywords, pricing fields, and variant structures. Clean data improves search ranking, conversions, Buy Box share, and overall marketplace performance. ### 2\. Why is Amazon data hygiene important for ranking? Amazon’s search algorithm relies heavily on metadata like titles, attributes, keywords, and images. Poor data—typos, missing attributes, weak bullet structures—reduces relevance signals and can cause ranking decline. ### 3\. What are common Amazon data hygiene problems? - Misspelled keywords - Missing sizes, colors, materials - Inconsistent variant logic - Weak or generic bullets - Missing or broken images - Keyword drift over time - Duplicate content across variants - Low attribute completeness ### 4\. How do you perform an Amazon product data audit? A good audit includes: - Scraping your Amazon product data - Checking titles, bullets, and attributes for consistency - Reviewing keyword placement and density - Checking image sequences - Verifying variant alignment - Comparing your metadata with top sellers - Running LLM‑based audits for deeper semantic issues ### 5\. How do you scrape Amazon product data for catalog audits? Use automated scraping tools, scripts, or APIs to collect data such as: - Title and bullet texts - Attributes and specifications - Pricing and discounts - Variant structures - A+ content - Image URLs and alt text This structured dataset becomes the foundation for LLM audits and competitive analysis. ### 6\. How do LLMs help improve Amazon data hygiene? LLMs like GPT‑4, Claude, and Gemini identify: - Typos - Missing attributes - Tone inconsistencies They produce actionable recommendations that strengthen listing quality. ### 7\. How often should Amazon listings be audited? High‑performing brands audit weekly or bi‑weekly. Categories with high competition or frequent changes may require daily monitoring to catch drift and competitor adjustments. ### 8\. How does poor data hygiene affect conversions? Weak bullets, missing details, inconsistent images, and unclear benefits reduce shopper confidence—leading to lower conversion rates, higher CPC, and decreased Buy Box wins. ### 9\. Can scraping Amazon product data help identify competitor trends? Absolutely. Scraping competitor listings reveals: - Title patterns - Keyword strategies - Image frameworks - Benefit order - Variant structures - Pricing logic - Catalog update frequency These patterns help you refine your own listings. ### 10\. How do I get started improving Amazon data hygiene? Begin by scraping your catalog, running an LLM audit, comparing metadata with top sellers, and creating a weekly improvement workflow. ### Hidden Ecommerce Profit Killers and How to Fix Them URL: https://www.blog.datahut.co/post/the-invisible-e-commerce-profit-killers-spot-the-flaws-your-standard-audits-fail-to-catch/ Last updated: 2026-09-07T09:43:53.000Z If we run an e-commerce business, we already know this: our website changes constantly. Products get added, removed, renamed, moved, and repriced. Developers ship updates. Merchandisers tweak content. Apps and integrations act unpredictably. And somewhere in all this movement, things quietly break. The scary part? Not in Google Search Console. Not in our SEO audits. Not in our automated QA checks. Not even in our analytics. The problems often worsen slowly. Teams do not notice until the bad pattern has repeated enough. This can hurt revenue or distort important product data. As founders, we like to believe we have a solid handle on our storefront. But once our catalog crosses a few hundred SKUs, reality hits fast. Complexity doesn’t grow; it multiplies. Every new product, integration, or content change becomes a potential failure point. Pages slip out of structure, templates break, metadata disappears, and variants drift. Suddenly, we’re managing a system that changes faster than our team can track. By the time we spot a problem, the damage has usually already touched revenue. This is where [web crawling](https://www.datahut.co/?ref=blog.datahut.co) becomes vital. This is exactly why modern e-commerce teams are starting to treat crawlers not as technical tools, but as revenue protection systems. ## Imagine We’re Running a 25,000-SKU Store This is where the real problems begin. Someone on our team removes a product, but a dozen pages still link to it. Our CDN changes a path, and 400 images go missing. A merchandiser updates a category but forgets the pagination structure. [Developers push a new build, and suddenly 80 URLs redirect through a 3-step chain.](https://www.conductor.com/academy/redirects/faq/redirect-chains/?ref=blog.datahut.co). A supplier feeds updates, and half the variants lose their descriptions. None of these triggers an alert. None of this gets flagged. But all of it affects sales. And that’s where a crawler becomes our most underrated ally. ## A Crawler Is Basically Our Most Reliable Intern (Who Never Sleeps) When we explain crawlers to other founders, we describe them like this: It moves through our store the same way a buyer would: - Browsing categories - Opening filters - Checking variants - Scrolling through recommendations - Navigating pagination It does this consistently and completely. It does not miss anything because web crawling forces it to check everything, as long as we give it superpowers by customizing it to fit our business logic. ## What a Crawler Finds in a Real E-Commerce Store These aren’t hypothetical issues; they’re the ones we see in the wild every week. ### 1\. Broken Product URLs Old links. Products that are discontinued. Redirects are missing. Soft 404s disguised as normal pages. These issues slip in quietly, but their impact is anything but small. When shoppers land on a product page that doesn’t work or, worse, looks like it works but leads nowhere, they lose trust instantly. This is the silent conversion killer, and it’s a well-documented e-commerce issue. Baymard Institute's UX research shows that unexpected error messages and poor validation flows cause many people to leave during checkout. See their breakdown here: [Baymard’s UX research on inline form validation](https://baymard.com/blog/inline-form-validation?ref=blog.datahut.co). ### 2\. Missing Product Images CDN restructuring? Bad upload? Botched migration? It takes only one of these behind-the-scenes issues to create a cascading failure across our storefront. Suddenly, 200+ PDPs load with broken thumbnails, missing hero images, or blank galleries—instantly making our products look unreliable or low-quality. Google warns that missing images hurt search visibility. They also affect whether our products can appear in Shopping results. See their [product image guidelines](https://developers.google.com/search/docs/appearance/structured-data/product?ref=blog.datahut.co#image-guidelines). For more on how dynamic image loading breaks, we covered this in our guide on [scraping dynamic websites with Playwright](https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/). ### 3\. Empty Product Descriptions Feeds break. Merchandisers forget things. Variants don’t inherit content. When any of those happen, we end up with product pages that look unfinished, inconsistent, or completely empty. This isn’t just a minor merchandising slip; it’s a direct hit to both visibility and revenue. Search engines rely heavily on descriptive, attribute-rich content to understand what a product is, who it’s for, and when to surface it. When descriptions vanish or never get populated, Google treats those pages as low-quality or irrelevant. Customers react the same way: they bounce, lose trust, and rarely convert. A simple content gap silently turns into lost traffic, weaker rankings, and fewer sales. Empty descriptions cost us both SEO and conversions. Google's guide says descriptive, attribute-rich content is important for finding products. See Google Product Content Guidelines. ### 4\. Orphaned Products This is a big one. Products exist, but nothing links to them. Google doesn’t find them. Customers don’t find them. Our revenue never sees them. It’s not because our products aren’t good; it’s because the pages that should showcase them are practically invisible. When a page has no internal links pointing to it, search engines can’t properly crawl or index it. That means Google can’t understand its relevance, can’t assign it value, and ultimately can’t rank it for the queries our customers are actively searching. On the user side, a page that isn't connected to our navigation or product clusters becomes a dead end, buried deep inside our site. The result? Lost visibility, lost sessions, and lost sales, without us ever realizing that internal linking was the silent culprit. Orphaned pages are a common e-commerce SEO problem. Aleyda Solis highlights this in her resources on internal linking at [LearningSEO.io](http://learningseo.io/?ref=blog.datahut.co). This ties directly into what’s explained in the [e-commerce technical SEO checklist](https://technicalseo.com/insights/blog/technical-seo-for-ecommerce-checklist-2022/?ref=blog.datahut.co). ### 5\. Broken Categories A category looks full inside the CMS but shows up empty on the live site. Pagination stops working. Filters load dead pages or return zero results even when products exist. These issues slip through easily because category logic is one of the most fragile parts of an e-commerce setup. Shopify’s own documentation warns that faceted navigation, filters, and collection rules can break with even small theme edits, bulk imports, or app conflicts. When this happens, customers think we’re out of stock or don’t carry what they need, so they leave. Broken categories disrupt discovery, reduce product visibility, and quietly drain revenue long before anyone notices something is wrong. Reference: [Shopify’s guide to collections and navigation](https://help.shopify.com/en/manual/online-store/menus-and-links?ref=blog.datahut.co). ### 6\. Slow or Heavy Pages We’d be surprised how often this happens. A single oversized image, extra app script, or small theme tweak can quietly slow our storefront. Shopify’s guidelines show even slight speed loss hurts conversions. Google’s Core Web Vitals highlight the same issue: slow product pages push mobile shoppers to bounce, hesitate, or abandon checkout, turning small delays into real revenue leaks for brands. Shopify’s own performance guidelines confirm how much lost speed equals lost conversions. Even Google’s Core Web Vitals data shows that slow product pages disproportionately hurt mobile checkouts. Reference: [Google’s Core Web Vitals overview](https://web.dev/vitals/?ref=blog.datahut.co). ### 7\. Duplicate Pages Common in stores with variants, multiple tagging systems, or inconsistent canonicals. [Google warns that duplicate content on similar product versions can split ranking signals.](https://developers.google.com/search/blog/2008/09/demystifying-duplicate-content-penalty?ref=blog.datahut.co) This can reduce our overall visibility. When many versions of the same product exist, like different colors, sizes, or small changes, Google may have trouble deciding which page to rank first. Instead of combining authority, the signals spread out over several almost identical URLs. Google's own documents strongly recommend using proper canonicalization. [Google's guide on canonicalization](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls?ref=blog.datahut.co) explains how a clear canonical URL helps search engines find the preferred version. It also helps keep ranking strength and stops our product pages from competing with each other. ### 8\. Redirect Chains No customer wants to jump through a 301 → 302 → final URL just to see a sneaker. Ahrefs’ own research on 301 redirects shows how redirect chains slow down crawl efficiency and dilute link equity: [Ahrefs’ guide to 301 redirects](https://ahrefs.com/blog/301-redirects/?ref=blog.datahut.co). ### 9\. Pricing Inconsistencies Across Variants This one hurts the most. On the surface, it looks like a small glitch, but it’s the kind of mistake that silently erodes trust and tanks conversions. Variant A shows a price of $49\. Variant B suddenly jumps to $59 for no clear reason. Then Variant C drops right back to $49 again, as if nothing happened. To a shopper, this feels inconsistent, confusing, and even suspicious. To a retailer, it’s a hidden revenue leak waiting to happen. No single team checks variant prices for consistency. Because of this, mismatched variant prices often slip through. Customers notice these mistakes right away. We’ve seen this firsthand during almost all large retail data analyses. Even Shopify warns merchants that inconsistent variant pricing confuses search engines and increases bounce rate: [Shopify Product Variants Guide](https://help.shopify.com/en/manual/products/variants?ref=blog.datahut.co). This is a massive trust killer—and web crawling catches it instantly. ## A Quick Anecdote From the Field (80,000+ Auto Parts) A while ago, we worked with a fast-growing auto parts retailer—a team of nearly 20 people managing a massive catalog of more than 80,000 SKUs. If we’ve ever worked in the auto-parts ecosystem, we know how messy and complex these catalogs can get. Each product isn’t just a standalone item; it comes with compatibility charts, year/make/model combinations, variant fitment differences, supplier-specific data feeds, and dozens of granular product attributes. Managing all this information manually becomes overwhelming very quickly, especially when the catalog keeps expanding and new suppliers push frequent updates. On the surface, the website looked great. But once we ran a deep crawl, the real picture emerged: - Thousands of old URLs still linked from internal pages - Discontinued parts that showed up in some categories but not others - Images missing on certain fitment variants - Redirected URLs buried three levels deep - Category filters that returned empty results The team was shocked. Until that moment, everyone had assumed someone else was checking these things. The merchandisers believed the development team was monitoring all the technical issues. The developers assumed the marketing team would catch anything broken on the customer-facing side. Marketing, on the other hand, thought the SEO tools would automatically flag missing pages, broken links, and product-page failures. In reality, no one was tracking the full picture. Each team only saw their own slice of the workflow, and the gaps between those slices allowed critical problems to slip through unnoticed for months. But with 80,000+ products and dozens of hands touching the catalog weekly, the truth was simple: Nobody had full visibility. Once we deployed a custom web crawler for them, issues that had been invisible for years surfaced within 48 hours. [And fixing those problems didn’t just improve SEO — it immediately reduced customer complaints and boosted conversions.](https://www.irjmets.com/uploadedfiles/paper/issue%5F2%5Ffebruary%5F2025/68198/final/fin%5Firjmets1740666099.pdf?ref=blog.datahut.co) That’s when it really hit us again: at scale, crawling stops being a technical task and becomes a business necessity. ## A Real Crawl Report (These Numbers Hurt) Let’s take a typical 8,000-SKU store. A typical monthly crawl uncovers: - 380 broken product URLs - 210 missing or broken images - 80 empty or incomplete descriptions - 29 slow pages - 14 orphaned products That’s 713 ways we are losing money without knowing. ## How a Crawler Actually Navigates Our Store Think of the crawler like an ultra-patient shopper: - Starts at the homepage - Goes category by category - Follows every link - Opens all variants - Scrolls dynamic sections - Captures everything it sees No complexity, no jargon—a web crawler behaves exactly like a customer. ## Crawler vs Scraper (Founder Edition) Here’s the simplest way we explain it to founders: A crawler finds the problems. A scraper describes the problems. The crawler maps our entire site, detects what’s broken, and shows us exactly where issues live—missing links, slow pages, empty descriptions, orphaned URLs. The scraper then digs deeper, pulling structured data that explains why those issues exist. Together, through structured web crawling, they give us visibility. Crawler says: Scraper says: We need both. ## Why This Matters (More Than We Think) These “small” issues add up. - Broken links destroy trust. - Missing images kill conversions. - Slow pages hurt rankings. - Orphaned products make inventory harder to track and surface. Crawlers automatically keep our storefront healthy. For more technical depth, our guide on [scaling web scraping from prototype to production](https://www.blog.datahut.co/post/scaling-web-scraping-from-prototype-to-production-challenges-explained/) goes into why monitoring matters as much as data collection. ## Why SEO Tools Don’t Catch These Problems This is something we learn the hard way. SEO tools: - Don’t apply filters - Don’t click variant selectors - Don’t scroll dynamic carousels - Rely too heavily on sitemaps - Don’t run hourly checks - Don’t understand our business logic They aren’t wrong; they’re just not built for e-commerce. Even Google’s own [e-commerce search best practices](https://developers.google.com/search/docs/specialty/ecommerce?ref=blog.datahut.co) emphasize clean linking structures most stores fail to maintain. ## Why We Need a Custom Crawler (Not a Standard Tool) Because our store is built our way. Generic tools can’t understand: - Our catalog structure - Our product relationships - Our variant logic - Our dynamic UI elements - Our staging environments - Our frequency of change A custom web crawler adapts to our rules—not the other way around. ## How Datahut Helps Most brands don’t have the time, infrastructure, or engineering depth to build this in-house. That’s where Datahut comes in. We build: - Fully custom crawlers - Designed for our business logic - Capable of crawling JS-rich stores - With daily/hourly monitoring - Plugged directly into our workflows - With alerting, integrations, dashboards We don’t manage proxies, queues, retries, or bot detection. All of that complexity disappears. Instead of spending hours juggling proxy pools, rotating IPs, handling region-based blocks, solving CAPTCHAs, or debugging why a site suddenly started throttling our requests, Datahut absorbs that entire operational burden. We don’t have to worry about concurrency limits, crawler crashes, queue failures, or whether our pipeline will scale when we add thousands of new URLs. Our system handles the invisible plumbing—network management, anti-bot mitigation, stabilizing request flows, and maintaining uptime—so our team never has to touch the messy parts. We simply get reliable, structured, ready-to-use data delivered exactly when we need it. If we want a crawling system built for our store—not for a generic SEO checklist—we can build the entire pipeline end-to-end. Get in touch with us - [Datahut](https://www.datahut.co/?ref=blog.datahut.co) ## What’s Next Now that we’ve seen how crawlers actually impact revenue, the next step is understanding how to build a crawler that works reliably. In Blog 2, we’ll cover: - How to build our own crawler in Python - How to avoid getting blocked - How to scale to thousands of URLs Stay tuned to [datahut](https://www.blog.datahut.co/) blog to read the upcoming blogs! ## FAQ SECTION 1\. What are invisible e-commerce profit killers? Invisible profit killers are small, often unnoticed issues—like slow product pages, hidden redirect chains, duplicate variations, or bad tracking setups—that quietly drain conversions and revenue. 2\. Why don’t standard e-commerce audits catch these issues? Traditional audits focus on surface-level SEO or UX checks. Invisible issues worsen slowly and only become clear when data, rankings, or revenue drop significantly. 3\. How can redirect chains affect my e-commerce revenue? Redirect chains increase page load time, weaken ranking signals, and disrupt user flow—leading to higher bounce rates and lower conversions. 4\. What types of duplicate content harm e-commerce stores? Duplicate product versions (sizes, colors, seasonal variants) and poorly managed faceted URLs can split ranking signals and confuse search engines. 5\. How can data tracking flaws hurt product decisions? Incorrect or inconsistent tracking leads to misleading performance metrics. This can result in wrong pricing decisions, poor inventory planning, and wasted ad spend. ### The Biggest GDPR Fines of 2025: What Got Companies in Trouble URL: https://www.blog.datahut.co/post/top-10-gdpr-fines-in-2025-a-data-driven-analysis/ Last updated: 2026-09-07T09:43:56.000Z ### Introduction : GDPR Fines of 2025 Yes, it’s over — the era of unchecked data collection, silent tracking, and unaccountable digital practices. The General Data Protection Regulation (GDPR) ended it for good, redefining how organizations collect, process, and protect the personal data of European Union citizens. A decade ago, user information was traded, tracked, and monetized with little scrutiny; privacy was an afterthought, not a business priority. That changed in 2018 with the enforcement of GDPR — now recognized as the world’s most comprehensive and consequential data privacy framework. There’s a saying in tech — “what you don’t track, you can’t improve.” But what happens when you track too much? A company eager to “understand its users” collects every click and scroll — the intent was insight, the outcome was investigation. GDPR reframed that idea, shifting focus from what data can do to what it should do. And if you’re still wondering what exactly the General Data Protection Regulation (GDPR) is — it’s the European Union’s landmark privacy law giving individuals control over their personal data while holding organizations accountable for how it’s collected, used, and shared.For a deeper understanding, explore[ ](https://www.blog.datahut.co/post/gdpr-compliance-what-is-it-and-how-your-must-prepare-for-it/)[DataHut’s in-depth guide on GDPR compliance and how to prepare for it](https://www.blog.datahut.co/post/gdpr-compliance-what-is-it-and-how-your-must-prepare-for-it/). Since GDPR enforcement began, one truth stands firm — data protection isn’t optional. Record fines show that compliance drives trust and resilience. To support ethical data collection, [DataHut’s guide to transparent web scraping](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/)[ ](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/)offers a clear path to compliance. This article reviews the top10 GDPR fines (2021–2024) and the lessons shaping data protection in 2025. 1. Why Was Meta (Facebook) Fined €1.2 Billion Under GDPR in 2023? In 2023,[ Meta was issued a €1.2 billion fine by the Irish Data Protection Commission (DPC)](https://www.edpb.europa.eu/news/news/2023/12-billion-euro-fine-facebook-result-edpb-binding-decision%5Fen?ref=blog.datahut.co). The DPC found that users’ personal data from Facebook was transferred to external servers without adequate safeguards, violating GDPR requirements on international data transfers. The record-breaking penalty followed a binding decision by the European Data Protection Board (EDPB). ➡️ GDPR Articles: 44–49 — Rules governing cross-border data transfers. Key Takeaway: Cross-border data transfers remain one of the most complex compliance challenges under GDPR. Organizations must implement enforceable safeguards and conduct continuous risk assessments when moving personal data outside the EU. 1. What Led to Amazon’s €746 Million GDPR Fine in 2021? Following Meta’s penalty ,In 2021,[ Amazon was issued a €746 million fine by Luxembourg’s Commission Nationale pour la Protection des Données (CNPD)](https://cnpd.public.lu/en/actualites/international/2021/08/decision-amazon-2.html?utm%5Fsource=chatgpt.com). The CNPD found that customer data was processed for targeted advertising without valid consent, [violating GDPR principles of lawfulness and transparency.](https://www.sec.gov/ix?doc=/Archives/edgar/data/0001018724/000101872421000020/amzn-20210630.htm&ref=blog.datahut.co) ➡️ GDPR Articles: 5(1)(a) and 6 — Fairness, transparency, and lawful processing. Key Takeaway: Clear and informed consent is essential for lawful data use. Even slight ambiguity in advertising practices can invite regulatory penalties. 1. How Did Instagram’s Public Data Settings Result in a €405 Million Fine? Soon, regulators turned their attention to social platforms. In 2022, [Instagram was issued a €405 million fine by the Irish Data Protection Commission (DPC).](https://www.edpb.europa.eu/news/news/2022/record-fine-instagram-following-edpb-intervention%5Fen?ref=blog.datahut.co) The DPC found that minors’ contact details were exposed through public business profiles and default visibility settings, violating GDPR requirements for the protection of children’s personal data. ➡️ GDPR Article: 6(1) — Lawfulness of processing (children’s data). Key Takeaway: Child data protection demands proactive design controls. Platforms targeting or accessible to minors must enforce strict default privacy settings and minimize public exposure. 1. Why Did Meta Face a €390 Million Penalty for Ad Consent Violations? In 2023, [Meta was issued a €390 million fine by the Irish Data Protection Commission (DPC)](https://www.dataprotection.ie/en/news-media/data-protection-commission-announces-conclusion-two-inquiries-meta-ireland?ref=blog.datahut.co). The DPC found that users were required to accept personalized advertisements to access Facebook and Instagram, violating GDPR principles regarding the lawful processing of personal data. The decision followed a binding determination by the European Data Protection Board (EDPB). ➡️ GDPR Article: 6(1) — Lawfulness of processing (invalid consent). Key Takeaway: Under GDPR, access cannot depend on ad consent. Users must have real freedom to opt out without losing service functionality. 1. How Did TikTok Breach GDPR Rules on Teen Privacy? The same year, [TikTok was issued a €345 million fine by the Irish Data Protection Commission (DPC).](https://www.dataprotection.ie/en/news-media/press-releases/DPC-announces-345-million-euro-fine-of-TikTok?ref=blog.datahut.co) The DPC found that TikTok set teen user accounts to public by default, allowing anyone to view or comment on their content, violating GDPR principles of data protection by design, by default, and data minimization. ➡️ GDPR Articles: 8 and 25 — Child consent and privacy by design and default. Key Takeaway: Integrating privacy early in design prevents bigger risks later. Regulators expect protection by design, not post-launch corrections. 1. Why Was LinkedIn Fined €310 Million by the Irish DPC in 2024? As regulators deepened their focus on ad targeting,[ LinkedIn was issued a €310 million fine by the Irish Data Protection Commission (DPC).](https://www.dataprotection.ie/en/news-media/latest-news?ref=blog.datahut.co) The DPC found that LinkedIn processed user data for targeted advertising without valid consent, violating GDPR requirements for lawful processing of personal data. ➡️ GDPR Article: 6(1)(a) — Lawfulness of processing based on cookie consent. 🔗[ ](https://www.dataprotection.ie/en/news-media/latest-news?ref=blog.datahut.co)[Irish DPC Press Release – LinkedIn Ireland Fine](https://www.dataprotection.ie/en/news-media/latest-news?ref=blog.datahut.co) Key Takeaway: Personalized advertising must be built on clear user consent. Without a legitimate basis, behavioral targeting erodes trust and invites regulatory risk. 1. What GDPR Violations Led Uber to a €290 Million Fine? As scrutiny expanded to cross-border platforms in 2024, [Uber was issued a €290 million fine by the Dutch Data Protection Authority (DPA)](https://www.reuters.com/technology/cybersecurity/dutch-privacy-watchdog-fines-uber-sending-drivers-data-us-2024-08-26/?utm%5Fsource=chatgpt.com). The DPA found that [Uber transferred EU users’ personal data to U.S. servers without adequate safeguards](https://www.edpb.europa.eu/news/news/2024/dutch-sa-imposes-fine-290-million-euro-uber-because-transfers-drivers-data-us%5Fen?ref=blog.datahut.co), violating GDPR Data Sharing provisions governing international data transfers. ➡️ GDPR Articles: 44–49 — International data transfer provisions. Key Takeaway: The enforcement underscored that contractual assurances are insufficient on their own. Robust technical and procedural measures are essential to meet GDPR adequacy requirements for data transfers. 1. How Did Meta’s Data Scraping Incident Lead to a €265 Million Fine? In 2022,[ Meta was issued a €265 million fine by the Irish Data Protection Commission (DPC)](https://www.euronews.com/next/2022/11/28/meta-hit-with-265-million-fine-by-irish-regulators-for-breaking-europes-data-protection-la?ref=blog.datahut.co). The DPC found that personal data of over 500 million Facebook users was scraped and published online due to insufficient security and privacy safeguards, violating GDPR privacy policy principles of integrity and confidentiality. ➡️ GDPR Articles: 25 and 5(1)(f) — Privacy by design and integrity/confidentiality of processing. Key Takeaway: This decision underscored that weak technical controls leading to mass data exposure are considered governance failures under GDPR’s security and integrity obligations. 1. Why Was Meta Penalized €251 Million for Data Protection Failures? In 2024, [Meta was among the notable cases in ](https://www.dataprotection.ie/en/news-media/press-releases/irish-data-protection-commission-fines-meta-eu251-million?ref=blog.datahut.co)[GDPR enforcement fines](https://www.dataprotection.ie/en/news-media/press-releases/irish-data-protection-commission-fines-meta-eu251-million?ref=blog.datahut.co)[, receiving a €251 million penalty from the Irish Data Protection Commission (DPC).](https://www.dataprotection.ie/en/news-media/press-releases/irish-data-protection-commission-fines-meta-eu251-million?ref=blog.datahut.co)The DPC found that Meta failed to implement adequate data protection measures at a system level, violating GDPR requirements for data protection by design and by default. ➡️ GDPR Article: 25 — Data protection by design and default. Key Takeaway: Compliance must operate at the architectural level rather than through isolated procedures. Embedding privacy by design within systems and workflows ensures resilience and consistent adherence to GDPR standards. 1. What GDPR Transparency Failures Cost WhatsApp €225 Million? Transparency became the next frontier of GDPR enforcement,In 2021,[ WhatsApp was issued a €225 million fine by the Irish Data Protection Commission (DPC)](https://www.dataprotection.ie/en/news-media/press-releases/data-protection-commission-announces-decision-whatsapp-inquiry?ref=blog.datahut.co). The DPC found that WhatsApp failed to clearly inform users and non-users about how their data was collected and shared with Facebook, violating GDPR Privacy Controls principles of transparency and fair processing of personal data. ➡️ GDPR Articles: 5(1)(a), 12, 13, and 14 — Transparency and information obligations. Key Takeaway: Transparency remains the foundation of effective data governance. Clear, layered, and accessible privacy disclosures strengthen accountability and help maintain user confidence. ![Top 10 GDPR Fines 2025](https://www.blog.datahut.co/content/images/2026/07/img-361.png.webp) Over 80% of these record fines were issued by the Irish Data Protection Commission (DPC), highlighting Ireland’s pivotal role in EU data enforcement. ## What These Fines Teach Businesses - Transparency and Consent Are Non-Negotiable: Most GDPR violations stem from unclear consent practices or data usage. - Privacy by Design Is Critical for privacy law: Weak system-level protections can lead to multimillion-euro penalties. - Cross-Border Data Transfers Need Safeguards: International processing requires documented compliance. - Children’s Data Is Highly Protected: Default public exposure of minors’ information is a red flag. - Continuous Compliance Is Essential: Regular audits and documentation prevent long-term exposure. ### Datahut Expertise At DataHut, we help businesses turn compliance into a competitive advantage — enabling responsible, GDPR-aligned data collection through transparent web scraping solutions. Partner with our experts at[ ](https://www.datahut.co/?ref=blog.datahut.co)DataHut to keep your operations ethical, efficient, and future-ready. ### Conclusion The enforcement of GDPR demonstrates that no organization is beyond regulatory accountability. Each case highlights a distinct compliance failure — from transparency and consent to cross-border data handling. As seen across these landmark decisions, regulators continue to reinforce that user trust and ethical data practices are the foundations of sustainable digital business. Organizations can prepare better by understanding real-world enforcement trends and applying compliant methods for data collection. For instance,[ ](https://www.blog.datahut.co/post/how-to-use-web-scraping-to-track-gdpr-fines-and-enforcement-cases/)[DataHut’s post on using web scraping to track GDPR fines and enforcement cases](https://www.blog.datahut.co/post/how-to-use-web-scraping-to-track-gdpr-fines-and-enforcement-cases/) explores how businesses can leverage data responsibly while staying fully compliant. As global data protection authorities strengthen coordination, organizations must view GDPR enforcement and data protection laws not as punishment, but as an evolving framework for safeguarding data breach and data subjects’ rights and digital accountability. To Ensure your web scraping and data collection practices stay GDPR-compliant —[ ](https://www.datahut.co/?ref=blog.datahut.co)[talk to our data experts today at ](https://www.datahut.co/?ref=blog.datahut.co)DataHut. ## Frequently Asked Questions (FAQ) Q1\. What is the largest GDPR fine to date? The largest fine to date is €1.2 billion, imposed on Meta (Facebook) in 2023 by the Irish Data Protection Commission (DPC) for unlawful data transfers to the United States. Q2\. Who enforces GDPR fines? GDPR fines are enforced by national Data Protection Authorities (DPAs) such as Ireland’s DPC, Luxembourg’s CNPD, and the Dutch DPA, under the coordination of the European Data Protection Board (EDPB). Q3\. How are GDPR fines calculated? The amount of a GDPR fine depends on the severity and duration of the violation, the company’s global turnover, and the degree of cooperation with regulators. Serious violations can result in penalties of up to 4% of annual global revenue. Q4\. Are small businesses also subject to GDPR? Yes — GDPR applies to any organization processing the personal data of EU citizens, regardless of business size or geographic location. Learn more in[ ](https://www.blog.datahut.co/post/when-does-gdpr-come-into-force/)[DataHut’s post on GDPR compliance for small businesses](https://www.blog.datahut.co/post/when-does-gdpr-come-into-force/). Q5\. How can companies prepare for GDPR compliance? Organizations should conduct data audits, document processing activities, and obtain explicit consent from users before collecting or sharing data. For practical guidance, read[ ](https://www.blog.datahut.co/post/gdpr-compliance-what-is-it-and-how-your-must-prepare-for-it/)[DataHut’s GDPR compliance preparation guide](https://www.blog.datahut.co/post/gdpr-compliance-what-is-it-and-how-your-must-prepare-for-it/). ### California vs New York Condo Prices: Data Insights URL: https://www.blog.datahut.co/post/california-vs-new-york-condo-prices-2025-1-400-homes-com-data-insights-revealed/ Last updated: 2026-09-07T09:43:57.000Z Buying a home—be it a house, condo, or co-op—in California or New York is not just a choice of location, but a high-stakes financial decision. At Datahut, [we scraped](https://www.blog.datahut.co/post/is-web-scraping-legal/) over 1,400 real estate listings from [Homes](https://www.homes.com/?ref=blog.datahut.co) to analyze how these two iconic states compare specifically in the condo and apartment market. This [Exploratory Data Analysis (EDA)](https://medium.com/@akshatsharma0610/a-tour-to-eda-exploratory-data-analysis-fafef76d38a7?ref=blog.datahut.co) reveals that from median purchase price and property taxes to the value per square foot and the cost of larger units, California's housing market is significantly more expensive than New York's, revealing stark coast-to-coast differences in affordability and long-term ownership burdens At [Datahut](https://www.datahut.co/?ref=blog.datahut.co), we scraped real estate listings from Homes to analyze how these two iconic housing markets compare. From median home prices to tax burdens and square footage value, this study reveals exactly how—and why—the cost of home ownership differs coast to coast. In this Exploratory Data Analysis (EDA), we dive into two sets of real estate listings scraped from Homes, focusing on California and New York—two of the most dynamic housing markets in the United States. Each dataset contains 700 property records, offering a balanced sample for meaningful comparison. The goal of this analysis is to explore, compare, and visualize housing trends across these two states—examining factors such as price distribution, property size, number of bedrooms and bathrooms, and location-based patterns etc. By leveraging [Plotly](https://towardsdatascience.com/dynamic-eda-for-qatar-world-cup-teams-8945970f16be/?ref=blog.datahut.co), an interactive visualization library, we aim to uncover insights that highlight how the real estate landscapes of California and New York differ in affordability, property features, and overall market composition. Want full raw data for your own analysis? Click here to view all the data ([California](https://www.dropbox.com/scl/fi/o702w69zwn51wpi736o6h/cleaned-data-california.csv?rlkey=eul75ithcimzrdr9n2ei8e3kz&st=2a6bzb8n&dl=0&ref=blog.datahut.co) , [New York](https://www.dropbox.com/scl/fi/qgu552y8vybyvkel3zp80/cleaned-data-newyork.csv?rlkey=362x3i2dyvzv6b3h5wnz4q67h&st=gyvztghd&dl=0&ref=blog.datahut.co)) ## [Homes](https://www.homes.com/?ref=blog.datahut.co) Data Insights- Median Condo Prices: ## Coast-to-Coast Comparison Thinking of buying a condo in the U.S.? One of the biggest decisions for home-buyers and real estate investors is choosing the right state — especially when comparing top markets like[ ](https://capstone72.com/california-vs-new-york-real-estate-market/?utm%5Fsource=chatgpt.com)[California vs. New York](https://capstone72.com/california-vs-new-york-real-estate-market/?utm%5Fsource=chatgpt.com). Condo prices in the United States vary widely based on location, demand, property size, and market trends. In this data-driven [real estate analysis](https://www.blog.datahut.co/post/scraping-property-data/), we break down median condo prices, affordability, and investment potential between California and New York — two of the most competitive U.S. housing markets. Whether you're a first-time buyer, property investor, or market analyst, this side-by-side comparison offers actionable insights backed by real scraped data from Homes. [image\_link](https://www.dropbox.com/scl/fi/1d9kc60u3hok0tix7jq7u/median%5Fhome%5Fprice%5Fby%5Fstate1.jpeg?rlkey=7ngzwbu4kps2fwfr8z4wiazg1&st=uuzohpen&dl=0&ref=blog.datahut.co), [data\_link](https://www.dropbox.com/scl/fi/8u3aila5jy6w7pdflcode/median%5Fhome%5Fprice%5Fby%5Fstate.csv?rlkey=5rmr0qvmv9qioqjeq4j581twe&st=fn426fa4&dl=0&ref=blog.datahut.co) This bar chart displays the median condo prices for two states — California and New York. Each bar represents the average price of condos in that state based on the data collected. - The blue bar represents California. - The orange bar represents New York. - The values above each bar show the exact median condo price in dollars. ### Key Insights - California has a much higher median condo price at $1,195,000. - New York’s median condo price is about $599,800, which is almost half of California’s. - This difference highlights how location and demand play a major role in [real estate](https://www.blog.datahut.co/post/data-scraping-in-real-estate-transform-housing-industry/) pricing. - The data suggests that the California condo market is significantly more expensive — likely due to high demand, limited supply, and popular cities like Los Angeles, San Francisco, and San Diego. - In contrast, New York State’s overall condo prices are more moderate, despite New York City having its own high-cost real estate areas. California continues to rank among the most expensive U.S. real estate markets, with [condo prices](https://www.bankrate.com/real-estate/median-home-price/?utm%5Fsource=chatgpt.com) and property taxes significantly higher than in New York. While New York is still considered a high-cost state, it offers comparatively better affordability, especially for buyers prioritizing value per square foot and lower annual property tax burdens. ## Exploring Condo Price Distributions Across States This section of the analysis explores the distribution of [Condo prices](https://www.redfin.com/blog/how-to-buy-a-condo/?ref=blog.datahut.co) in both locations. ![This section of the analysis explores the distribution ofCondo pricesin both locations.](https://www.blog.datahut.co/content/images/2026/07/img-2.jpeg.webp) [image\_link](https://www.dropbox.com/scl/fi/w6wi447wirrjuj9up7pun/price%5Fdistribution%5Fhistogram1.jpeg?rlkey=e26p6dq3ay0b0wio26b3xcv08&st=fejwfjgu&dl=0&ref=blog.datahut.co), [data\_link](https://www.dropbox.com/scl/fi/ols5ug6kyh5sxaw6rrfau/price%5Fdistribution%5Ftable.csv?rlkey=hm416lsxfuf59o95yfbxq44i1&st=jzuo4cth&dl=0&ref=blog.datahut.co) This histogram compares how condo prices are distributed in California (red bars) and New York (green bars). - The horizontal axis (x-axis) shows the price of homes in dollars. - The vertical axis (y-axis) represents the number of properties that fall within each price range (count). - Both distributions are plotted on the same graph, allowing us to see where prices overlap and where they differ. This type of plot helps reveal whether most condos are clustered around a certain price range or if there are many properties in higher or lower ranges. ### Key Insights - Majority of condos in both states fall within the lower price brackets (under $5 million). This suggests that while luxury condos exist, most listings are in more affordable ranges. - New York shows a tighter cluster of prices, meaning condo prices are more concentrated within a smaller range. Most properties fall below $10 million. - California’s distribution is more spread out, indicating that while many condos are in lower price ranges, there are also a noticeable number of high-value properties extending well above $10 million. - The long “tail” seen in California’s distribution shows the presence of ultra-luxury properties, going as high as $50 million in some areas. - New York’s count is higher in the lower range, but the presence of extremely high-priced condos is less frequent compared to California. The condo price distribution clearly highlights how different the California and New York housing markets truly are. - California shows a wider price range, especially in the luxury segment, indicating strong demand for high-end condos in cities like Los Angeles, San Francisco, and San Diego. - New York displays tighter price clustering, suggesting a more stable and mature condo market with fewer extreme price spikes outside premium NYC districts. These findings provide valuable real estate insights for buyers, investors, and analysts aiming to understand state-wise property trends and identify opportunities in the U.S. condo market. ## Understanding Condo Value Through Square Footage In this section, we examine how the size of condos (measured in square footage) influences the median price across two of the most competitive real estate markets in the U.S. ![In this section, we examine how the size of condos (measured in square footage) influences the median price across two of...](https://www.blog.datahut.co/content/images/2026/07/img-3.jpeg.webp) [image\_link](https://www.dropbox.com/scl/fi/5dab2437cbch84q4gv5c9/Top-5-Common-Sqft-vs-Median-Price1.jpeg?rlkey=6h93d7258zsgrrb8h779kcawt&st=u2dchpgp&dl=0&ref=blog.datahut.co)[,](https://www.dropbox.com/scl/fi/5dab2437cbch84q4gv5c9/Top-5-Common-Sqft-vs-Median-Price1.jpeg?rlkey=6h93d7258zsgrrb8h779kcawt&st=u2dchpgp&dl=0&ref=blog.datahut.co) [data\_link](https://www.dropbox.com/scl/fi/xm261kmlc7wvnwewubdhg/Top-5-Common-Sqft-vs-Median-Price%5Ftable.csv?rlkey=eb81013gj522czhlijd5fl9d5&st=d2zryrrq&dl=0&ref=blog.datahut.co) The bar chart compares the median condo prices for the top five most common square footage ranges (Sqft) in California and New York. - The x-axis shows the square footage (condo size). - The y-axis represents the median condo price in dollars. - Each pair of bars compares the same condo size between the two states. - Teal bars represent California, and purple bars represent New York. This visualization helps us understand whether larger condos always lead to higher prices — and how those price patterns differ between the two states. ### Key Insights - For most condo sizes, California condos are priced higher than condos in New York, even when the square footage is similar. - Around 780–800 sqft, both states show high prices, but California’s prices are still slightly higher — reflecting the state’s overall higher cost of housing. - The largest gap appears around 670–750 sqft, where California’s median price exceeds New York’s by a noticeable margin. - Interestingly, New York shows some high-priced outliers (for example, at \~790 sqft), which might represent premium urban apartments, likely influenced by location (e.g., NYC metro areas). - Square footage alone doesn’t fully explain condo price differences — factors like neighborhood, demand, and amenities play major roles in these variations. The California condo market consistently maintains higher median prices across nearly all common condo sizes. This reflects the state’s strong housing demand, limited property supply, and premium land values in cities such as Los Angeles, San Francisco, and San Diego. In contrast, New York condo prices show more variation — generally lower overall but with occasional spikes in high-value urban listings, particularly in Manhattan and nearby metro areas. Interestingly, the relationship between condo size and price isn’t always linear. Smaller condos in high-demand neighborhoods can cost more per square foot than larger units in less competitive areas. Overall, California demonstrates a stronger price premium per square foot, driven by factors such as limited housing supply, location desirability, and intense market competition in major real estate hubs. ## How Much Do Homeowners Pay? Which State Has Higher Property Tax This analysis focuses on understanding the median estimated annual property taxes in 2 states. The goal of this analysis is to compare how property taxes differ between these two states, giving potential homeowners or investors an idea of the cost of maintaining a home in each location. ![This analysis focuses on understanding the median estimated annual property taxes in 2 states. The goal of this analysis is...](https://www.blog.datahut.co/content/images/2026/07/img-4.jpeg.webp) [image\_link](https://www.dropbox.com/scl/fi/j1x8nova255g3jl77gfx5/median%5Fannual%5Ftaxes%5Fby%5Fstate1.jpeg?rlkey=95xzghsbi4xate8qxssnff3c0&st=elpq97nb&dl=0&ref=blog.datahut.co), [data\_link](https://www.dropbox.com/scl/fi/4wdk2mai69cyltvglkjvv/median%5Fannual%5Ftaxes%5Fby%5Fstate.csv?rlkey=mr8n9egg2b4z5u1wxoznyu1s1&st=v2s633pz&dl=0&ref=blog.datahut.co) This horizontal bar chart shows the median estimated annual property taxes for California and New York. Each bar represents the midpoint (median) value of property taxes estimated across the available condos in that state. What the Chart Shows - The x-axis represents the Median Annual Taxes ($) — the average midpoint of yearly taxes homeowners are expected to pay. - The y-axis lists the States being compared — California and New York. - Each colored bar represents the tax amount for that state, and the exact value is displayed inside the bar. ### Key Insights - California has a significantly higher median annual tax of $10,746. - New York, on the other hand, has a median annual tax of $5,352 — roughly half of California’s median tax value. - This large difference indicates that property ownership costs are higher in California, even though both states are known for high living expenses. - Factors that may contribute to this difference include property valuation methods, state tax policies, and regional housing market trends. From this comparison, it’s clear that California homeowners pay much more in property taxes than those in New York, on average. For buyers considering relocation or investment, this kind of insight helps in budget planning and financial decision-making. ## Do Larger Condos Mean Higher Monthly Payments? CA vs NY Cost Analysis This analysis explores how estimated monthly condo payments vary with the number of bedrooms. The goal is to understand how affordability shifts as condo size increases in these two states, helping buyers or renters make more informed housing budget decisions. ![This analysis explores how estimated monthly condo payments vary with the number of bedrooms. The goal is to understand how...](https://www.blog.datahut.co/content/images/2026/07/img-5.jpeg.webp) [image\_link](https://www.dropbox.com/scl/fi/5lzo2ja8uljixrdoncu2s/Top-Estimated-Payment-vs-Number-of-Beds-by-State1.jpeg?rlkey=bvn7t3oqsk2zir69tismoov8q&st=qw2wdovs&dl=0&ref=blog.datahut.co), [data\_link](https://www.dropbox.com/scl/fi/7a7pecoj7a1cujihflifm/Top-Estimated-Payment-vs-Number-of-Beds-by-State.csv?rlkey=yttvopcu12g4kvg3z1aqd5tyb&st=71x0pe17&dl=0&ref=blog.datahut.co) This bar chart compares the top estimated monthly payments for homes with different numbers of bedrooms. Each bar represents the highest estimated payment recorded for that specific number of bedrooms. What the Chart Shows - The x-axis represents the number of bedrooms (from 1 to 6). - The y-axis represents the estimated payment per month ($) — the expected top monthly cost for properties with the given number of bedrooms. - The two colored bars represent California (green) and New York (orange). ### Key Insights - For 1–2 bedroom homes, both states have relatively lower estimated payments, but California consistently shows slightly higher costs. - From 3 bedrooms onward, the payment gap widens significantly, with California’s payments rising much faster than New York’s. - The highest estimated payment is for 6-bedroom homes in California ($93,087/month), compared to New York’s $51,138/month — showing almost a double difference. - This pattern highlights how larger homes in California tend to be far more expensive, reflecting the state’s higher real estate prices and luxury housing demand. - Overall, as the number of bedrooms increases, both states see higher costs, but California’s growth rate is much steeper, suggesting premium property markets and higher valuation per square foot. This analysis highlights a clear trend in the U.S. housing market — the cost of living and home-ownership expenses rise much faster in the California housing market compared to New York real estate as property size increases. For home-buyers and real estate investors looking into larger condos or multi-bedroom homes, California property prices demand a significantly higher budget. In contrast, New York housing prices show a more gradual and moderate increase in overall costs as the number of bedrooms expands, offering relatively better value for larger living spaces. ## How Condo Age Impacts Price per Square Foot This analysis explores the relationship between the year a home was built and its price per square foot. The purpose of this analysis is to understand how property age affects pricing trends in each state, and whether newer homes are priced higher or lower per square foot compared to older ones. ![This analysis explores the relationship between the year a home was built and its price per square foot. The purpose of this...](https://www.blog.datahut.co/content/images/2026/07/img-6.jpeg.webp) [image\_link](https://www.dropbox.com/scl/fi/btm2svuj4cdn11jw4mjtd/price%5Fper%5Fsqft%5Fvs%5Fyear%5Fbuilt1.jpeg?rlkey=u5dgo577mvya7mrk8toyh0vg9&st=lltxyngj&dl=0&ref=blog.datahut.co), [data\_link](https://www.dropbox.com/scl/fi/m99cs0uj26saa3hrkxi0l/price%5Fper%5Fsqft%5Fvs%5Fyear%5Fbuilt.csv?rlkey=fhqmeh5ltcfrsgfuxhvzgilyc&st=zkrk625q&dl=0&ref=blog.datahut.co) This scatter plot visualizes how home prices per square foot vary depending on when the homes were built. Each dot represents a property listing from California or New York. What the Chart Shows - The x-axis represents the Year Built, showing how old or new each property is. - The y-axis represents the Price per Square Foot ($) — how much one square foot of property costs. - Each colored dot shows a property in either California (red) or New York (green). ### Key Insights - California properties generally have higher prices per square foot than those in New York, regardless of the construction year. - Many of the older homes (built before 1950) show a wide variation in prices, which could be due to location, renovation quality, or historic value. - Recently built homes (after 2000) tend to have higher prices per square foot, especially in California — suggesting that modern constructions command a premium. - In New York, most homes cluster between $1,000–$3,000 per sqft, indicating more price stability across different build years. - The scattered high-price points in California show that certain luxury or high-demand areas push property values sharply upward. This real estate analysis shows that newer condos cost more per square foot, driven by demand for modern amenities and updated infrastructure. However, California condos remain significantly pricier than New York, highlighting the impact of location-driven market value. Pricing is shaped by more than home age — key factors include state-wise housing demand, urban development, and local market trends. These insights help home buyers, real estate investors, and market analysts compare long-term value and make data-informed decisions in the [U.S. condo market](https://www.redfin.com/us-housing-market?ref=blog.datahut.co). Want similar data-driven real estate insights? Get in touch with [Datahut](https://www.datahut.co/?ref=blog.datahut.co) to extract, analyze, and visualize property data from any real estate platform — effortlessly. ### FAQs 1\. How were the California and New York condo prices analyzed in this study? The condo prices were analyzed using data scraped from [Homes](https://www.homes.com/?ref=blog.datahut.co) listings. This included property prices, square footage, location, and amenities. Tools like Datahut’s real estate web scraping solutions help automate this process to generate reliable and large-scale datasets. 2\. Why do condo prices differ so much between California and New York? Price variations are influenced by several factors like demand, location desirability, property taxes, and income levels. California’s coastal cities and New York’s Manhattan region both attract premium buyers, but local regulations and housing supply create distinct pricing patterns. 3\. Can real estate businesses use web scraping to track price trends like these?Absolutely. With ethical and compliant real estate data scraping, businesses can monitor pricing fluctuations, inventory changes, and buyer demand. Datahut offers scalable scraping solutions that transform raw listing data into actionable market insights. 4\. Is the data from [Homes](https://www.homes.com/?ref=blog.datahut.co) reliable for market analysis? Yes, [Homes](https://www.homes.com/?ref=blog.datahut.co) is a reputable source for property listings. However, to ensure accurate insights, it’s crucial to clean, validate, and analyze the scraped data — a process that Datahut specializes in through its end-to-end data extraction and analysis services. 5\. How can I get similar data for my real estate business? You can request a custom data scraping project with [Datahut](https://www.datahut.co/?ref=blog.datahut.co). Whether you need condo prices, rental trends, or agent data, Datahut can collect and structure real-time property information from major listing sites to support data-driven decision-making. ### Extracting Product Data from AllMachines Easily URL: https://www.blog.datahut.co/post/how-to-scrape-data-from-allmachines/ Last updated: 2026-09-07T09:43:59.000Z Did you ever think about how comparison websites get the prices and details of the same product from so many online stores? There’s a pretty little trick called web scraping that does it. You can think of web scraping as almost sending a tiny robot to various websites to collect similar information and extract titles, prices and descriptions. Over the years, that robot has gotten very intelligent! The advent of new technologies like headless browsers (browsers that run in the background - they don’t have a window), and workarounds to [avoid being blocked](https://www.blog.datahut.co/post/how-to-maintain-anonymity-when-web-scraping-at-scale-expert-tips/) by sites to prevent scraping has made the techniques for [web scraping ](https://www.blog.datahut.co/post/python-web-scraping-tutorial/)more powerful and reliable. In this blog I will take you through a project where we use web scraping to scrape data from a site called AllMachines - a site that lists farming equipment. Why farming tools? Well, just like comparing phones, or comparing laptops helps a person make a buying decision or understand the market, comparing tractors and other machines should help in that process too. We divided the project into two simple sections: 1)Collecting links to products: First, we collected all of the web addresses (URLs) for the product pages on the site. 2)Collecting details for products: Then we visited each of those product pages and collected important information such as product name, features, specifications, etc... This two-part process makes it easier to manage and is handy if we need to modify or extend one of the sections in the future. Now, let us move into the code and see where all these sections fit together to create our smart farming equipment data collector! ## Links Collection This web scraping project was designed to scrape product listings from a website called AllMachines, which features farming equipment. The aim is straightforward: scrape through categories like tractors, harvesters, balers and other farm machinery, and scrape the product links into a database. Once the scraper locates these links, it saves them in a lightweight database called SQLite - almost like a small notebook to save information to use on the product links we have scraped. To make this process run smoothly, we utilized two powerful pieces of technology: [Playwrigh](https://playwright.dev/?ref=blog.datahut.co)t and [BeautifulSoup.](https://beautiful-soup-4.readthedocs.io/en/latest/?ref=blog.datahut.co) The scraper is sophisticated enough to handle some tricky web features, for example infinite scrolling (where more products appear below the current page as you scroll down) as well as having built-in error handling and logging so in the event of an error we know how to ascertain the problem and fix it. ### Import Section ``` import time import sqlite3 import logging import traceback from bs4 import BeautifulSoup from playwright.sync_api import sync_playwright ``` Before we dive into the actual scraping, let’s go over the tools (or libraries) we will be using in the code. You can think of these as applications or tools, which help us accomplish specific tasks. Here’s the list of the libraries that we have imported: 1. time: This library helps us pause between actions. You can think of it as a short break so that we don’t overwhelm the website with too many requests at once. 2. sqlite3: This library will provide us with a simple database, where we can save the product links and other information we scrape. You can think of it as a saving note for something we can look back on. 3. logging: This library will keep track of what happens in the course of the script. It is extremely helpful for seeing if things are functioning correctly or if the web-scraper has stopped running. 4. traceback: If our code generates an error, this library will show us exactly where the code broke down, so we can diagnose the issue more effectively. Next, here are the real stars of the show: 1. BeautifulSoup: This library allows us to read it and extract the information from the code of a webpage (HTML). This is like scanning a recipe and pulling off just the list of ingredients. 2. Playwright: This one is really neat. It allows our code to interact with sites like a human would - clicking buttons, scrolling on pages, closing popups. This is most useful when dealing with a lot of Javascript from a webpage (which is responsible for features we take for granted, like sliders, infinite scrolling, or even pop-up menus). ### Logging Setup ``` # Configure logging to write logs to a file with a specific format logging.basicConfig( filename='scraper.log', filemode='w', # Overwrite the log file each time the script runs level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s' ) ``` Logging is pretty much your code's diary, where you note what happens while it runs. It's really important because if something goes wrong, your log can provide context for what happened and when. In our scraper, we have implemented everything so that all those notes (or "logs") are saved in a file called scraper.log. Here are a couple of key things about our implementation: - filemode='w': This means that every time we run the code, we start a new log file. We don't come back to an enormous log file filled with outdated messages that we don't need any longer. - Logging level is INFO: We are going to log all relevant info, and leave out tiny details that we don't need to know (unless it's deep debugging). We intentionally structured the log messages in an organised fashion. Every log message has: - The timestamp of when the event occurred - The event type (such as info/warning/error) - The message describing what happened This way of structuring the log helps you simply look through the logs later, find a problem, or just see how it went. By doing this from the start, we now have a scraper that almost tells its own story, which isn't any more helpful than it being there when we improve or update the scraper or fix bugs. ### Database Functions ``` def init_db(db_name="allmachines_products.db"): """ Initialize the SQLite database and create the `product_links` table if it doesn't exist. This function connects to the SQLite database specified by `db_name`. If the database file does not exist, it will be created. The function also ensures that the `product_links` table is created with the required schema to store product categories and URLs. Args: db_name (str): The name of the SQLite database file. Defaults to "allmachines_products.db". Returns: sqlite3.Connection: A connection object to interact with the SQLite database. Logs: - INFO: When the database is successfully initialized and the table is created. - ERROR: If there is an error during database initialization. Raises: Exception: If there is an error during database initialization, it logs the error and raises it. """ try: conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_links ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT UNIQUE ) """) conn.commit() logging.info("Initialized database and created table if not exists.") return conn except Exception as e: logging.error(f"Database initialization error: {e}") logging.debug(traceback.format_exc()) raise ``` The init\_db function is like preparing our storage box before we go and find anything. And in this particular case, the "box" refers to a SQLite database that we are going to use in order to store the product links that we find from scraping. What this function does is the following: First, it creates a connection to the database file. If this file does not existed, it will create one for us. This file is where we will store the data we collect. Second, it creates a table named product\_links. A table, you could think of it like a spreadsheet, and it has rows and columns, which is a place to store our information. In this product\_links table, there are three columns: 1. id: A number that increments each time we add a new entry (like row number in Excel). 2. category: Will tell us what type of equipment it is, like "tractor" or "harvester". 3. url: A link to the actual product on the web. We have also added a UNIQUE constraint on the URL column to prevent saving the same product URL more than once. If the scraper finds the same URL a second time, it will not save a record again. Do we see how this is useful? We added error handling here as well. In case of some failure regarding the database setup, the function will log an error so it's easier to find out what was wrong. By putting the setup inside of a separate function, we created a cleaner method of setting up our schemas, and organized our code. If you think about it, it's a bit like organizing your tools in a labeled box before you start a project; it makes everything neat and orderly. ``` def save_links_to_db(conn, category, links): """ Save the extracted product links to the SQLite database. This function inserts the extracted product links into the `product_links` table. Duplicate entries are ignored using the `INSERT OR IGNORE` SQL statement. Args: conn (sqlite3.Connection): The SQLite database connection object. category (str): The category of the products (e.g., "tractors", "combine-harvesters"). links (set): A set of product links to be saved in the database. Returns: None Logs: - INFO: The number of links saved for the given category. - ERROR: If there is an error during the database operation. """ try: cursor = conn.cursor() for link in links: # Insert links into the database, ignoring duplicates cursor.execute("INSERT OR IGNORE INTO product_links (category, url) VALUES (?, ?)", (category, link)) conn.commit() logging.info(f"Saved {len(links)} links for category '{category}' to the database.") except Exception as e: logging.error(f"Database save error for category {category}: {e}") logging.debug(traceback.format_exc()) ``` The save\_links\_to\_db function is where we actually save the links we found for all the products into our database. Sort of like taking everything you wrote down in your notepad and typing in a real file that won't get lost. This function needs three things to work: 1. The database connection we set up 2. The category with the equipment (tractors, harvesters, etc.) 3. A set of product links we just scraped from the website. Here's how it works: - The function loops through each link one after the other and tries to insert into the database. - We perform a neat little trick with the database: INSERT OR IGNORE. This tells the database: "hey, try to add this link, but if it already exists, ignore it." This means we avoid saving the same product/item more than once and we don't have to deal with annoying errors. - Instead of saving (committing) each link individually, it saves (commits) them all at once at the end. This is much more efficient—like putting everything in one box and carrying it, instead of taking up to 10 trips. As per all of our scraper, we added logging here as well. So every time we save links, we also log how many were saved and from which category. That way, we can track how thing are progressing. And like all of the other functions, if something goes wrong, it doesn't crash the entire app. It can catch an error, log what happened, then move on. So even if our scraper encounters a hiccup it remain strong. ### Scraper Functions ``` def load_all_products(page): """ Perform infinite scrolling on the page to load all products. This function simulates infinite scrolling by repeatedly scrolling to the bottom of the page until no new content is loaded. It is used to ensure that all products are loaded on pages with lazy-loading or infinite scrolling mechanisms. Args: page (playwright.sync_api.Page): The Playwright page object representing the browser page. Returns: None Logs: - INFO: Each scroll action performed. - INFO: When no more content is loaded, indicating the end of scrolling. - WARNING: If there is an error during the scrolling process. """ try: previous_height = None while True: # Scroll to the bottom of the page page.evaluate("window.scrollTo(0, document.body.scrollHeight)") logging.info("Scrolled to the bottom of the page.") time.sleep(5) # Wait for new content to load # Get the current height of the page current_height = page.evaluate("document.body.scrollHeight") if previous_height == current_height: # Stop scrolling if no new content is loaded logging.info("No more content to load.") break previous_height = current_height except Exception as e: logging.warning(f"Error during infinite scrolling: {e}") logging.debug(traceback.format_exc()) ``` The load\_all\_products function helps us out with something called infinite scrolling - which is sometimes difficult to scrape. Let's clarify. I imagine you have encountered a website where there are more products that keep loading as you scroll down. That is an example of infinite scrolling, where instead of loading all of the items at once, the website provides more items while you're scrolling down. For users this creates a very nice experience, but for scrapers it means we have to simulate human scrolling otherwise we miss a lot of content! This is what this function does. It automatically scrolls down to the bottom of the page, waits a few seconds, and then asks the question: "Um, is the page any taller?" (which is means new products were added) If the page does not get taller after we scroll, it is signal that there are no more products to load. We made it to the bottom! On each scroll, there’s also a five second pause to allow the website to load more products, as you would do naturally waiting a moment while scrolling for the next items to suddenly appear. In addition, this function can be tough and clever. If something unexpected happens, for example the scroll fails to work unexpectedly, If for whatever reason the page does not load properly. This function won’t crash everything entirely, it manages to catch those errors and log them, enabling you to check later after that wrap up. To keep it short: This function ensures we are not missing any products that are hidden behind infinite scrolling. Simply put, it would be like taking your time to scroll down a shopping site patiently until you have looked at every item sitting on the shelf. ``` def extract_product_links(html): """ Extract product links from the HTML content of the page. This function parses the HTML content using BeautifulSoup and extracts all product links from the specified CSS selector. It constructs full URLs for relative links. Args: html (str): The HTML content of the page as a string. Returns: set: A set of unique product links extracted from the page. Logs: - INFO: The number of links extracted from the page. - ERROR: If there is an error during the HTML parsing process. """ links = set() try: soup = BeautifulSoup(html, "html.parser") # Select all anchor tags within the specified CSS selector for a in soup.select(" div.flex.justify-between.items-center > a"): href = a.get("href") if href: # Construct the full URL if the link is relative full_url = "https://www.allmachines.com" + href if href.startswith("/") else href links.add(full_url) logging.info(f"Extracted {len(links)} product links from page.") except Exception as e: logging.error(f"HTML parsing error: {e}") logging.debug(traceback.format_exc()) return links ``` Alright, let’s talk about the heart of our scraper — the part where we actually grab the product links from the page. That’s what the extract\_product\_links function does. Imagine you’ve just scrolled all the way down on an online store’s page (like we did in the previous step). Now, your screen is filled with tons of farming equipment listings. What we need to do next is pick out the links that lead to each of those products. That’s where this function comes in. First, it takes the full web page (in HTML format — kind of like the skeleton of a webpage) and feeds it into a tool called BeautifulSoup. This tool is like a super-smart highlighter — it helps us easily spot and pull out specific parts of the webpage. Think of it as using "Find" in a document, but with extra powers. In this case, we’re looking for special pieces of HTML code — anchor tags () inside certain boxes (or
    s) that are styled in a specific way. These are the boxes that AllMachines uses to hold product links. We tell BeautifulSoup, “Hey, find any anchor tags that live inside boxes with these class names: flex, justify-between, and items-center.” (Class names are just labels that websites use to organize their layout.) Once we find those tags, we grab the href attribute from each one — that’s the part that holds the actual URL. But sometimes these links are only partial, like:/product/tractor-123To fix that, we add the website’s main URL at the beginning, so it becomes a full, working url. One more cool trick here — we use a set to store all the links. Why? Because a set automatically removes duplicates. So even if the same product appears twice on the page, we’ll only keep it once. And of course, just like the other functions, we’ve added some error handling. That means if something goes wrong while reading the page, the whole scraper won’t crash — it will simply log the error so you can check it later. Oh, and yes — it also logs how many links it found, which is super helpful to track how things are going. ``` def scrape_category(category, url, conn): """ Scrape a specific category page for product links and save them to the database. This function navigates to the category page, performs infinite scrolling to load all products, extracts product links from the page, and saves them to the database. Args: category (str): The category of the products (e.g., "tractors", "combine-harvesters"). url (str): The URL of the category page to scrape. conn (sqlite3.Connection): The SQLite database connection object. Returns: None Logs: - INFO: The start and completion of scraping for the given category. - ERROR: If there is an error during the scraping process. """ try: with sync_playwright() as p: # Launch the browser browser = p.chromium.launch(headless=False) context = browser.new_context() # Create a new browser context page = context.new_page() # Open a new page logging.info(f"Scraping category: {category} | URL: {url}") page.goto(url) # Navigate to the category URL time.sleep(3) # Wait for the page to load load_all_products(page) # Perform infinite scrolling html = page.content() # Get the page content product_links = extract_product_links(html) # Extract product links save_links_to_db(conn, category, product_links) # Save links to the database browser.close() # Close the browser except Exception as e: logging.error(f"Error scraping category '{category}': {e}") logging.debug(traceback.format_exc()) ``` Let’s break down what the scrape\_category function does — it’s kind of like the team leader in charge of collecting product links from one specific equipment category (like tractors or harvesters). Here’s the step-by-step story: When this function starts, it opens up a browser window — just like you would when using Chrome or Firefox. We’re using a tool called Playwright to do this. It lets us control the browser automatically with code (like a robot clicking and scrolling for us). And since we’re setting headless=False, the browser actually opens up on your screen — which is super helpful while you're testing or debugging, so you can watch the scraper in action. Next, the browser goes to a category page — say, the one for “Tractors.” Once the page loads, the scraper starts scrolling down automatically — just like a real user. This is important because many websites only load more products when you scroll, a feature called infinite scrolling. So, we keep scrolling until everything is loaded. After that, we grab all the HTML (the website’s behind-the-scenes code), and pass it to a helper function that picks out all the product links on the page. It’s kind of like scanning through a grocery list and pulling out all the items you need. Once we have the links, we save them into a database using another function. This way, we have a nice organized list of links that we can come back to later. Finally, the browser closes down, so we’re not using up unnecessary memory or keeping windows open in the background. It’s like washing your hands and tidying up your desk after you're done. And what if something goes wrong? Don’t worry — this function also includes error handling. That means if something fails (like a page doesn’t load properly), it won’t crash the whole process. Instead, it’ll log the error, skip that category, and move on to the next one. So your scraper keeps running smoothly even if one piece goes a little off track. ### Main Script ``` def main(): """ Main function to scrape all categories and save product links to the database. This function initializes the database, iterates through all categories and their URLs, scrapes each category page for product links, and saves the extracted links to the database. Steps: 1. Initialize the SQLite database. 2. Iterate through all categories and their URLs. 3. Scrape each category page for product links. 4. Save the extracted links to the database. Logs: - INFO: The start and completion of the scraping process. - CRITICAL: Any critical errors encountered during execution. Returns: None """ category_urls = { #hint:tractor section can be is filtered and scraped because of high number of products and to handle infinite scrolling "tractors": "https://www.allmachines.com/tractors/view-all", "combine-harvesters": "https://www.allmachines.com/combine-harvesters/view-all", "balers": "https://www.allmachines.com/balers/view-all", "forage-harvesters": "https://www.allmachines.com/forage-harvesters/view-all", "combine-headers": "https://www.allmachines.com/combine-headers/view-all", "forage-headers": "https://www.allmachines.com/forage-headers/view-all", "rakes": "https://www.allmachines.com/rakes/view-all", "tedders": "https://www.allmachines.com/tedders/view-all", "specialty-crop-harvesters": "https://www.allmachines.com/specialty-crop-harvesters/view-all" } try: conn = init_db() # Initialize the database for category, url in category_urls.items(): scrape_category(category, url, conn) # Scrape each category conn.close() # Close the database connection logging.info("All categories scraped and saved to SQLite successfully.") except Exception as e: logging.critical(f"Fatal error in main(): {e}") logging.debug(traceback.format_exc()) if __name__ == "__main__": main() ``` When you are beginning your journey into web scraping using Python, it can be hard to keep track of everything going on in a script. One of the most important parts of the script — and one of the pieces that ties everything together — is the main() function. You can think of it as the control center for your scraper, as it is what directs the code for what to do, when to do it, and in what order it should be done. In our case, for our web scraping project, the main() function starts off with setting up a list of the categories we would like to scrape. Each category (for example, tractors, excavators and forklifts) has its own webpage on the AllMachines website. We setup each category with its specific URL as a dictionary — somewhat like taking note of what you need to pick up from different supermarkets to be as efficient as possible, when you actually do your shopping! After we prepare the list, the function opens a connection to a database where all of the product details we scrape will be stored (like names, prices, or specs). Then it processes each category in order and passes the URL to the function called scrape\_category(). The scrape\_category() function will do all of the real scraping, visiting the page, collecting the data and saving it. While the function called main() is just the hub for it all to keep moving. Once all categories have been scraped, the main() function closes the database connection. This is crucial to keep it neat and tidy so that data doesn't get corrupt or lost. It is like closing the lid on a storage box after packing everything away neatly. A neat feature of this script is the last line, if \_\_name\_\_ == "\_\_main\_\_": .This may look odd to you now, but this just means, "only run this script when we're running it as a script." If someone else wants to reuse some of our code — only the scraping function for example — they can import this script and not have the whole scraping process happen. It's just a nice way to make your code semantic and reusable. Structuring our code in this manner makes it simple to update or add new categories down the road. The script is neat, easy to follow, and not likely to include bugs. For people learning how to build real scraping tools, this is a really good example of keeping things simple and practical. So if you're new to Python or scraping, don't worry! main() is something to be understood, not feared. It is simply there to ensure that everything runs in order (like a recipe). Once you understand how to run it in your IDE and how it works the rest of the code will make a lot more sense! ## Data Collection Now we enter the second phase of our web scraping journey— this is where the action takes place! By this point, the scraper will visit every URL in our lists. But it is NOT just going to simply grab what is visible to it — it will patiently wait until the entire page has been fully loaded. This is key since, many sites today use javascript to load data dynamically — meaning the information is not necessarily all available upon page opening. What kinds of things do we collect? Everything from product names and descriptions, prices, catchy marketing phrases, unique scores or ratings, important product features, and even deep technical specs, basically everything you want to know about before making a purchase on a piece of agriculture equipment! In the background, we keep track of everything we scrape using SQLite — think of it as a notebook that tracks the things we've already scraped! With this data organization, we can easily resume where we left off if something goes wrong and need to pause or restart. In addition, it safely stores all the data in such a fashion that we can access it later for analysis, as reports, or funnel it into other apps. So with smart tools and good organization, we can ensure that our scraping is much more than a one-time solution. Our scraping solution is reliable, easy-to-maintain, and resilient so we can scale as the website grows. This means that whether we are tracking 10 products or 10,000 products, our scraper will work, and that's a major accomplishment when working with the scaled data in the agriculture and machinery space! ### Import Section ``` import sqlite3 import logging import traceback import time import json from bs4 import BeautifulSoup from playwright.sync_api import sync_playwright ``` In this step, we are going to make only a small change to our list of imported libraries: we're adding the json library. This little addition really helps us when dealing with complex and nested pieces of data! Think about it like packing a messy drawer into a neat labeled box. This library will help us take all the detailed information about our product and put it into a standard format (string) that is easy to save in our database. Other than that, all the other imports remain the same as we had during our previous step when we were capturing the product links. So there's nothing super new or complicated about this step - simply a useful upgrade to accommodate more products and more detailed data! ### Logging Configuration ``` # Configure logging to write logs to a file with a specific format logging.basicConfig( filename='product_scraper.log', filemode='w', # Overwrite the log file each time the script runs level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s' ) ``` In this step, we will make only a minor adjustment to the way we do logging - and it is simply a matter of organization. Instead of writing to the same log file as we did in the first phase, we will write to a new log file called product\_scraper\_phase2.log. This will help us log everything in a separate and organized way, so that we can keep track of what is going on, and fix problems if there are any. Other than simply changing the name of the log file, everything else about the logging remains absolutely unchanged - it has the same format, the same behavior, and it records in exactly the same way. This is a small change that will help keep us organized as our scraper works through different stages of the process! ### Database Functions Now let's discuss the most important part of the scraping project; the database functions. We need functions that will take our scraper from a simple one-off script to a fully functional tool that can execute again and again without losing anything. These functions allow us to track what the scraper has accomplished. By utilizing these functions, we are not simply running requests to grab data, we are building a system that remembers what URLs it has crawled through and that keeps track of thousands of product URLs and all the information we collect about those products - including names, prices, features - all safely stored in one tidy SQLite database file. Another advantage of using a SQLite database versus a complicated server, is that it runs immediately, straight out of the box. SQLite is lightweight and ideal for a project like this. And these database functions are so much more than storing data; they also create an intelligent, organized environment that can deal with interruptions, resume where it left off, and keep our existing, collected data safe and accessible. ``` def connect_db(db_name="allmachines_products.db"): """ Establish a connection to the SQLite database. Args: db_name (str): Name of the SQLite database file to connect to. Defaults to "allmachines_products.db". Returns: sqlite3.Connection: An active connection to the specified SQLite database. Example: >>> conn = connect_db() >>> # Perform operations with conn >>> conn.close() """ return sqlite3.connect(db_name) ``` At first glance, the connect\_db() function might seem like no big deal — just a tiny piece of the code that connects to the database. But don’t let its simplicity fool you! This little function plays a huge role in keeping our project clean, consistent, and easy to manage. By putting all the database connection logic in one place, we create a single point of control. That means the rest of our code can just call connect\_db() whenever it needs to talk to the database — no need to repeat the same connection code over and over again. Why is this a big deal? Imagine you decide to change the database file name, or even switch to a completely different database system someday. If your connection code is scattered everywhere, you’d have to dig through every file to update it. But with connect\_db(), you just make the change once — and the rest of the project still works like a charm. It’s a simple trick, but one that makes your code easier to maintain and way more flexible in the long run. ``` def ensure_scraped_column(conn): """ Ensure the 'scraped' column exists in the product_links table. This function checks if the 'scraped' column exists in the product_links table. If it doesn't exist, the function adds the column with a default value of 0, representing that URLs have not been scraped yet. Args: conn (sqlite3.Connection): An active database connection. Raises: Exception: If there's an error checking for or adding the column. Notes: - Uses PRAGMA table_info to query table structure - Default value of 0 indicates not scraped - Logs success or failure of the operation """ try: cursor = conn.cursor() cursor.execute("PRAGMA table_info(product_links)") columns = [row[1] for row in cursor.fetchall()] if 'scraped' not in columns: cursor.execute("ALTER TABLE product_links ADD COLUMN scraped INTEGER DEFAULT 0") conn.commit() logging.info("Added 'scraped' column to product_links table.") except Exception as e: logging.error(f"Error ensuring scraped column: {e}") logging.debug(traceback.format_exc()) ``` The ensure\_scraped\_column() function is a great example of smart and safe coding — also known as defensive programming. Instead of assuming that everything in the database is already perfect, this function double-checks the setup before moving forward. Here’s what it does: it looks at the structure of the database table (kind of like peeking under the hood) using a special SQLite feature called PRAGMA. If it finds that the “scraped” column is missing — which is important for tracking progress — it simply adds it. No fuss, no errors, no need for you to run a separate setup script. Pretty cool, right? This makes the scraper super flexible. Whether it’s your very first run or you’re continuing after a break, the function makes sure everything is in place so the process can run smoothly. The "scraped" column it adds starts off with a value of 0 for each URL, meaning "not scraped yet." Later on, this little flag helps the scraper figure out exactly where it left off, so it can resume without repeating anything or missing a step. It’s a small function with a big impact — making the whole system more reliable, user-friendly, and able to recover gracefully from interruptions. ``` def init_data_table(conn): """ Initialize the product_data table to store scraped product details. Creates the product_data table if it doesn't exist. This table stores all the product information extracted from product pages, including metadata, descriptive content, and structured data in JSON format. Args: conn (sqlite3.Connection): An active database connection. Table Schema: - id: Primary key, auto-incremented integer - category: Text field for product category classification - url: Text field with unique constraint for product page URL - title: Text field for product title - description: Text field for product description - price: Text field for product price - tagline: Text field for product tagline - equipment_score: Text field for AllMachines Equipment Score - highlights: Text field storing product highlights as JSON - specifications: Text field storing product specifications as JSON Raises: Exception: If there's an error creating the table. """ try: cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT UNIQUE, title TEXT, description TEXT, price TEXT, tagline TEXT, equipment_score TEXT, highlights TEXT, specifications TEXT ) """) conn.commit() logging.info("Initialized 'product_data' table.") except Exception as e: logging.error(f"Error initializing data table: {e}") logging.debug(traceback.format_exc()) ``` ``` def init_error_table(conn): """ Initialize the error_urls table to store URLs that failed to scrape. Creates the error_urls table if it doesn't exist. This table tracks URLs that encountered errors during the scraping process, allowing for retry attempts or manual investigation. Args: conn (sqlite3.Connection): An active database connection. Table Schema: - id: Primary key, auto-incremented integer - category: Text field for product category - url: Text field with unique constraint for failed URL Raises: Exception: If there's an error creating the table. """ try: cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS error_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT UNIQUE ) """) conn.commit() logging.info("Initialized 'error_urls' table.") except Exception as e: logging.error(f"Error initializing error table: {e}") logging.debug(traceback.format_exc()) ``` The init\_data\_table() and init\_error\_table() functions do more than just create database tables — they’re actually setting the stage for how we organize and manage the complex data we collect while scraping. Let’s break it down. The product\_data table is built in a really smart way. For simple stuff like product names and prices, it uses regular database fields (think: neat columns like in a spreadsheet). But for more complicated information — like technical specs or a list of features — it uses JSON, a format that’s perfect for storing detailed or layered data. This mix gives us the best of both worlds: we can easily search and filter simple data, while still having room to store complex details without needing to force everything into a strict structure. Then there’s the error\_urls table — and this is a lifesaver! Sometimes, when scraping the web, a page might not load properly or its layout might suddenly change. Instead of crashing or giving up, the scraper calmly writes down that URL into this table so we can check it later or try again. It’s like keeping a list of “things to follow up on” instead of just forgetting them. Together, these two functions make the scraper reliable, flexible, and resilient — all things you want when dealing with the ever-changing, sometimes unpredictable world of web pages. It’s a big step up from a basic script and makes the project feel much more professional and production-ready. ### HTML Parsing Functions The HTML parsing functions serve as the scrapers' brains - they are the ones really thinking and ripping off usable information from the untidy web pages. You can imagine a web page as a huge bowl of alphabet soup. The functions are the surgeons (or soup sleuths!) that are good at finding the right words or ingredients we care about that seem to float around aimlessly. Each function has its own task - maybe one pulls out the product name, another, the price, while another might dig at the technical specifications. They were created to be intelligent and flexible, so even if the website changed a tiny bit (which happens all of the time!), the functions can typically still figure it out without breaking. This section of the code is where the real magic happens - taking a pile of HTML and turning it into clean data we can use for whatever purpose, be it analysis, display, or storage. This is the step that takes chaos and creates clarity, ensuring that the information we have collected is accurate, complete, and easily used. ``` def parse_title(soup): """ Extract the product title from the BeautifulSoup object. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: The product title text if found, None otherwise. Notes: - Uses a specific CSS selector to locate the title element - Returns None if the element is not found or an error occurs - Strips whitespace from the extracted text """ try: return soup.select_one("body > main > div > div.flex.items-start.justify-start.mt-2 > div > h1").text.strip() if soup.select_one("body > main > div > div.flex.items-start.justify-start.mt-2 > div > h1") else None except Exception as e: logging.error(f"Error parsing title: {e}") logging.debug(traceback.format_exc()) return None def parse_description(soup): """ Extract the product description from the BeautifulSoup object. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: The product description text if found, None otherwise. Notes: - Targets a specific div with class containing 'three-line-clamp' - Returns None if the element is not found or an error occurs - Strips whitespace from the extracted text """ try: return soup.select_one("div.three-line-clamp.lg\:line-clamp-none.\!overflow-hidden > p").text.strip() if soup.select_one("div.three-line-clamp.lg\:line-clamp-none.\!overflow-hidden > p") else None except Exception as e: logging.error(f"Error parsing description: {e}") logging.debug(traceback.format_exc()) return None def parse_price(soup): """ Extract the product price from the BeautifulSoup object. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: The price text if found, None otherwise. Notes: - Uses CSS selector to find price element in a specific div structure - Returns None if the element is not found or an error occurs - Strips whitespace from the extracted text """ try: return soup.select_one("div > div > div.mb-5 > div").text.strip() if soup.select_one("div > div > div.mb-5 > div") else None except Exception as e: logging.error(f"Error parsing price: {e}") logging.debug(traceback.format_exc()) return None ``` There are some fundamental parsing functions used in the scraper such as: parse\_title(), parse\_description(), and parse\_price(), which are among the most important parts of our code. The basic parsing functions in the scraper are the pieces of code that kill a web page and pull out the specific pieces of information that we care about the most — what the product is, what the product does, and how much it costs. To accomplish this, we will use rules that are known as CSS selectors, the web equivalent of GPS, to tell the scraper exactly where the individual pieces of information are located on the webpage. Finding selectors can be an exercise in investigations - we really had to search through the underlying code of the AllMachines website and attempt each selector path to determine the best approaches to achieving success. Although many of these functions appear quite simple, there was a lot of thinking behind them. Also, websites can change at any time (and they frequently change), so we put the calls to each function, wrapped in a safety net, called a try-except block. This way, if just a tiny thing breaks one function - like the product price is missing or moved - the whole scraper will not crash. Instead, it will log the problem for us to come back to later and then continue and keep going. This makes the scraper more robust and much less likely to fail just because one important aspect was missing. ``` def parse_tagline(soup): """ Extract and format the product tagline from the BeautifulSoup object. This function locates the tagline div, extracts text from spans within it, and formats them with a comma after the first span's text. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: The formatted tagline text if found, None otherwise. Format Details: - If multiple spans exist: "{first_span_text}, {remaining_text}" - If only first span exists: just the first span text - If no spans but div exists: all text inside the div Notes: - Handles complex nested structure with multiple spans - Special formatting with comma after first span's text - Returns None if the target div is not found or an error occurs """ try: tagline_div = soup.select_one("div.flex.flex-wrap.items-center.gap-2.mt-4.text-sm.leading-4") if tagline_div: # Extract all span elements inside the div spans = tagline_div.find_all("span") if spans: # Get the text of the first span first_span_text = spans[0].text.strip() # Get the remaining text inside the div (excluding the first span) remaining_text = tagline_div.get_text(separator=" ", strip=True).replace(first_span_text, "", 1).strip() # Combine the first span text with the remaining text, separated by a comma tagline = f"{first_span_text}, {remaining_text}" if remaining_text else first_span_text return tagline else: # If no spans are found, return all text inside the div return tagline_div.get_text(separator=" ", strip=True) return None except Exception as e: logging.error(f"Error parsing tagline: {e}") logging.debug(traceback.format_exc()) return None def parse_equipment_score(soup): """ Extract the AllMachines Equipment Score from the BeautifulSoup object. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: The equipment score text if found, None otherwise. Notes: - Targets a specific div with classes related to the score display - The score element has specific styling (semibold font, background color) - Returns None if the element is not found or an error occurs """ try: score_element = soup.select_one("div.flex.items-center.gap-2.text-sm.font-semibold.w-fit.py-2.bg-white.relative.z-10") if score_element: return score_element.text.strip() return None except Exception as e: logging.error(f"Error parsing equipment score: {e}") logging.debug(traceback.format_exc()) return None def parse_highlights(soup): """ Extract and format product highlights as a JSON array of key-value pairs. This function locates the highlights section, extracts key-value pairs from list items, and returns them as a JSON-formatted string. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: JSON string containing highlight key-value pairs if found, None otherwise. JSON Format: [{"key1": "value1"}, {"key2": "value2"}, ...] Notes: - Targets a specific section with class 'my-8.lg\:my-14.lg\:\!my-9.relative' - Each highlight is represented as a dictionary with a single key-value pair - Returns None if the highlights section is not found or an error occurs """ try: highlights_section = soup.select_one("section.my-8.lg\:my-14.lg\:\!my-9.relative") # Locate the section if highlights_section: highlights = [] list_items = highlights_section.select("ul > li") # Select all
  • elements for li in list_items: key = li.select_one("span.justify-self-center").text.strip() if li.select_one("span.justify-self-center") else None value = li.select_one("div.items-center.card-body").text.strip() if li.select_one("div.items-center.card-body") else None if key and value: highlights.append({key: value}) # Add key-value pair to the list return json.dumps(highlights) # Convert the list to a JSON string return None except Exception as e: logging.error(f"Error parsing highlights: {e}") logging.debug(traceback.format_exc()) return None def parse_specifications(soup): """ Extract and format product specifications as a structured JSON array. This function locates the specifications section (third div with class 'scroll-mt-40'), extracts data from tables within it, and organizes it into a hierarchical JSON structure. Args: soup (BeautifulSoup): Parsed HTML of the product page. Returns: str or None: JSON string containing structured specification data if found, None otherwise. JSON Format: [ { "Table Title 1": [ {"Row Header 1": "Value 1"}, {"Row Header 2": "Value 2"} ] }, { "Table Title 2": [ {"Row Header 1": "Value 1"}, {"Row Header 2": "Value 2"} ] } ] Notes: - Targets the third div with class 'scroll-mt-40' - Each table in the specifications section becomes a key in the output - Table headers become dictionary keys with row values - Returns None if the specifications section is not found or an error occurs """ try: # Locate all divs with the class 'scroll-mt-40' specifications_sections = soup.select("div.scroll-mt-40") # Ensure there are at least three divs and select the third one if len(specifications_sections) >= 3: specifications_section = specifications_sections[2] # Select the third div specifications = [] # Find all tables within the third div tables = specifications_section.select("table.w-full") for table in tables: # Get the table title from the thead title = table.select_one("thead").get_text(strip=True) if table.select_one("thead") else None if not title: continue # Skip tables without a title # Parse the rows in the tbody rows = table.select("tbody > tr") table_data = [] for row in rows: key = row.select_one("th").get_text(strip=True) if row.select_one("th") else None value = row.select_one("td").get_text(strip=True) if row.select_one("td") else None if key and value: table_data.append({key: value}) # Add key-value pair to the table data # Add the table title and its data to the specifications if table_data: specifications.append({title: table_data}) return json.dumps(specifications) # Convert the list to a JSON string return None except Exception as e: logging.error(f"Error parsing specifications: {e}") logging.debug(traceback.format_exc()) return None ``` Now let’s talk about the real magic—the advanced parsing functions. These parts of the scraper go beyond just grabbing simple text; they actually understand how information is organized on a web page and then carefully rebuild that structure into clean, usable data. Take the parse\_tagline() function, for example. Sometimes websites use multiple layers (like several tags inside each other) to style or highlight different parts of a product’s tagline. This function smartly navigates through that tangled structure and extracts the full tagline exactly as it was meant to be seen. Then there are functions like parse\_highlights() and parse\_specifications(). These are even more impressive. They don’t just pull random pieces of text—they figure out how everything fits together. Imagine looking at a table of features on a product page, with labels on the left and values on the right. These functions understand that relationship and convert the whole thing into something structured, like a neat dictionary or a JSON object. That way, when we later want to use this data—for analysis, filtering, or even showing it in an app—it’s already organized and easy to work with. This process of turning messy, semi-organized website code into clean, structured data is one of the most powerful things our scraper does. It’s what turns a bunch of web pages into something truly useful. ``` def parse_product_page(html): """ Parse the complete product page HTML to extract all relevant product data. This function serves as a centralized parser that coordinates the extraction of all product details by creating a BeautifulSoup object and passing it to specialized parsing functions for each data component. Args: html (str): Raw HTML content of the product page. Returns: tuple: A 7-tuple containing: - title (str or None): Product title - description (str or None): Product description - price (str or None): Product price - tagline (str or None): Product tagline - equipment_score (str or None): AllMachines Equipment Score - highlights (str or None): JSON string of product highlights - specifications (str or None): JSON string of product specifications Notes: - Creates a BeautifulSoup object using the 'html.parser' - Delegates extraction of each component to specialized functions - Returns None for any component that fails to parse - Logs any errors that occur during parsing """ try: soup = BeautifulSoup(html, 'html.parser') title = parse_title(soup) description = parse_description(soup) price = parse_price(soup) tagline = parse_tagline(soup) equipment_score = parse_equipment_score(soup) highlights = parse_highlights(soup) specifications = parse_specifications(soup) return title, description, price, tagline, equipment_score, highlights, specifications except Exception as e: logging.error(f"HTML parsing error: {e}") logging.debug(traceback.format_exc()) return None, None, None, None, None, None, None ``` At the center of our scraping setup is a function called parse\_product\_page()—and honestly, it’s the real conductor of the orchestra. Think of it like this: when you’re cooking a big meal, instead of trying to make every dish in one huge pot, you use different pans and tools for different tasks—one for boiling, one for frying, one for baking. That’s exactly what this function does! Rather than trying to handle everything itself, it lets other smaller, specialized functions (like parse\_title(), parse\_price(), and so on) do their specific jobs. What makes this setup so smart is that it builds the BeautifulSoup object—the thing that reads and organizes the webpage’s HTML—only once. Then it passes this organized version of the page to each of the helper functions. This saves time and avoids doing the same work over and over. Also, if anything changes on the website—for example, if the way product titles are displayed is updated—you only need to fix the one small function that handles titles. Everything else keeps working smoothly. That makes the whole scraper way easier to maintain and less likely to break. And here’s another cool thing: if one part of the page is broken or weird, the scraper doesn’t crash. It just skips that piece and keeps going. This means we can still collect lots of useful information even from pages that aren’t perfect. In short, parse\_product\_page() ties everything together in a neat, reliable way—and it’s a big reason why this scraper works so well behind the scenes. ### Scraping Functions The scraping functions are where all the action happens—they’re basically the heart of our whole system. This is the moment when the scraper goes out into the wild (a.k.a. the internet), visits real web pages, and starts collecting information. Imagine a robot that not only knows how to browse a website like a human but can also pick out the exact details we care about—like product names, prices, and features—and then neatly save them into a database. That’s what these functions do. They’re smart enough to handle today’s modern, often tricky websites (some of which don’t even load all their content right away), and they do it over and over again across thousands of different pages. What’s really impressive is how well all the parts work together—browser automation helps open and load the pages, HTML parsing figures out what information to grab, and the database stores everything safely. It’s like a carefully choreographed dance, where each step has to be in sync. If even one move goes off, things could fall apart. But thanks to this setup, the whole process runs smoothly and reliably. ``` def scrape_product_data(conn, url, category): """ Scrape product data from a given URL and save it to the database. This function orchestrates the complete scraping process for a single URL: launching a browser, navigating to the page, extracting content, parsing data, and saving results to the database. It also handles error conditions and updates scraping status. Args: conn (sqlite3.Connection): An active database connection. url (str): The URL of the product page to scrape. category (str): The category of the product. Process Flow: 1. Launch a headless Chromium browser using Playwright 2. Navigate to the target URL and wait for page load 3. Get the page content and parse product details 4. If successful (title exists), save data and mark URL as scraped 5. If unsuccessful, record the URL in the error table 6. Close the browser Notes: - Uses Playwright's synchronous API for browser automation - Sets a generous 60s timeout for page navigation - Includes a 3-second wait for JavaScript-loaded content - Records failed URLs for potential retry """ try: with sync_playwright() as p: browser = p.chromium.launch(headless=True) # Launch browser context = browser.new_context() page = context.new_page() logging.info(f"Scraping URL: {url}") page.goto(url, timeout=60000) # Navigate to the URL time.sleep(3) # Wait for the page to load html = page.content() # Get the page content title, desc, price, tagline, equipment_score, highlights, specifications = parse_product_page(html) # Parse the product details if title: save_product_data(conn, category, url, title, desc, price, tagline, equipment_score, highlights, specifications) # Save product data mark_as_scraped(conn, url) # Mark the URL as scraped else: save_error_url(conn, category, url) # Save the URL to the error table browser.close() # Close the browser except Exception as e: logging.error(f"Error scraping URL {url}: {e}") logging.debug(traceback.format_exc()) save_error_url(conn, category, url) # Save the URL to the error table in case of failure ``` The scrape\_product\_data() function is where our scraper really shows off what it can do. Instead of just grabbing raw webpage code (like simple scrapers that send basic requests), this one uses a powerful tool called Playwright. Think of Playwright like a mini web browser that our program controls—it opens the webpage, waits for everything to load (just like how you’d wait for all the pictures and buttons to appear when you open a site), and then gets to work collecting the data. It even runs JavaScript—just like a real browser—so we don’t miss out on content that loads a little later or gets added by scripts running in the background. This is super helpful for modern websites that don’t show everything right away. The function is designed to be smart and responsible too. It opens the browser, does its job, and then neatly closes everything—even if something goes wrong. This avoids wasting memory or accidentally leaving background tasks running. There are also some clever timing tricks in place. For example, the scraper is patient—it gives each page up to 60 seconds to load (just in case the internet is slow or the site is heavy), and then waits an extra 3 seconds to make sure any late-loading bits of the page have time to show up before we start extracting data. This little detail really helps us get complete and accurate results every time. ``` def save_product_data(conn, category, url, title, description, price, tagline, equipment_score, highlights, specifications): """ Save the scraped product data to the database. Inserts all product information into the product_data table. Uses INSERT OR IGNORE to handle potential duplicate URLs gracefully (avoiding constraint violations). Args: conn (sqlite3.Connection): An active database connection. category (str): The category of the product. url (str): The URL of the product page. title (str): The product title. description (str): The product description. price (str): The product price. tagline (str): The product tagline. equipment_score (str): The AllMachines Equipment Score. highlights (str): JSON string of product highlights. specifications (str): JSON string of product specifications. Notes: - Uses INSERT OR IGNORE to handle potential URL uniqueness constraint violations - Logs success or failure of the insertion operation - Commits the transaction to ensure data is saved """ try: cursor = conn.cursor() cursor.execute(""" INSERT OR IGNORE INTO product_data (category, url, title, description, price, tagline, equipment_score, highlights, specifications) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) """, (category, url, title, description, price, tagline, equipment_score, highlights, specifications)) conn.commit() logging.info(f"Saved product data for URL: {url}") except Exception as e: logging.error(f"Error saving product data for {url}: {e}") logging.debug(traceback.format_exc()) ``` The database functions that support our scraper are designed with smart data-handling habits. One great example is the save\_product\_data() function. This function takes the product details we just scraped and makes sure they're safely saved into our database. But here’s the clever part: it uses a special SQL trick called "INSERT OR IGNORE". That means if we’ve already saved data for a certain product (maybe because we ran the scraper before or had to restart it), the database won’t freak out or crash—it’ll just skip over that entry and move on. No duplicates, no errors. This makes our scraper more reliable and allows it to pause, restart, or rerun without messing up the data or starting from scratch. It’s a small detail, but it’s what makes the whole system more professional and stress-free to use. ``` def mark_as_scraped(conn, url): """ Mark a URL as successfully scraped in the product_links table. Updates the 'scraped' column to 1 for the specified URL to indicate that it has been successfully processed and should not be scraped again. Args: conn (sqlite3.Connection): An active database connection. url (str): The URL to mark as scraped. Notes: - Sets the 'scraped' flag to 1 to indicate successful processing - Commits the transaction to ensure data is saved - Logs success or failure of the update operation """ try: cursor = conn.cursor() cursor.execute("UPDATE product_links SET scraped = 1 WHERE url = ?", (url,)) conn.commit() logging.info(f"Marked as scraped: {url}") except Exception as e: logging.error(f"Error updating scraped status for {url}: {e}") logging.debug(traceback.format_exc()) ``` The mark\_as\_scraped() function plays an important role in helping our scraper keep track of its progress. After it finishes collecting data from a product page, this function updates the database to mark that URL as "scraped." This is done by changing a special flag (which we set up earlier) from 0 to 1\. It’s a simple idea, but very powerful—it tells the system, “Hey, we already got data from this one, no need to visit again.” Thanks to this tracking system, the scraper can easily pause and resume without repeating work or missing anything. It’s like crossing off items on a checklist, ensuring nothing gets lost or done twice. ``` def save_error_url(conn, category, url): """ Record a URL that failed to scrape in the error_urls table. Inserts the failed URL and its category into the error_urls table for potential retry or manual investigation. Args: conn (sqlite3.Connection): An active database connection. category (str): The category of the product. url (str): The URL that failed to scrape. Notes: - Uses INSERT OR IGNORE to handle potential URL uniqueness constraint violations - Logs the operation with WARNING level (not as serious as ERROR) - Commits the transaction to ensure data is saved """ try: cursor = conn.cursor() cursor.execute(""" INSERT OR IGNORE INTO error_urls (category, url) VALUES (?, ?) """, (category, url)) conn.commit() logging.warning(f"Saved error URL: {url}") except Exception as e: logging.error(f"Error saving to error_urls for {url}: {e}") logging.debug(traceback.format_exc()) ``` The save\_error\_url() function shows a smart way to handle problems when something goes wrong. Instead of just writing an error message and moving on, this function actually saves the problem URL into a special database table. This means you can easily go back later, see which pages failed, and try scraping them again or figure out what caused the issue. Together with the other two database functions (save\_product\_data() and mark\_as\_scraped()), it follows best practices by always saving changes (committing) after updating the database. This helps protect your data even if the program crashes partway through. On top of that, all three functions use consistent error logging, which makes it much easier to understand what went wrong and fix it—just like a well-prepared toolkit for real-world issues. ### Main Driver Functions The main driver functions act like the directors of the entire scraping project. They don’t just run other parts of the code—they carefully guide when and how each part should work, making sure everything happens in the right order. These functions are what bring the whole system together, turning many smaller pieces into one smooth and automated process. Instead of just being a bunch of separate tools, the scraper works like a well-organized machine because of these drivers. They control the flow of the program from start to finish, making the system reliable, easy to understand, and ready for repeated use. ``` def get_unscraped_urls(conn):  """ Retrieve all unscraped URLs from the product_links table. Queries the database for all URLs that have not been scraped yet (where scraped = 0) along with their categories. Args: conn (sqlite3.Connection): An active database connection. Returns: list: A list of tuples, each containing (category, url) for unscraped URLs. Query Details: - Selects category and URL from product_links table - Filters for records where scraped = 0 - Returns all matching records as a list of tuples """ cursor = conn.cursor() cursor.execute("SELECT category, url FROM product_links WHERE scraped = 0") return cursor.fetchall() ``` The get\_unscraped\_urls() function is a smart way to manage the scraping workflow. Instead of going through all the URLs every time or keeping track of progress manually, this function checks the database and returns only the URLs that haven’t been scraped yet. You can think of it like a to-do list that updates itself automatically. This design has some great benefits: you can stop the scraper and restart it later without losing any progress; you can run multiple scraping sessions without repeating work; and the scraper always knows what’s left to do. The function itself is short and simple, but very powerful—it just runs one SQL query to find the right URLs based on a “scraped” flag we added earlier. This shows how smart database planning at the beginning can make everything else easier and more reliable later. ``` def main(): """ Main function to orchestrate the entire scraping process. This function serves as the entry point for the script and coordinates the overall workflow: 1. Connect to the database 2. Ensure required database structure (tables, columns) 3. Retrieve unscraped URLs 4. Process each URL to extract product data 5. Clean up resources Process Flow: 1. Establish database connection 2. Ensure database has required structure (ensure_scraped_column) 3. Initialize product_data and error_urls tables 4. Retrieve list of unscraped URLs from database 5. Iterate through URLs, scraping each one 6. Close database connection when complete Notes: - Logs the number of URLs found for processing - Logs completion of the scraping process - Database connection is properly closed after processing """ conn = connect_db() # Connect to the database ensure_scraped_column(conn) # Ensure the 'scraped' column exists init_data_table(conn) # Initialize the product_data table init_error_table(conn) # Initialize the error_urls table urls = get_unscraped_urls(conn) # Get all unscraped URLs logging.info(f"Found {len(urls)} unscraped URLs to process.") for category, url in urls: scrape_product_data(conn, url, category) # Scrape each URL conn.close() # Close the database connection logging.info("Scraping finished for all URLs.") if __name__ == "__main__": main() ``` The main() function acts like the conductor of an orchestra, making sure every part of the scraper works together in the right order. It follows a clear three-step structure: setup, processing, and cleanup. First, in the setup phase, it connects to the database and makes sure all the necessary tables and columns are ready. Then, in the processing phase, it gets the list of URLs that still need to be scraped and goes through them one by one—loading each page, extracting the data, and saving it. Finally, in the cleanup phase, it closes the database connection to free up system resources. This clean structure makes the function easy to understand and maintain. If you ever want to add new steps—like sending notifications or saving logs—you can easily do it in the right section without breaking anything. Even though the main() function looks simple, it’s the key piece that ties everything together and helps new developers quickly understand how the whole scraping system works. ## Conclusion This AllMachines web scraper is much more than just a script—it’s a full-featured data collection system built with care, precision, and real-world usability in mind. Instead of just grabbing a few lines of text, it turns messy web pages into clean, structured data that’s ready for analysis or use in other apps. The smart design is visible in every part of the scraper. If something goes wrong during scraping, the program doesn’t crash—it logs the issue clearly and keeps going. The database is designed not only to store data efficiently but also to keep track of what’s already been scraped, which makes the scraper perfect for long-term, large-scale use. One of the standout strengths is the modular approach: each function focuses on one task, making the code easier to read, maintain, and update as the website changes. Using Playwright to control a real browser and BeautifulSoup to read the HTML means this tool handles everything from JavaScript-heavy pages to tricky nested data. Most importantly, this scraper is built like a professional tool. It can pause and resume, logs any problems, and keeps your data safe with smart transaction handling. That makes it ideal for scraping thousands of products over time. As the AllMachines website updates its products and prices, this scraper is ready to keep your dataset current—making it a powerful tool for market research, business analysis, or anything else you might need. AUTHOR I’m Shahana, a Data Engineer at [Datahut](http://datahut.co/?ref=blog.datahut.co), where I build reliable and scalable data pipelines that turn unstructured web content into clean, usable datasets—especially for use cases like e-commerce, product intelligence, and market research. In this blog, I walked through a real-world scraping project where we collected detailed product information from AllMachines, a site that lists farming equipment. Using tools like Playwright, BeautifulSoup, and SQLite, we created a scraper that handles dynamic pages, avoids detection with smart user agent handling, and stores everything in an organized format for future analysis. At Datahut, we focus on building [web scraping](https://www.blog.datahut.co/post/web-scraping-challenges-you-need-to-know/) solutions that are not just effective, but also practical and robust—designed to handle real-world websites, recover from errors, and scale as needed. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. FAQ section ### FAQ 1: What is web scraping and how does it work? Web scraping is the process of automatically extracting data from websites using scripts or tools. It allows you to collect information like product titles, prices, and descriptions from multiple pages efficiently. 👉 Learn more about [what web scraping is and how it works](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/). ### FAQ 2: Which tools are used for scraping data from AllMachines? In this project, we used Playwright and BeautifulSoup — two powerful Python libraries for handling dynamic websites and parsing HTML content.Explore our [Python web scraping tutorial](https://www.blog.datahut.co/post/python-web-scraping-tutorial/) to understand how to use these tools effectively. ### FAQ 3: How do you handle infinite scrolling while web scraping? Websites like AllMachines may use infinite scrolling to load more products dynamically. To scrape such pages, you can use tools like Playwright to simulate user scrolling. Read how to [build smart and resilient web scrapers for dynamic websites](https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/) that handle infinite scrolling seamlessly. ### FAQ 4: How do you prevent being blocked during web scraping? Many websites detect and block bots during scraping. You can prevent this by rotating proxies, adding random delays, and mimicking human browsing patterns.Check out our expert guide on [how to maintain anonymity when web scraping at scale](https://www.blog.datahut.co/post/how-to-maintain-anonymity-when-web-scraping-at-scale-expert-tips/) for best practices. ### FAQ 5: What are common challenges faced in web scraping projects? Some of the biggest challenges in web scraping include dynamic website structures, CAPTCHAs, and legal compliance. Each of these can be managed with the right setup and ethical data collection approach. Discover more about [web scraping challenges you need to know](https://www.blog.datahut.co/post/web-scraping-challenges-you-need-to-know/). ### Fix Unit Economics by Cutting SKUs | Nestlé Strategy URL: https://www.blog.datahut.co/post/want-to-fix-your-unit-economics-do-what-nestle-did-start-saying-no-to-more-skus/ Last updated: 2026-07-23T07:48:31.000Z In 2021, Nestlé made a bold move that few companies of its size dare to make. They didn’t launch a new product. They started deleting them. [Project TASTY](https://www.nestle.com/sites/default/files/2022-11/investor-seminar-2022-latam.pdf?ref=blog.datahut.co) — Nestlé’s global SKU rationalization program — was launched to simplify the company’s portfolio and improve unit economics. Here’s what they discovered: 👉 34 % of Nestlé’s SKUs contributed just \~1 % of sales. 👉 Only 11 % of SKUs generated \~80 % of revenue. The logic was clear but courageous: if one-third of your 100,000 SKUs drive only 1% of sales, it’s time to clean house. By 2025, Nestlé reported over CHF 1.2 billion in savings — the result of trimming low-margin, low-rotation SKUs and focusing on the high performers. The outcome? ✅ Higher on-shelf availability (up to 97%) ✅ Improved service levels ✅ Stronger margins despite inflationary pressure As Nestlé CEO Mark Schneider explained, a leaner portfolio was “win-win-win”: - “Consumers get their popular items, retailers get fast-moving products, and we gain less complexity, more operational efficiency, and better on-shelf availability.” - That’s the power of data-driven simplification — cutting with clarity, not emotion. Behind this transformation lies a data story — one built on analytics, digital tools, and tough decisions guided by facts, not feelings. Because whether you’re Nestlé or a challenger brand, sometimes the smartest move is to delist strategically — not to shrink, but to simplify. And simplification, done with data, compounds into clarity, speed, and profitability. ![Project TASTY — Nestlé’s global SKU rationalization program](https://www.blog.datahut.co/content/images/2026/07/img-405.png.webp) ### The Retail Overload Problem Retailers today face a paradox: they have more data than ever, yet margins are thinner than ever. The biggest culprit? Over-assortment. Every new product SKU adds cost — procurement, storage, marketing, and shelf space. Yet only a fraction truly drive revenue. At Datahut, we’ve seen this repeatedly across retail and e-commerce: product counts balloon, complexity rises, and profits quietly erode. That’s where data-driven range rationalization comes in — the process of refining your assortment using internal and competitor data to strengthen your unit economics. As Harvard Business Review notes,[ ](https://hbr.org/2012/11/which-products-should-you-stock?ref=blog.datahut.co)[“getting product assortment right isn’t easy, yet it’s absolutely critical to retail success”](https://hbr.org/2012/11/which-products-should-you-stock?ref=blog.datahut.co) — which is exactly why analytics and web-scraped competitor insights now sit at the heart of modern retail strategy. ### Why Range Rationalization Matters Most retailers assume adding more products equals more choice and happier customers. In reality, it often means higher costs, cluttered shelves, and confused shoppers. Each extra SKU introduces: - Additional supply-chain handling - Increased working capital lock-up - Forecasting and replenishment complexity - Marketing dilution When only 20 % of SKUs drive 80 % of profits, trimming the long tail becomes the smartest way to boost margins without raising prices. A McKinsey case study on analytical assortment optimization found that a grocer cut SKUs by 36 % and lifted sales and margins by 1–2 %, showing how KPI-driven delisting directly improves unit economics. ### The Profitability × Customer Commitment Lens At Datahut, we encourage retailers to evaluate every SKU across two axes: ![At Datahut, we encourage retailers to evaluate every SKU across two axes:](https://www.blog.datahut.co/content/images/2026/07/img-406.png.webp) This framework clarifies which SKUs deserve your attention — and which don’t. McKinsey’s[ ](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/solutions/periscope/solutions/category-solutions/assortment-optimization?ref=blog.datahut.co)[Periscope assortment optimization brief](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/solutions/periscope/solutions/category-solutions/assortment-optimization?ref=blog.datahut.co) calls this approach “identifying which categories are under- or over-represented to optimize the assortment and increase sales.” ### How Competitor Data Strengthens Range Decisions The smartest retailers don’t make assortment decisions in isolation. They combine internal sales data with competitor-scraped data to see what’s working across the market. With retail data scraping and competitive assortment analysis, you can uncover insights such as: #### 1\. SKU Density & Category Mapping Web scraping lets you map how competitors distribute SKUs across categories — for instance, how many “premium organic snacks” a rival carries versus “everyday snacks.” This helps identify whether your range is over- or under-represented in profitable categories, echoing McKinsey’s advice that[ ](https://www.mckinsey.com/industries/industrials-and-electronics/our-insights/category-managements-next-horizon-how-distributors-can-outperform?ref=blog.datahut.co)[“category managers today should build strategies based on customers’ needs and willingness to pay”](https://www.mckinsey.com/industries/industrials-and-electronics/our-insights/category-managements-next-horizon-how-distributors-can-outperform?ref=blog.datahut.co). #### 2\. Price-to-Value Positioning By collecting real-time competitor pricing, packaging, and size data, you can benchmark your price-value equation. However, as HBR warns in its[ ](https://hbr.org/2023/11/a-step-by-step-guide-to-real-time-pricing?ref=blog.datahut.co)[guide to real-time pricing](https://hbr.org/2023/11/a-step-by-step-guide-to-real-time-pricing?ref=blog.datahut.co), “retailers that use simple heuristics like scraping lowest competitor prices miss significant opportunities.” The takeaway: intelligent pricing models must blend competitor data with demand signals, not chase the cheapest tag. #### 3\. Discount & Promotion Patterns Web-scraped data reveals how often competitors discount products and by how much. Platforms like[ ](https://www.bcg.com/x/product-library/merch-ai-solutions?ref=blog.datahut.co)[BCG X’s Merch AI](https://www.bcg.com/x/product-library/merch-ai-solutions?ref=blog.datahut.co) show how AI tools synthesize such internal + external data to create “customer-centric assortments and cost advantage,” often lifting margins 1–3 %. #### 4\. Stock Availability & Demand Signals Tracking competitor out-of-stock (OOS) events offers a goldmine of demand insights. If rivals repeatedly run out of a variant, it signals unmet demand — a theme echoed in[ ](https://www.walmartdataventures.com/content/walmartdataventures/en%5Fus/insights/articles/enhanced-assortment-deep-dive.html/?ref=blog.datahut.co)[Walmart Data Ventures’ Assortment Deep Dive](https://www.walmartdataventures.com/content/walmartdataventures/en%5Fus/insights/articles/enhanced-assortment-deep-dive.html/?ref=blog.datahut.co), which stresses building “customer-centric assortments by gaining insights into shopper purchasing behaviors.” #### 5\. New Product Velocity Competitor monitoring shows how fast others launch or retire SKUs — critical for innovation pacing. McKinsey’s[ ](https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/how-to-predict-your-competitors-next-move?ref=blog.datahut.co)[“How to Predict Your Competitor’s Next Move”](https://www.mckinsey.com/capabilities/strategy-and-corporate-finance/our-insights/how-to-predict-your-competitors-next-move?ref=blog.datahut.co) advises focusing on one or two competitors’ pricing and product portfolios to anticipate actions instead of spreading efforts too thin. When stitched together, these insights help retailers improve unit economics by reducing waste, focusing on profitable categories, and aligning pricing with real market behavior. ### Web Scraping: The Data-First Advantage Modern web scraping for retail isn’t about collecting random data — it’s about structured, compliant, and high-fidelity extraction that powers decision-making. At Datahut, our infrastructure enables: - Automated capture of competitor product, pricing, and availability data across thousands of SKUs daily - Extraction of category-level intelligence such as product titles, images, variants, and metadata - Integration of clean, structured datasets into BI tools, pricing engines, or inventory models - Continuous monitoring of market shifts for pricing, promotions, and stock changes Deloitte describes similar AI agents in retail as “digital workers that ingest sales trends, customer browsing behavior, and competitors’ movements in real time”, showing how automated data feeds enable dynamic assortment tuning. Likewise, NRF reports that[ ](https://nrf.com/blog/what-can-ai-do-retailers-and-customers?ref=blog.datahut.co)[AI assistants can become “experts in product assortment” with deep semantic understanding of every product](https://nrf.com/blog/what-can-ai-do-retailers-and-customers?ref=blog.datahut.co), reinforcing the power of automated intelligence in merchandising. ### Unit Economics: The Data-First Path Optimizing unit economics isn’t just about cost cuts — it’s about improving contribution margins through smarter assortment. McKinsey’s Online Grocery Fulfillment study notes that[ ](https://www.mckinsey.com/industries/retail/our-insights/achieving-profitable-online-grocery-order-fulfillment?ref=blog.datahut.co)[“the challenge boils down to unit economics”](https://www.mckinsey.com/industries/retail/our-insights/achieving-profitable-online-grocery-order-fulfillment?ref=blog.datahut.co) — a $100 basket can swing from +$4 in-store profit to –$13 online once logistics costs hit. That’s why assortment discipline and competitor benchmarking matter. Datahut’s retail data extraction and pricing-intelligence solutions help brands: - Collect real-time competitor pricing and product data - Analyze SKU-level profitability - Identify assortment gaps and redundant SKUs - Feed structured data directly into pricing or inventory models Bain & Company similarly points out that even Amazon’s cashier-less prototype “is likely to face challenging unit economics” ([Bain Insight](https://www.bain.com/insights/retail-holiday-newsletter-2016-2017-4/?ref=blog.datahut.co)), underscoring that innovation still hinges on profitability per unit sold. Meanwhile,[ ](https://newsroom.ibm.com/2024-01-15-NRF-2024-IBM-Reports-Generative-AI-Can-Bridge-the-Consumer-Expectation-Gap-with-Unified,-Integrated-Shopping-Experiences?ref=blog.datahut.co)[IBM’s 2024 NRF report](https://newsroom.ibm.com/2024-01-15-NRF-2024-IBM-Reports-Generative-AI-Can-Bridge-the-Consumer-Expectation-Gap-with-Unified,-Integrated-Shopping-Experiences?ref=blog.datahut.co) shows AI tools now “optimize store-level assortments” by analyzing each store’s mix to maximize sales — the future of unit-economic optimization in action. ![More SKUS= More Profit](https://www.blog.datahut.co/content/images/2026/07/img-88.jpg.webp) ### AI and Algorithmic Retailing Retail is entering the age of Algo Retailing, where algorithms balance customer need and profitability at every SKU. As MIT Sloan Management Review explains, AI techniques use large data sources “to decide on store-level assortment and drive deep localization at scale” (MIT Sloan). Kroger’s[ ](https://www.8451.com/knowledge-hub/technology/optimizing-store-space-with-machine-learning/?ref=blog.datahut.co)[84.51° case study](https://www.8451.com/knowledge-hub/technology/optimizing-store-space-with-machine-learning/?ref=blog.datahut.co) illustrates this in practice: its ML-driven planogram optimizer predicts category sales for each layout, yielding about $18 million in extra annual sales through better assortment and shelf allocation. These examples prove that web-scraped competitor data, when fused with AI and machine learning, can turn retail complexity into measurable profit. ### The Courage to Simplify The hardest part of assortment optimization isn’t the analysis — it’s the courage to act. Retailers often know which products underperform but hesitate to delist them. Simplification isn’t shrinkage — it’s strategy. As McKinsey emphasizes in its category-management research, success means “building strategies based on customer needs and willingness to pay” — not on what suppliers push. Or as Walmart Data Ventures puts it, the goal is to “build a customer-centric assortment by gaining insights into shopper purchasing behaviors.” With clean, compliant, and competitive web data guiding decisions, retailers can finally move from intuition-based merchandising to intelligent assortment design. ### Final Thought In modern retail, success isn’t defined by how much you sell — but by how efficiently you sell it. Using competitor data scraping, AI-driven assortment analytics, and data-first decisioning, you can build a product range that’s lean, profitable, and aligned with customer demand. If you’re ready to uncover what to keep, cut, or create —Datahut can help you get there. FAQs FAQ 1: What is Nestlé’s Project TASTY? Project TASTY is Nestlé’s global cost-efficiency and portfolio simplification initiative launched in 2021\. Its goal is to reduce operational complexity, streamline the company’s vast product assortment (over 100,000 SKUs), and drive structural savings—all without compromising product quality or taste (hence the name TASTY). The program focuses on eliminating underperforming SKUs and optimizing recipes and packaging, helping Nestlé improve service levels, cut waste, and reinvest in core brands and innovation. FAQ 2: How many SKUs did Project TASTY target, and why? Nestlé revealed that about one-third of its SKUs accounted for only 1% of total sales—a classic long-tail problem. Project TASTY was designed to “cut the tail” by removing thousands of low-rotation, low-margin, or duplicative SKUs that were consuming disproportionate resources. The rationale was to free up supply chain capacity, improve in-stock rates for high-performing products, and simplify manufacturing, logistics, and marketing operations globally. FAQ 3: What results has Project TASTY delivered? Since its rollout in 2021, Project TASTY has generated over CHF 2 billion in cumulative savings (as of early 2025), helping Nestlé mitigate severe cost inflation and protect profit margins. SKU reductions and recipe/packaging optimizations boosted supply chain efficiency and improved product availability—raising shelf availability of top-selling items from \~95% toward a 99% target. Nestlé also saw organic growth improve after cutting low-value SKUs, with executives calling the initiative “a significant boost” to their growth model. FAQ 4: What is the Profitability × Customer Commitment matrix, and how does it help retailers optimize their assortments? The Profitability × Customer Commitment matrix is a two-axis model that helps retailers classify SKUs based on their financial contribution and customer relevance. Each SKU is evaluated across: - Profitability: How much margin it generates, considering real costs. - Customer Commitment: How essential it is to target shoppers (e.g., loyalty, frequency of purchase, substitution risk). This framework creates four clear quadrants — star performers, margin drainers, overlooked gems, and dead weight — enabling data-driven decisions to invest, optimize, promote, or delist SKUs. It transforms assortment planning from intuition-driven to margin-optimized. FAQ 5: How does competitor data improve assortment and pricing decisions? Retailers that rely only on internal data risk missing market context. By scraping structured competitor data — including SKUs, pricing, promotions, stock levels, and launch velocity — retailers can: - Benchmark assortment density by category. - Identify pricing gaps or opportunities. - Detect emerging trends (e.g., frequent OOS events signaling unmet demand). - Monitor innovation cycles and promotional aggression. This external intelligence, when fused with internal performance data, helps build customer-centric, high-margin assortments aligned with real-time market dynamics. FAQ 6: Why is simplifying the assortment critical to improving unit economics? Over-assorted portfolios often include low-velocity SKUs that drain working capital, complicate operations, and dilute brand focus. Simplifying the assortment — based on both profitability and customer value — improves unit economics by: - Reducing warehousing and supply-chain complexity. - Increasing shelf availability for top-performing SKUs. - Lowering per-unit costs through higher production frequency. - Enhancing pricing power by focusing on high-demand products. This isn’t about cutting for the sake of cutting — it’s about designing a leaner, smarter range that delivers better margins and stronger customer relevance. Related posts ### How Data-Driven Storytelling Builds Brand Trust URL: https://www.blog.datahut.co/post/how-data-driven-storytelling-builds-brand-trust-and-purpose-in-2025/ Last updated: 2026-09-07T09:44:02.000Z ## Introduction – The Shift from Marketing to Meaning In 2025, the most trusted [brands](https://www.blog.datahut.co/post/web-scraping-vs-api/) aren’t just the ones with the best products, they’re the ones with the most transparent stories. And often, those stories begin with data. Modern consumers no longer buy products; they buy into values. [Edelman’s 2024 ](https://www.edelman.com/trust/2024/trust-barometer?ref=blog.datahut.co)[Trust Barometer](https://www.edelman.com/trust/2024/trust-barometer?ref=blog.datahut.co)[ ](https://www.edelman.com/trust/2024/trust-barometer?ref=blog.datahut.co)revealed that 68% of global consumers make buying decisions based on shared beliefs and trust, not just price or convenience. People expect brands to reflect their values in behavior, not words. They want measurable action — sustainability goals, ethical sourcing, fair treatment — and they want those claims backed by facts. Meanwhile, the rise of digital platforms and open data has shifted the landscape. Product and operational data are no longer private assets — they are public touchpoints of integrity. This transformation marks a new phase in brand-building: marketing has evolved from crafting messages to proving meaning. ## The Rise of Data-Driven Storytelling Data-driven storytelling merges analytics and authenticity. It turns raw data into emotional, value-driven narratives that connect with consumers who increasingly scrutinize every claim. In a [Deloitte 2025 Global Marketing Trends report](https://www.deloitte.com/global/en/about/press-room/deloitte-globals-2025-predictions-report.html?ref=blog.datahut.co), 73% of consumers said they are more likely to engage with brands that “show how their products are made.” Transparent storytelling rooted in verifiable data satisfies this demand by turning behind-the-scenes operations into accessible, human-centered stories. Imagine a cosmetics brand showcasing lab-tested clean-ingredient data or a coffee company revealing its carbon-neutral supply chain. These aren’t marketing stunts; they are truth-telling mechanisms that give data emotional meaning. Marketers are now curators of proof. Real-time product details — from sourcing and pricing to user reviews and carbon metrics — become chapters in a narrative of accountability. This shift elevates storytelling from subjective creativity to measurable credibility. ![Data-Driven Storytelling](https://www.blog.datahut.co/content/images/2026/07/img-446.png.webp) ## Real-World Examples - Patagonia demonstrates how verified environmental data can become brand storytelling gold. Each product is connected to an interactive supply-chain map showing where materials originate. The company reports its sustainability metrics publicly and ties them to clear performance indicators. For Patagonia, transparency is not an afterthought — it’s the product. - Nike leverages performance, innovation, and athlete testing data to tell stories that fuse science with aspiration. Its product R&D insights fuel narratives about human achievement, illustrating that every gram reduced or fabric enhanced is driven by measurable research. - Everlane pioneered “Radical Transparency,” publishing factory details, cost breakdowns, and sourcing impact. By sharing actual cost-to-consumer data, the company reframed what honesty in fashion looks like and built enduring trust with ethically-minded shoppers. - Lenskart and ASOS collect and analyze customer-review data at scale to refine product messaging and address concerns proactively. ASOS’s transparency campaign, for instance, uses aggregated customer sentiment and quality insights to improve both product design and communication. ![Brands and storytelling- patagonia, everlane, nike, ASOS](https://www.blog.datahut.co/content/images/2026/07/img-447.png.webp) ## How Web Data Feeds Brand Purpose Internal metrics aren’t the only storytelling assets. Ethically collected web data—such as public product details, competitor benchmarks, and social sentiment—provides a richer, more contextualized foundation for brand narratives. According to [PwC’s ](https://www.pwc.com/us/en/services/consulting/business-transformation/library/2025-customer-experience-survey.html?ref=blog.datahut.co)[Consumer Intelligence Report](https://www.pwc.com/us/en/services/consulting/business-transformation/library/2025-customer-experience-survey.html?ref=blog.datahut.co)[, 87% of customers sa](https://www.pwc.com/us/en/services/consulting/business-transformation/library/2025-customer-experience-survey.html?ref=blog.datahut.co)y they research brand credibility via public information before making a purchase. That external scrutiny pushes companies to be more forthcoming. By scraping publicly available data responsibly, brands can evaluate industry transparency, identify differentiators, and detect market gaps. Imagine a sustainable footwear startup analyzing open product data from competitors to discover that consumers favor clear climate-footprint labeling. The brand could then feature verified emission data on its own packaging and website — transforming analytics into storytelling that resonates with real demand. “By analyzing publicly available product data, brands can identify which features resonate most with consumers — and turn those findings into value-driven stories.” Such insight-driven storytelling helps companies shift from brand-centric narratives to consumer-informed ones. It’s not about what the brand wants to say — it’s about what the data shows people care about. ## Building Trust Through Transparency Transparency isn’t merely a communication tool; it’s an economic asset. According to Edelman, 81% of consumers say they must trust a brand to purchase from it. That trust grows when brands demonstrate accountability using verifiable data. Transparency builds long-term resilience because it converts vulnerability into authenticity. When companies admit challenges — for instance, showing that supply-chain emissions are still being reduced — consumers perceive realism rather than weakness. This “show-your-work” approach mirrors academic honesty: every claim is supported by data, reinforcing credibility. [Deloitte’s 2025 Insights](https://www.deloitte.com/us/en/insights/industry/telecommunications/connectivity-mobile-trends-survey.html?ref=blog.datahut.co) on Ethical Consumption found that 63% of customers prefer brands that disclose their environmental or social challenges, as it signals genuine intent to improve. The implicit message: transparency is the new creative currency. Consider the difference between saying “we’re sustainable” and publishing a live environmental dashboard linked to product life cycles and supplier audits. The latter transforms transparency into an ongoing dialogue. Trust grows when audiences are treated as active participants in a brand’s journey. Data-driven storytelling invites that participation, encouraging customers to follow along, engage with updates, and hold brands accountable. ## The Road Ahead: Purpose-Driven Data Strategy As the line between marketing and ethics blurs, data openness will become a cornerstone of brand strategy. Rather than guarding product or supplier information, forward-thinking brands in 2025 and beyond will use it as a differentiation tool. Future consumers will expect real-time access to performance and sustainability metrics. Data from energy use, water consumption, ethical labor audits, and product life cycles will become public indicators of brand character. Transparency will be not just “nice to have” but the price of market entry. According to [Accenture Strategy’s latest Purpose Report](https://www.accenture.com/us-en/insights/consulting/empowered-consumer?ref=blog.datahut.co), brands that demonstrate measurable commitment to transparency experience up to 3.5x higher customer loyalty scores. That loyalty translates into tangible growth — proving that purpose and profit now move in parallel. For marketers, the call to action is clear: treat data not just as an analytical tool but as a storytelling voice. Develop purpose-driven data strategies that unite marketing, sustainability, and data science functions. Empower teams to translate numbers into narratives that express human impact. The real challenge ahead isn’t [collecting more data](https://www.blog.datahut.co/post/why-retailers-should-invest-in-web-scraping-product-matching-and-bi/) — it’s communicating it clearly, [ethically](https://www.blog.datahut.co/post/is-web-scraping-legal/), and empathetically. Brands that master this will become trusted interpreters of their own truth. ## Conclusion Product data is no longer confined to back-end systems or pricing algorithms — it’s the foundation of how purpose and trust are built in the digital age. By bringing transparency, ethics, and emotion into harmony, data transforms from a cold metric into a living story. The brands of tomorrow will be those that can show — not just say — why they exist, what they value, and how they act. In this new era of [data-driven storytelling](https://www.blog.datahut.co/post/what-web-scraping-can-and-cant-do-for-you/), truth has become the most valuable brand asset. When shared openly, it not only inspires consumers but also holds companies accountable to the values they promote. ![datahut's data solutions ](https://www.blog.datahut.co/content/images/2026/07/img-448.png.webp) To explore how [Datahut](https://www.datahut.co/?ref=blog.datahut.co)’s data solutions can help your organization uncover, structure, and use insights to tell honest, trust-building stories, start your purpose-driven transformation today. ### FAQs 1\. What is data-driven storytelling and why is it important for brands in 2025? Data-driven storytelling is the practice of using verified data to create authentic, value-based brand narratives. In 2025, it’s essential because consumers demand transparency — they trust brands that back their claims with facts and measurable actions. 2\. How does data transparency build brand trust? Transparency builds trust by proving accountability. When brands openly share data about their sourcing, sustainability, or performance, they show honesty and credibility — qualities that strengthen long-term consumer relationships. 3\. What are examples of brands using data-driven storytelling successfully? Brands like Patagonia, Nike, Everlane, ASOS, and Lenskart use data storytelling effectively. They disclose metrics such as supply-chain origins, carbon impact, or customer review insights to engage audiences with facts, not fluff. 4\. How can web data enhance brand purpose and storytelling? Ethically sourced web data provides real-world context for brand decisions. It helps brands benchmark competitors, analyze consumer sentiment, and discover new transparency opportunities — transforming analytics into meaningful stories. 5\. How can businesses start using data to build brand trust? Start by identifying what data aligns with your brand values — such as ethical sourcing or sustainability metrics — and share it openly through dashboards, reports, or campaigns. Partnering with data experts like Datahut can help structure and communicate those insights effectively. ### GDPR Enforcement Tracker: Monitor GDPR Fines & Cases URL: https://www.blog.datahut.co/post/how-to-use-web-scraping-to-track-gdpr-fines-and-enforcement-cases/ Last updated: 2026-09-07T09:44:04.000Z Why do you need to scrape this data in the first place? Well take a look at the top [10 GDPR fines](https://www.blog.datahut.co/post/top-10-gdpr-fines-in-2025-a-data-driven-analysis/) . You'll be relived you have a resource like this In today’s digital world, organizations handle massive amounts of personal data, and protecting that information has become a serious responsibility. Data protection is no longer just a [legal ](https://www.blog.datahut.co/post/web-scraping-e-commerce-websites-top-five-legal-battles-and-learnings/)phrase tucked away in regulations - it has become a practical necessity for organizations in today’s digital world. With the introduction of the General Data Protection Regulation (GDPR) in 2018, companies across Europe and beyond have been held to higher standards in handling personal data. When these standards are not met, the consequences often come in the form of fines and penalties. Tracking these enforcement actions is made easier by the [GDPR Enforcement Tracker](https://www.enforcementtracker.com/?ref=blog.datahut.co), an online database maintained by CMS (CMS Law). It collects and organizes details of fines issued by data protection authorities across the European Union (EU) and the European Economic Area (EEA). Instead of just learning GDPR theory, the tracker shows how the law is applied in real situations - highlighting real cases, real penalties, and real lessons. Each case tells a story: which country issued the fine, the type of violation, the amount imposed, and the GDPR article involved. Fines range from smaller penalties against individuals or local organizations to multimillion-euro actions against global tech companies. This makes the tracker a valuable resource for understanding enforcement patterns, seeing where regulators are most active, and identifying common compliance mistakes. While not every fine is made public, the database provides a structured and regularly updated overview. Its columns—ETid, Country, Date of Decision, Fine, Controller or Processor, Quoted Article, Type, and Source—turn individual cases into a dataset that can reveal broader [insights](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/). In this blog, we will explore how to turn this rich online resource into a usable dataset. From [scraping](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/) the data, handling missing values, to cleaning and structuring it for analysis, this step-by-step guide will demonstrate how even a large, multi-page table of enforcement cases can be transformed into actionable information. By the end, readers will understand not only the mechanics of data extraction but also the insights this dataset can reveal about GDPR enforcement across Europe. ## Smarter Way to Collect Data When dealing with large websites filled with structured information, manually copying data is simply not practical. Imagine scrolling through hundreds of enforcement cases, each listing details like the country, the decision date, the fine amount, and the GDPR article cited. Doing this by hand would take hours and still risk mistakes. This is where web scraping offers a far smarter and more reliable way forward. The GDPR Enforcement Tracker website is a good example. It publishes detailed records of data protection enforcement across Europe, but the information sits in long tables spread across multiple pages. Instead of treating it like just another static webpage, a scraper can transform it into a structured dataset. Every row of the table—containing the case ID, the organization involved, the penalty imposed, and even the link to the official source—gets extracted automatically. ### Step 1: Collecting Data from a Paginated Table The only step in this scraping project is figuring out where the data lives and how it is structured. In this case, the GDPR Enforcement Tracker website presents information in a tabular format. Each row of the table represents an enforcement case, containing details such as the case ID, the country, the date of the decision, the fine amount, the type of violation, and even a link to the original source. Now, the table isn’t just a single page—it actually holds thousands of entries (2,839 in this case). To make things manageable, the site uses pagination: only a portion of the data is visible at once, and a “Next” button at the bottom of the page lets you move forward. While this makes browsing easier for humans, it adds an extra layer of complexity for a scraper. The scraper not only has to read the data from the current page but also click “Next” repeatedly until all entries are collected. Think of it like flipping through a big book with multiple chapters. You don’t want to read just the first chapter—you want the entire story. In the same way, the scraper patiently goes through every page of the table, ensuring no case is left behind. Once this step is complete, you have a comprehensive list of all enforcement cases ready to be stored and analyzed. ### Step 2: Cleaning the Data Once all the rows have been scraped and stored, the next challenge is ensuring the data is clean and reliable. Raw data from websites often looks neat at first glance, but when you dig deeper, you’ll notice small issues—missing values, inconsistent formatting, or incomplete entries. These problems may seem minor, but they can make analysis difficult later on. In the GDPR Enforcement Tracker dataset, most of the values were well-structured because the source itself is a clean table. However, a few columns had gaps. For example, some cases did not list the exact fine amount, while others had the quoted GDPR article marked as “Unknown”. These missing or unclear values needed attention before moving ahead. So fill the missing values with “N/A”. For this task, [OpenRefine](https://openrefine.org/?ref=blog.datahut.co) proved handy. It’s a beginner-friendly tool designed for cleaning and organizing messy data. Think of it like a workshop table where you can polish and fix raw pieces before building something useful. With OpenRefine, it was easy to standardize entries and handle missing values. Of course, when dealing with much larger files or more complex cleaning tasks, Python libraries like pandas are often better suited. Pandas lets you programmatically check for gaps, replace missing values, and even reformat entire columns with just a few lines of code. But for the current dataset, OpenRefine did the job effectively. In short, this step ensures the dataset is trustworthy. Clean data means fewer headaches later, whether the goal is statistical analysis, visualization, or sharing insights with others. ## Essential Tools for Smooth and Efficient Data Extraction When it comes to collecting structured data from websites, the right combination of tools can make the process smooth, reliable, and surprisingly fast. Think of it as building a small team where each member has a specific role: some handle browsing, others extract information, and a few organize and store the results. In this setup, [Playwright](https://playwright.dev/?ref=blog.datahut.co) takes the lead. It’s a Python library that can open web pages, navigate through links, click buttons, and even scroll tables—just like a human user would. By automating these actions, Playwright removes the need to manually copy data, which can be tedious and error-prone, especially when the website has thousands of rows or multiple pages. To store the data efficiently, sqlite3 is used. It’s a lightweight database that organizes information neatly without the need for a separate server. Each scraped row—like case ID, country, fine amount, and source link—can be saved safely and retrieved whenever needed. Meanwhile, logging keeps track of every action and any errors that occur, providing a clear trail of what the scraper did and helping debug issues if something goes wrong. Other supporting tools also play key roles. The time module introduces small pauses to ensure web pages load properly, re helps clean or extract specific patterns from text, and urljoin ensures all links are converted into absolute URLs, so nothing gets lost during navigation. Altogether, this combination forms a resilient and organized toolkit for web scraping. Each library contributes its strength—automation, data storage, logging, or text processing—allowing even large-scale extraction from dynamic, paginated websites to be handled efficiently and accurately. Using these tools together transforms a static webpage into a structured, analyzable dataset ready for deeper insights. ### Importing the Right Tools Before we dive into any kind of scraping or automation, we first need to set up the tools that will help us do the job. Think of this step as gathering all the ingredients before you start cooking. If something is missing, you’ll get stuck halfway through the recipe. In Python, these “ingredients” come in the form of libraries. Each library brings its own special abilities to the table, and together, they make our project possible. Let’s walk through the ones we’re using here and why they matter. ``` # Importing libraries from playwright.sync_api import sync_playwright import sqlite3 import logging import time import re from urllib.parse import urljoin """Dependencies: - Playwright: For browser automation. - sqlite3: For local database management. - logging: For logging actions and errors. - time & urllib.parse: For delay handling and URL joining. - re: Regular expressions (optional usage).""" ``` The first one, Playwright, is the star of our setup. It allows us to control a browser automatically, almost like a robot clicking buttons and navigating pages for us. With Playwright, we can open websites, scroll through them, and capture information—without having to do everything manually. Next comes sqlite3, which is a built-in library in Python. Think of it as a small notebook where we can neatly store all the data we collect. Instead of having everything scattered in memory, we’ll keep it safe in a local database file that we can query later. The logging library is like our personal diary for this project. It keeps track of what happens while the code runs—whether things are going smoothly or if any errors pop up. This makes troubleshooting much easier, especially when the project grows bigger. Then we have time, which we’ll use to add delays. Sometimes, when scraping, rushing through pages too quickly can raise suspicion or even block us from accessing the site. Adding small pauses makes our automation look more natural. The re library handles regular expressions. You can think of it as a magnifying glass we use to spot patterns in text. For example, if we want to pick out numbers or clean up messy strings, regex helps us do it. Lastly, urljoin from urllib.parse comes into play when we need to handle web links. Websites often provide relative URLs, and this tool helps us turn them into complete, usable links. By combining all these libraries, we get a powerful yet lightweight toolkit: Playwright for automation, sqlite3 for storage, logging for tracking, time for delays, re for pattern matching, and urljoin for building links. With these in place, we’re ready to move on to the next stage of our project. ### Setting Up Logging Once we have our tools imported, the next important step is to make sure we keep track of what our program is doing. Imagine running a long scraping script overnight and waking up in the morning only to see it crashed halfway through. Without a record of what happened, it would feel like trying to solve a mystery with no clues. This is exactly why logging is so useful. In Python, logging acts like a journal for your program. Every time something important happens—whether it’s a success, a warning, or an error—it writes an entry into a log file. Later, when you want to understand how your scraper behaved, you can just read through that file. ``` # Logging Setup logging.basicConfig(    filename="gdpr_scraper.log",    level=logging.INFO,    format="%(asctime)s - %(levelname)s - %(message)s" ) """ Logging Setup Configuration This section configures the logging module to record events and errors during the scraping process. - Logs will be saved to the file named "gdpr_scraper.log". - Logging level is set to INFO, meaning that INFO, WARNING, ERROR, and CRITICAL messages will be recorded. - Log messages will include:    - Timestamp of the log entry (when the event occurred).    - Logging level (INFO, ERROR, etc.).    - The actual message describing the event or error. Purpose: The log file helps in tracking the progress of the scraper and diagnosing any issues that occur during execution. For example, each time a row is successfully saved or an error happens during scraping, it will be logged.""" ``` With this configuration, all log messages will be saved into a file called gdpr\_scraper.log. This file acts like a black box recorder for your scraper—it captures everything you tell it to. We’ve set the logging level to INFO, which means the file will record not just important errors, but also general progress updates. For example, if the scraper successfully saves a new case to the database, we’ll log that. If it runs into an error, that too will be logged. The format parameter makes sure each log entry contains three pieces of information: 1. When it happened (the timestamp). 2. What type of event it was (INFO, WARNING, ERROR, etc.). 3. The actual message describing what occurred. Together, these details make debugging much easier. Instead of staring at your code wondering why it broke, you can open the log file and retrace exactly what happened step by step. So in short, while logging might look like a small setup, it’s actually one of the most important parts of any scraping project. It helps us stay in control, even when the script is running on its own. ### Defining Constants and Preparing the Database Now that we’ve set up our logging system, the next step is to decide where our scraper will look and where it will store the information it collects. Think of this as setting the destination before starting a journey, and also keeping a notebook ready to record everything you find along the way. ``` # Constants & DB Setup URL = "https://www.enforcementtracker.com/" DB_NAME = "gdpr_enforcement_tracker1.db" TABLE_NAME = "gdpr_cases" """ Constants used for scraping and database setup: - URL: Source website for GDPR enforcement tracker data. - DB_NAME: Name of the SQLite database file. - TABLE_NAME: Name of the table where case data will be stored. """ ``` In this part, we set three constants to guide our scraper. - URL tells the scraper where to look for data—the GDPR Enforcement Tracker website. - DB\_NAME is the name of the local SQLite file where we’ll save everything, like a notebook for our data. - TABLE\_NAME is the section inside that notebook where the cases will be stored. Defining them at the top makes the code cleaner and easier to change later if needed. ### Creating the Database Table Once we know where to store our data, the next step is to prepare the actual space inside the database. Think of it like setting up an empty table in a notebook before you start writing in it. Without that table, you wouldn’t know where each piece of information should go. ``` # Database Functions  def create_table():    """Create the SQLite table to store GDPR enforcement cases if it doesn't exist.      1. Connects to the SQLite database specified by DB_NAME.       2. Executes a SQL command to create a table named TABLE_NAME with specified columns if it doesn't already exist.      3. Commits the changes and closes the connection    Table Columns:    - ETid: Unique identifier for the enforcement decision.    - Country: Country where the enforcement was issued.    - DateOfDecision: Date when the decision was made.    - Fine: The fine imposed in the case.    - ControllerProcessor: Name of the controller or processor involved.    - QuotedArt: Quoted GDPR Article relevant to the decision.    - Type: Type of enforcement decision.    - Source: URL of the source document or reference.    """    conn = sqlite3.connect(DB_NAME)    conn.execute(f"""        CREATE TABLE IF NOT EXISTS {TABLE_NAME} (            ETid TEXT,            Country TEXT,            DateOfDecision TEXT,            Fine TEXT,            ControllerProcessor TEXT,            QuotedArt TEXT,            Type TEXT,            Source TEXT        )    """)    conn.commit()    conn.close() ``` The function create\_table() takes care of this preparation. Here’s what it does step by step: 1. It connects to the database file we defined earlier (gdpr\_enforcement\_tracker1.db). If the file doesn’t exist yet, SQLite will quietly create it for us. 2. It runs a command that says: “Make me a table called gdpr\_cases with the following columns—ETid, Country, DateOfDecision, Fine, ControllerProcessor, QuotedArt, Type, and Source.” 3. Each column has a clear purpose. For example, ETid is like a unique ID for the case, Fine records the penalty, and Source points to the official document or link. 4. Finally, it saves the changes and closes the connection to the database. By doing this, we’ve set up a neat structure that ensures all the scraped data will land in the right place. If we skip this step, the scraper would have nowhere to save its findings, much like trying to write notes without having a page ready. ### Saving Scraped Data into the Database Once our scraper starts collecting information, the next challenge is deciding how to store it safely. Imagine going to a library, taking notes on dozens of books, but then forgetting to file those notes properly—you’d end up with a mess. That’s why we need a dedicated function to neatly save each piece of data into our database. ``` # Scraping Functions  def save_row_to_sqlite(row):    """    Save a single scraped row into the SQLite database.       Args:        row (list): List containing values for the enforcement decision:            [ETid, Country, DateOfDecision, Fine, ControllerProcessor, QuotedArt, Type, Source]       Logs the saved row ETid for traceability.    """    conn = sqlite3.connect(DB_NAME)    conn.execute(f"""        INSERT INTO {TABLE_NAME}        (ETid, Country, DateOfDecision, Fine, ControllerProcessor, QuotedArt, Type, Source)        VALUES (?, ?, ?, ?, ?, ?, ?, ?)    """, row)    conn.commit()    conn.close()    logging.info(f"Saved row: {row[0]} | source: {row[7]}") ``` The function save\_row\_to\_sqlite() does exactly this. Each row of scraped data, which contains details like the case ID, country, fine amount, and source link, is passed into this function. The function then opens a connection to our SQLite database, inserts the row into the correct table, and closes the connection once the job is done. What makes this step especially useful is the logging. Every time a row is successfully saved, the scraper records the event in the log file with the case ID and source link. Think of this like a “receipt” for each entry—it helps us trace what has already been stored and makes it easier to debug if something goes wrong. By separating the saving process into its own function, we also make the code cleaner and more reusable. Anytime we want to save new rows, we don’t have to rewrite the database commands—we just call this function. It’s like setting up a reusable storage box where every new record automatically finds its place. ### Extracting the Source Link from Each Case When we scrape the GDPR Enforcement Tracker website, each row in the table represents a case. Along with details like the fine amount or the country, there’s often a source link—a URL pointing to the official decision or a related document. Capturing this link is important, because it allows us to trace back to the original source for verification or deeper reading. ``` # Scraper Logic  def extract_source_from_tr(tr, etid_href=None):    """    Extract the source URL from a table row element representing a GDPR enforcement case.    This function attempts to find a relevant hyperlink (source document or reference)    associated with the enforcement case in a table row (`tr`) by following these steps:    1. First, it looks for an anchor () tag with the class 'blau', which is typically       used for the primary source link on the webpage.       - If found, it processes the href attribute:           - Converts protocol-relative URLs (starting with "//") to absolute by adding "https:".           - Converts relative URLs (starting with "/") into absolute URLs using the base URL.           - Returns the absolute URL directly if it’s already complete.    2. If no 'blau' class anchor is found, the function searches through all anchor () tags in the row.       - It skips the ETid link (provided as `etid_href`) to avoid using it as the source.       - It returns the first link that:           - Starts with "http" (absolute URL), or           - Ends with ".pdf" (document file link).    3. If no suitable link is found, it returns an empty string.    Args:        tr (ElementHandle): A Playwright ElementHandle representing one table row () of the GDPR enforcement cases table.        etid_href (str, optional): The ETid hyperlink (href) to skip, typically the link to the case itself.    Returns:        str: The absolute URL of the source document or reference if found, otherwise an empty string.    Notes:        - Uses urljoin to handle relative URLs.        - Logs any unexpected errors during extraction but continues execution.    """    try:        # Prefer anchor with class 'blau'        a = tr.query_selector("a.blau")        if a:            href = a.get_attribute("href") or ""            href = href.strip()            if href:                if href.startswith("//"):                    return "https:" + href                if href.startswith("/"):                    return urljoin(URL, href)                return href        # Fallback: first anchor that isn’t the ETid link        anchors = tr.query_selector_all("a")        for link in anchors:            href = (link.get_attribute("href") or "").strip()            if not href or href == etid_href:                continue            if href.startswith("http") or href.endswith(".pdf"):                return href    except Exception as e:        logging.error(f"extract_source_from_tr error: {e}")    return "" ``` That’s where the function extract\_source\_from\_tr() comes in. Its job is simple but essential: look at a single table row on the webpage and try to find the right link to save. Here’s how it works. First, the function checks if the row has a special link marked with the CSS class "blau". On this website, that usually means the main source link. If such a link exists, the function carefully examines its format. Some links may be written in shorthand (like starting with // or /), so the function fixes them by adding the base website address to make them full, usable URLs. If the link is already complete, it just returns it as is. But what if the "blau" link isn’t there? In that case, the function doesn’t give up. It goes through all the other links in the row, skipping over the case ID link (since that just points back to the case itself). From the remaining links, it picks the first one that looks valid—either a regular website link (starting with http) or a document link (ending with .pdf). And if nothing useful is found, the function simply returns an empty string. This way, the scraper doesn’t break; it just moves on, while also recording the error in the log file for later review. In plain terms, this function acts like a careful librarian. It goes through the details of each case, finds the most trustworthy source link, cleans it up if needed, and hands it over to be saved. Without this step, our scraper would have the basic case details but no direct trail to the official documents—which would be like having book summaries without being able to check the actual books. ### Scraping GDPR Enforcement Cases from a Webpage Now that we have our database ready, let’s move on to the main part of the project: actually scraping the data from the Enforcement Tracker website. This is where our function scrape\_page(page) comes in. Think of this function as a worker that looks at one page of the site, reads through the table of cases, and carefully copies each row into our database. ``` # Main Scraper Logic  def scrape_page(page):    """Scrape table data from the current page, row by row, saving each row to DB    Scrape all GDPR enforcement cases from the current page of the enforcement tracker table.    This function performs the following actions:    1. Selects all rows from the penalties table on the current page.    2. For each row:        - Extracts the following information:            - ETid: The unique case identifier.            - Country: Country where the enforcement was issued (from an image 'alt' attribute).            - DateOfDecision: Date the decision was made.            - Fine: The fine amount imposed (as displayed).            - ControllerProcessor: Name of the controller or processor involved.            - QuotedArt: The GDPR article cited in the decision.            - Type: Type of the enforcement case.            - Source: Link to the source document or relevant reference.        - Handles missing data gracefully to avoid breaking the process.        - Saves the extracted row data into the SQLite database.    3. Logs the number of rows found and any errors encountered while processing rows.    Notes:        - Uses the helper function `extract_source_from_tr(tr, etid_href)` to extract the source URL.        - Each row is saved to the database using `save_row_to_sqlite(row)`.        - Any parsing or saving error for individual rows is logged, but the process continues for remaining rows.    """    table_rows = page.query_selector_all("table#penalties tbody tr")    logging.info(f"Found {len(table_rows)} rows on current page")    for idx, tr in enumerate(table_rows, start=1):        try:            tds = tr.query_selector_all("td")            # Map according to actual HTML structure you provided            etid_el = tds[1].query_selector("a") if len(tds) > 1 else None            etid = etid_el.inner_text().strip() if etid_el else ""            etid_href = (etid_el.get_attribute("href") or "").strip() if etid_el else ""            country = ""            if len(tds) > 2:                country_img = tds[2].query_selector("img")                if country_img:                    country = country_img.get_attribute("alt").strip()            date_decision = tds[4].inner_text().strip() if len(tds) > 4 else ""            fine = tds[5].inner_text().strip() if len(tds) > 5 else ""            controller = tds[6].inner_text().strip() if len(tds) > 6 else ""            quoted_art = tds[8].inner_text().strip() if len(tds) > 8 else ""            case_type = tds[9].inner_text().strip() if len(tds) > 9 else ""            source = extract_source_from_tr(tr, etid_href=etid_href)            row = [etid, country, date_decision, fine, controller, quoted_art, case_type, source]            save_row_to_sqlite(row)        except Exception as e:            logging.error(f"Error parsing/saving row {idx}: {e}") ``` The Enforcement Tracker website organizes GDPR penalty cases in a big table. Each row of this table is like a record card, holding information such as the case ID, the country where it happened, the fine amount, the law article quoted, and a source link. What our function does is simple: go row by row, collect all this information, and store it safely. When the function runs, the first thing it does is find all the rows inside the penalties table. Imagine you have a stack of papers, and you ask, “How many papers are in this stack?” That’s what the logging info line does—it tells us how many rows (or “papers”) were found on the current page. This is useful because it helps us track whether the website structure has changed or whether the scraper is working as expected. Next, the function goes through each row one at a time. For every row, it looks for specific pieces of information. For example: - The ETid, which is like a case number. - The country, which is taken from the little flag image shown in the table. - The date of decision, so we know when the ruling was made. - The fine amount, which is the headline figure most people care about. - The controller or processor, which is just the company or entity that got penalized. - The quoted GDPR article, which tells us which law they broke. - The type of case, describing what kind of violation it was. - And finally, the source link, which points to the original legal document or article. One nice thing about this function is that it has been written carefully to handle missing data. For example, sometimes a field may not exist in a row. Instead of crashing, the scraper just saves it as an empty string and keeps moving. This makes the process robust, like a person who doesn’t stop writing notes just because one page in a book has a smudge. After gathering all the details, the function saves the row into the database using another helper function called save\_row\_to\_sqlite(row). That way, every case we scrape is permanently stored and can be analyzed later without needing to scrape again. Of course, web scraping can sometimes be unpredictable—maybe a row has unusual formatting, or the page takes too long to load. That’s why the function also has error handling. If something goes wrong while scraping a particular row, it doesn’t stop the whole process. Instead, it logs the error and simply moves on to the next row. This ensures that one bad record doesn’t ruin the entire run. In short, scrape\_page(page) is like a careful note-taker: it goes through each enforcement case, copies down all the important details, and files them neatly into our database. This is the heart of the scraper, because without it, we wouldn’t have any data to analyze later. ### Bringing Everything Together: The Main Function Up until now, we have looked at how to scrape one page at a time and how to store each case in the database. But a scraper is only truly useful when it can run smoothly from start to finish—moving across multiple pages, collecting all the cases, and finally wrapping up neatly. That’s exactly what our main() function is designed to do. You can think of it as the “conductor of the orchestra,” making sure every part of the scraper plays in harmony. ``` # Main Execution def main():    """    Main function to run the GDPR Enforcement Tracker web scraper.    This function performs the following steps:       1. Initializes logging to track the progress and any errors during scraping.    2. Creates the SQLite database table if it doesn’t already exist, to store scraped data.    3. Launches a browser using Playwright in non-headless mode       (so you can see the browser actions for debugging or monitoring).    4. Opens a new browser page and navigates to the GDPR Enforcement Tracker website.    5. Waits for the penalties table to load before starting the scraping process.    6. Repeatedly scrapes all rows of the current penalties table and saves them into the database.    7. Checks for the "Next" button to move to the next page of results:       - If the "Next" button is disabled or not found, the scraper stops.       - Otherwise, it clicks "Next" and continues scraping the next page.    8. After reaching the last page, it closes the browser session.    9. Logs that the scraping process finished successfully.    Purpose:    This function automates data collection from a paginated GDPR enforcement table,    storing each case into a local SQLite database for analysis.    Notes:    - `headless=False` is used for visibility during development/debugging.    - A small time delay is added after loading each page to ensure the table is fully loaded.    - Errors during scraping are logged but do not stop the overall process.    """    logging.info("Starting GDPR Enforcement Tracker scraper")    create_table()    with sync_playwright() as p:        browser = p.chromium.launch(headless=False)        page = browser.new_page()        page.goto(URL)        page.wait_for_selector("table#penalties tbody tr")        time.sleep(2)        while True:            scrape_page(page)            # Check if "Next" button is disabled            next_button = page.query_selector("a#penalties_next")            if next_button is None or "disabled" in (next_button.get_attribute("class") or ""):                break            else:                next_button.click()                page.wait_for_selector("table#penalties tbody tr")                time.sleep(2)        browser.close()    logging.info("Scraper finished successfully.") ``` The first thing the function does is set up logging. Logging is like keeping a diary of what the program is doing. Every important step—like starting the scraper, saving rows, or finishing the run—is written down. If something goes wrong, these notes help us trace where the problem happened. Next, the function calls create\_table(). This ensures that the SQLite database is ready to receive data. If the table already exists, the function won’t overwrite it—it just makes sure everything is in place. Imagine preparing a filing cabinet before you start sorting papers into it. Once the database is ready, the scraper launches a browser using Playwright. Here, the browser is started in non-headless mode. That means you can actually see the browser window opening, loading pages, and clicking through results. This is especially helpful for beginners, because you can watch the scraper in action and confirm that it’s working as expected. Later, once you’re confident, you could switch to headless mode for faster and quieter runs. The browser then navigates to the GDPR Enforcement Tracker website, where all the penalty cases are listed. Before doing anything else, the function waits until the table of penalties has fully loaded. This pause is important—without it, the scraper might try to grab data before the page is ready, which would cause errors. A small delay is also added to give the page extra time to settle. Now comes the main loop. The scraper repeatedly calls the scrape\_page(page) function, which extracts the rows of data from the current page and saves them into the database. Once it finishes a page, it looks for the “Next” button. If the button is disabled or missing, that means we’ve reached the last page and the scraper can stop. Otherwise, it clicks the button, waits for the next page to load, and continues scraping. This cycle repeats until every single page of cases has been processed. Finally, when the last page is done, the browser closes and the scraper logs a message to say that everything finished successfully. At this point, all the scraped cases are neatly stored in the database, ready for analysis. In simple terms, the main() function is like a project manager: it sets up the workspace, launches the browser, makes sure each page is scraped in order, and closes everything down once the job is complete. Without it, our scraper would just be a collection of disconnected parts. With it, the whole process runs from start to finish, hands-free. ### The Entry Point of the Script Every program needs a clear starting point, a place where the instructions begin. In our scraper, that role is played by the following block of code: ``` # Entry Point  if name == "__main__":    main() """ Entry Point of the Script This block ensures that the script’s main function runs only when the script is executed directly, and not when it is imported as a module in another script. - `if name == "__main__":`    This condition checks if the script is being run as the main program.    - If true, it calls the `main()` function to start the scraper.    - If the script is imported elsewhere, this block is skipped,      so the `main()` function does not run automatically. """ ``` At first glance, this might look a little mysterious. But here’s what it really means. When Python runs a file, it sets a special variable called name. If the file is being run directly (like typing python scraper py in the terminal), this variable is set to "\_\_main\_\_". That’s Python’s way of saying, “This is the main script, go ahead and start from here.” So when the condition if name == "\_\_main\_\_": is true, the program calls the main() function, which in our case starts the entire scraping process. Why is this helpful? Because sometimes you may want to reuse parts of your script in another project. If you import this file into another Python script, the main() function will not run automatically. Instead, only the specific functions you call will execute. Think of it as a safety switch—it prevents the whole scraper from starting up unexpectedly when you only need one small piece of it. In simple terms, this block tells Python: “Only run the scraper if this file is opened directly. Otherwise, stay quiet.” It’s a neat little way to keep your code flexible and well-behaved. ## Conclusion Exploring the GDPR Enforcement Tracker through scraping reveals how structured data can turn a static webpage into a rich source of insights. By systematically collecting and cleaning each row of enforcement cases, it becomes possible to see which countries are most active in issuing fines, which types of violations occur most frequently, and how significant the penalties can be—from minor fines to multi-million-euro actions against major organizations. The dataset also highlights the most commonly cited GDPR articles and the variety of entities affected, giving a clear picture of real-world compliance challenges. This approach demonstrates that with careful data gathering and processing, complex regulatory information can be transformed into a format that is easier to analyze, understand, and learn from. ## Libraries and Versions Name: playwright , Version: 1.48.0 Name: sqlite3 (built-in, comes with Python, no separate versioning) Name: urllib (specifically urllib.parse, Built-in (standard Python library)) Name: plotly (if you’re doing visualization later), Version: 5.24.1 AUTHOR I’m Anusha P O, Data Science Intern at Datahut. I specialize in building automated data collection pipelines that transform raw online information into structured, analysis-ready datasets. In this blog, I dive into a crucial topic in today’s digital era — data protection and GDPR enforcement. Organizations across the globe are responsible for handling vast amounts of personal data, and with the General Data Protection Regulation (GDPR) setting strict standards since 2018, non-compliance can lead to serious fines and legal consequences. To bring this topic to life, I’ll walk you through how we can use the GDPR Enforcement Tracker to collect, clean, and structure real-world enforcement case data. This includes details like fines, countries, violation types, and quoted articles—turning a complex multi-page table into meaningful insights. At Datahut, we build smart, scalable scraping solutions that power business decisions with reliable data. If your organization wants to leverage public web data for research, compliance tracking, or competitive intelligence, feel free to connect with us through the chat widget. Let’s transform raw data into actionable intelligence. Related posts: ### FAQs 1\. Can web scraping be used to track GDPR fines? Yes. Web scraping can collect publicly available information from regulatory websites, news portals, and official announcements to track GDPR fines issued to companies. 2\. Is it legal to scrape GDPR fine data? Yes, as long as the data is publicly available and no personal or sensitive information is collected unlawfully. Always ensure compliance with website terms of service and applicable data protection laws. 3\. How can tracking GDPR fines benefit businesses? Tracking GDPR fines helps businesses understand common compliance pitfalls, benchmark their own practices, and implement proactive measures to avoid violations and financial penalties. 4\. What tools are recommended for scraping GDPR fine data? Popular tools include Python libraries like Beautiful Soup, Scrapy, Requests, and automated browsers like Playwright or Selenium. These tools help extract structured information efficiently from websites and reports. 5\. How can I ensure accuracy and reliability while scraping GDPR fines? - Cross-verify scraped data from multiple sources. - Automate regular updates to capture new fines. - Clean and structure the data to maintain consistency and reliability. ### How to Bypass Cloudflare When Web Scraping (curl_cffi Guide) [2026] URL: https://www.blog.datahut.co/post/web-scraping-without-getting-blocked-curl-cffi/ Last updated: 2026-09-07T09:44:06.000Z If your Python scrapers keep running into 403s, CAPTCHA walls, or that maddening situation where requests just quietly get throttled, here's the thing most people miss: it's usually not your code, and it's often not even your IP. It's your TLS fingerprint. The requests library basically announces itself as Python the second it opens a connection, well before the server has read a single one of your headers. curl\_cffi fixes that at the exact layer where the detection happens. In this guide we'll walk through what it does, what it can't do, and how we use it in production, with code you can copy and run as-is. One thing we want to be straight about up front, because a lot of articles aren't: curl\_cffi beats TLS and network-layer detection, and that's the gate on most e-commerce and data sites. What it doesn't do is run JavaScript. So on its own, it won't clear Cloudflare Turnstile or those "Checking your browser…" challenges. We'll come back to that. ## Why requests Gets Blocked For years, requests was the obvious choice: stable, readable, easy to teach. We still reach for it constantly. But the web kept evolving and the library, by design, mostly didn't. When requests opens an HTTPS connection, the TLS handshake it produces looks nothing like a browser's. And anti-bot systems are reading four things long before your headers even come into play: - TLS fingerprint (JA3/JA4): requests sends a static OpenSSL handshake that Cloudflare has seen millions of times, so it gets recognized instantly. (JA3 is the original TLS-fingerprinting standard, [open-sourced by Salesforce in 2017](https://github.com/salesforce/ja3?ref=blog.datahut.co); JA4 is its [successor from FoxIO](https://github.com/FoxIO-LLC/ja4?ref=blog.datahut.co). Both are what anti-bot systems compute on your handshake, and [Cloudflare documents using them directly in Bot Management](https://developers.cloudflare.com/bots/additional-configurations/ja3-ja4-fingerprint/?ref=blog.datahut.co).) - HTTP version: requests only speaks HTTP/1.1, while real Chrome negotiates HTTP/2\. That mismatch alone is enough to flag you. - Header order and casing: browsers send headers in a specific, consistent order, and requests doesn't match it. - ALPN negotiation: the protocol-negotiation sequence isn't what any real browser would send. You can spoof the User-Agent all day, but the handshake underneath still says "Python." It's a disguise with the wrong voice. (For the wider picture of how sites detect and block automated traffic, see our guide on [how to bypass anti-scraping tools](https://www.blog.datahut.co/post/web-scraping-how-to-bypass-anti-scraping-tools-on-websites/).) ## What curl\_cffi Actually Is curl\_cffi is a Python HTTP client built on a fork of curl-impersonate, wired up through CFFI (a fast bridge between C and Python). Rather than just patching headers, it swaps out the TLS stack so the handshake itself matches a real browser, right down to cipher suites, extension ordering, GREASE values, supported curves, and HTTP/2 settings. So when you tell it to impersonate Chrome, your script sends the exact encrypted handshake Chrome would. From Cloudflare's Bot Management or Akamai's point of view, the connection reads as a real browser session at the network layer. Two things, in our experience, make it practical and not just a neat trick: 1. The API mirrors requests. Porting an existing scraper is usually a one-line change to your import. (If you're newer to the ecosystem, our [web scraping in Python guide](https://www.blog.datahut.co/post/web-scraping-in-python/) covers the fundamentals curl\_cffi slots into.) 2. It's fast. Sitting on libcurl's C core, it holds its own against aiohttp and leaves pure-Python clients behind once you're doing serious volume. ## Quick Start Install it first (heads-up: the package name uses a hyphen, the import uses an underscore, which is easy to trip over): ``` pip install curl-cffi ``` Here's the smallest Cloudflare-safe request you can make: ``` from curl_cffi import requests resp = requests.get("https://www.amazon.com/", impersonate="chrome") print(resp.status_code) print(resp.http_version) # HTTP/2, a browser-like negotiation ``` A small but important habit: use impersonate="chrome", not a pinned version like chrome124\. The generic alias always resolves to the newest profile your installed version supports, so your scraper doesn't quietly rot as browsers move on. We've seen plenty of "it used to work" tickets that traced back to nothing more than a stale pinned profile. Now compare that to plain requests, which on the same URL usually gets the door slammed: ``` import requests resp = requests.get("https://www.amazon.com/") print(resp.status_code) # commonly 403, or a 200 that's actually a bot page ``` And while we're here, that last comment matters. A 200 does not mean you won. Plenty of sites hand back a 200 with a CAPTCHA or a "verify you're human" page sitting in the body. Always look at what came back, not just the status code. ## A Realistic Production Request In an actual pipeline you'll want a persistent session (so cookies and connections get reused), proxies, and headers that look the part: ``` from curl_cffi import requests session = requests.Session(impersonate="chrome") headers = { "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Referer": "https://www.google.com/", } proxies = { "http": "http://user:pass@proxy-host:8080", "https": "http://user:pass@proxy-host:8080", } resp = session.get("https://httpbin.org/anything", headers=headers, proxies=proxies) print(resp.json()) ``` Since the session negotiates the way Chrome does, the connection holds up as authentic end to end. Not just on request one, but across the whole cookie-backed session. And here are a few things the beginner tutorials tend to skip, which we've learned the hard way: - Sessions aren't thread-safe. Give each thread its own session, or move to AsyncSession (below). Share one session across threads and you'll get the kind of intermittent failures that eat an afternoon. - Actually verify the fingerprint landed. In staging, route traffic through a local mitmproxy and log the JA4 hash of your outbound connections, then confirm it matches the browser you're claiming before you ship. This is how you catch "profile drift," where a target tightens its detection so your once-good profile silently stops matching, before it starts quietly degrading your data. - Pin the library version and audit it quarterly. Browser TLS parameters shift with each major release, so lock curl-cffi in your dependency file and revisit it when a new Chrome stable lands. ## Async Scraping for High Volume For I/O-bound pipelines hammering hundreds of URLs, async beats threading, and it sidesteps the thread-safety headache entirely: ``` import asyncio from curl_cffi.requests import AsyncSession async def fetch(session, url): resp = await session.get(url, impersonate="chrome") return resp.status_code, len(resp.text) async def main(urls): async with AsyncSession() as session: tasks = [fetch(session, u) for u in urls] return await asyncio.gather(*tasks) urls = ["https://httpbin.org/anything"] * 20 print(asyncio.run(main(urls))) ``` Add async support with pip install "curl-cffi\[asyncio\]". The async session carries the exact same browser-grade fingerprint as the sync one, and the async with block makes sure connections get closed cleanly. (The [official asyncio quick-start](https://curl-cffi.readthedocs.io/en/latest/quick%5Fstart.html?ref=blog.datahut.co) is the canonical reference if you want to go deeper.) ## requests vs curl\_cffi There's really one row that decides everything here, and it's the last technical one: neither library runs JavaScript. That single fact draws the line around what curl\_cffi can and can't do for you. ## How curl\_cffi Compares to Other HTTP Clients requests isn't your only option, and curl\_cffi isn't the only client that can fake a fingerprint. So here's how the field actually shakes out in 2026\. We've split it into clients that can impersonate TLS fingerprints and the standard ones that can't. A few honest takeaways from working with these: - httpx and aiohttp are great clients that just don't help here. They're modern and quick, but they still hand over a Python TLS fingerprint, so against a Cloudflare TLS gate they fall over for the same reason requests does. Save them for APIs and unprotected targets, not anti-bot work. - primp is genuinely faster if throughput is your bottleneck, and it's worth benchmarking on your own workload. It's also picked up an AsyncClient and, unusually, lets you choose the impersonated OS independently of the browser. We dig into it just below. - tls-client and CycleTLS lean Go/Node. Both are solid, but calling them from a Python codebase means a binding layer and the friction that comes with it. If you're already in Python, that friction rarely earns its keep over curl\_cffi. - hrequests is the interesting one because it can fall back to a real browser, which starts to chip away at the JS-challenge gap we'll get to, though you pay for it with a heavier dependency. For most Python teams, curl\_cffi lands in the sweet spot: maturity, broad browser coverage, HTTP/2 and HTTP/3, async, and an API that's basically a drop-in. primp is the one we'd A/B test against it when performance really counts. And one caveat that covers every row in the top half of that table: TLS impersonation is the foundation, not the whole building. Even a flawless Chrome fingerprint won't get you past a JavaScript challenge, which is precisely where curl\_cffi and its alternatives all hit the wall. ## The Alternative Worth Knowing: [primp](https://github.com/deedy5/primp?ref=blog.datahut.co) Most curl\_cffi write-ups wave vaguely at "other libraries exist" and move on. We think primp deserves a real look, because it's the one mainstream alternative that goes toe-to-toe with curl\_cffi on its own turf, and it has a couple of tricks curl\_cffi simply doesn't. primp ("Python Requests IMPersonate") is a Python binding to the Rust rquest library. It impersonates the same set of things, TLS/JA3/JA4 and HTTP/2 fingerprints, but it's built for raw speed, and it bills itself as the fastest impersonating client in Python. The basic usage is clean: ``` import primp client = primp.Client(impersonate="chrome_146", impersonate_os="windows") resp = client.get("https://tls.peet.ws/api/all") print(resp.status_code) ``` It also has an async client: ``` import asyncio import primp async def main(): async with primp.AsyncClient(impersonate="chrome_146") as client: resp = await client.get("https://tls.peet.ws/api/all") print(resp.status_code) asyncio.run(main()) ``` Two things it genuinely does better than curl\_cffi: 1. Independent OS selection. In curl\_cffi the operating system is baked into each browser profile, so you get whatever OS that profile ships with, take it or leave it. primp lets you set impersonate\_os (windows, macos, linux, android, ios, or random) separately from the browser. That extra knob matters when a target cross-checks the OS implied by your TLS fingerprint against your User-Agent and Client Hints. 2. Speed. Being Rust under the hood, it's tuned for throughput and tends to come out on top in head-to-head benchmarks. Where curl\_cffi still wins for most teams: it's more mature, has a bigger community and far more documentation, ships HTTP/3 fingerprints, and its API is closer to a true requests drop-in. primp's API is requests-flavored but not identical, since you spin up a Client rather than calling module-level functions the same way. So, bottom line: make curl\_cffi your default. Reach for primp when you're throughput-bound at serious volume, or when you specifically need to split the impersonated OS from the browser profile, which is a genuine edge case curl\_cffi can't handle today. ## Where curl\_cffi Stops: JS Challenges curl\_cffi handles the network layer, full stop. What it can't touch is anything that runs code inside the browser. So if you hit a "Checking your browser…" screen, or you notice a cf\_clearance cookie getting set after a short wait, that page is running a JavaScript challenge (Cloudflare's IUAM or Turnstile), and TLS impersonation on its own won't get you through. When that happens, you've basically got three options: 1. Go hybrid. Let a real browser ([Playwright](https://www.blog.datahut.co/post/scraping-amazon-reviews-playwright-python/), Nodriver) solve the challenge once, grab the cf\_clearance cookie, then hand it to curl\_cffi for all the fast follow-up requests. You pay the browser tax once instead of on every call. (This is the path you take when a page needs full rendering, the same approach in our walkthrough on [scraping a dynamic website with Python](https://www.blog.datahut.co/post/scrape-a-dynamic-website-using-python/).) 2. Use a solver. Services like CapSolver or 2Captcha can handle the challenge token programmatically. 3. Offload it entirely. A managed unblocking API takes fingerprinting, challenges, and proxies off your plate completely. For the big chunk of sites protected only by TLS fingerprinting, meaning most e-commerce catalogs, pricing pages, and listing data, curl\_cffi on its own gets the job done. ## Choosing a Browser Profile A few rules of thumb that have saved us a lot of debugging: - Default to impersonate="chrome" (or "safari"). The aliases follow the latest supported fingerprints on their own, so you don't have to think about it. - If you must pin a version, pin a recent one. Profiles run from ancient (chrome99) to current, and the newest ones track the latest stable releases. An old pinned profile in 2026 is, ironically, its own detection signal. - Watch for version gaps. curl\_cffi only adds a new profile when the fingerprint actually changes, so if a version number looks "missing," its fingerprint just didn't change in any meaningful way. Reach for the nearest available one and match the headers. (The [official list of supported impersonate targets](https://curl-cffi.readthedocs.io/en/latest/impersonate/targets.html?ref=blog.datahut.co) is the source of truth here.) ## Still Getting Blocked? A Checklist If your fingerprint is right and you're still getting blocked, work down this list. It's roughly the order we troubleshoot in: 1. Datacenter IPs. Cloudflare scores datacenter ranges as low-trust regardless of how perfect your fingerprint is. Switch to residential proxies. This is the single most common fix. (Our [guide to using proxies for web scraping](https://www.blog.datahut.co/post/a-guide-to-using-proxies-for-web-scraping/) breaks down the residential-versus-datacenter trade-off.) 2. Stale profile. Update the library (pip install -U curl-cffi) and use the generic chrome alias. 3. Request rate. Real users don't load 100 pages in two seconds. Add randomized 1 to 3 second delays. 4. Thin headers. Add Accept-Language, Accept-Encoding, and a plausible Referer. 5. Session/IP mismatch. Reusing one session across rotating IPs is inconsistent and flaggable. Keep one session per IP. 6. It's a JS challenge. If none of the above helps and you see a challenge page, you've hit the JS wall, so pair with a browser or solver as above. ## When curl\_cffi Is the Right Tool Reach for it when you're running into 403s or CAPTCHAs on TLS-fingerprinting sites, when you want HTTP/2 performance at volume, or when you're running scraping APIs that have to keep their success rates up. For lightweight public data or a simple API, honestly, plain requests is still fine, with no need to drag in the extra machinery. The bigger picture: reliable large-scale scraping comes from browser-accurate networking, not from brute-force retries or firing up a headless browser for every single page. curl\_cffi hands you most of a browser's stealth at a tiny fraction of what Playwright or Puppeteer cost in resources, and you keep the heavy tools in reserve for the genuinely JS-gated pages that actually need them. (If you're still mapping out your stack, our roundup of [web scraping tools](https://www.blog.datahut.co/post/web-scraping-tools/) and our walkthrough on [building a web crawler from scratch](https://www.blog.datahut.co/post/how-to-build-a-web-crawler-from-scratch/) are good companions to this piece.) ## FAQ 1. Is curl\_cffi a drop-in replacement for requests? Pretty much. Swap import requests for from curl\_cffi import requests, add impersonate="chrome" to your calls, and most of your existing code keeps running untouched. 2. Does it bypass Cloudflare Turnstile? .No. Turnstile is a JavaScript challenge and curl\_cffi has no JS engine. Use it for TLS-gated sites, and pair it with a browser or solver when you hit a JS challenge. 3. Which profile should I use? impersonate="chrome". It resolves to the latest supported fingerprint on its own, so you're never tracking version numbers or accidentally pinning something stale. 4. Why am I still blocked even though I'm impersonating? Usually, in this order: datacenter IPs, too high a request rate, thin headers, a stale library, or a JavaScript challenge that the network layer just can't clear. Run the checklist above. 5. Is it legal to use? The library itself is open-source and perfectly legal. Whether a given scrape is lawful comes down to the site's Terms of Service, its robots.txt, and the laws that apply to you (including GDPR if personal data is involved), not the tool you picked. Check robots.txt, respect crawl delays, and stay away from anything behind a login. (We cover this in more depth in [is web scraping legal?](https://www.blog.datahut.co/post/is-web-scraping-legal/)) ## Conclusion: When to Build, and When to Hand It Off curl\_cffi is the right first move against TLS-layer detection. It's mature, it's fast, it's close to a drop-in for requests, and for the big share of sites that are gated only by fingerprinting, it's often everything you need. Pair it with residential proxies, sane request pacing, and the generic chrome profile, and you'll clear most of what currently blocks requests. But we'd rather be honest about what comes after that first win. A reliable scraper was never really one library. It's an ongoing operation. You're rotating and vetting proxy pools, keeping fingerprint profiles current as browsers ship new versions, handling the JavaScript challenges curl\_cffi can't, watching for silent fingerprint drift, and babysitting the ban-and-retry logic the hard targets force on you. Any one of those is manageable. Stacked together, at scale, they quietly turn into a full-time engineering job, and that maintenance load tends to grow faster than the data you're actually collecting. That's really the decision hiding behind the tooling. If scraping is your product, building all of this in-house makes sense. If you just need the data and not the infrastructure headache, the math usually points the other way. That second path is what we do at [Datahut](https://datahut.co/?ref=blog.datahut.co). We run managed, fully compliant web scraping as a service. We handle the fingerprinting, the proxy management, the anti-bot challenges, and the pipeline reliability, so your team ends up with clean, structured data on a schedule instead of a growing pile of broken scrapers. Whether you're tracking competitor pricing, keeping an eye on product assortment and stock, or building a data feed across hundreds of protected sites, we deliver the output and absorb the upkeep. So if your scrapers are spending more time getting unblocked than actually collecting data, [come talk to us at Datahut](https://datahut.co/?ref=blog.datahut.co). We'll scope it with you and tell you honestly whether managed extraction is the right call for your case. References & Further Reading All primary sources, meaning official project documentation and the original creators of the underlying standards, not third-party vendors: - curl\_cffi official documentation: [https://curl-cffi.readthedocs.io/en/latest/](https://curl-cffi.readthedocs.io/en/latest/?ref=blog.datahut.co) - curl\_cffi source and release notes (GitHub): [https://github.com/lexiforest/curl\_cffi](https://github.com/lexiforest/curl%5Fcffi?ref=blog.datahut.co) - curl-impersonate, the underlying engine curl\_cffi binds to: [https://github.com/lexiforest/curl-impersonate](https://github.com/lexiforest/curl-impersonate?ref=blog.datahut.co) - primp source and documentation (GitHub): [https://github.com/deedy5/primp](https://github.com/deedy5/primp?ref=blog.datahut.co) - JA3, the original TLS client fingerprinting standard, open-sourced by Salesforce in 2017 (BSD-3-Clause): [https://github.com/salesforce/ja3](https://github.com/salesforce/ja3?ref=blog.datahut.co) - JA4 / JA4+, the successor standard from FoxIO, by JA3's original author John Althouse: [https://github.com/FoxIO-LLC/ja4](https://github.com/FoxIO-LLC/ja4?ref=blog.datahut.co) - JA4 explainer, John Althouse's write-up of the JA4+ methods: [https://medium.com/foxio/ja4-network-fingerprinting-9376fe9ca637](https://medium.com/foxio/ja4-network-fingerprinting-9376fe9ca637?ref=blog.datahut.co) - Cloudflare, official JA3/JA4 fingerprint documentation: [https://developers.cloudflare.com/bots/additional-configurations/ja3-ja4-fingerprint/](https://developers.cloudflare.com/bots/additional-configurations/ja3-ja4-fingerprint/?ref=blog.datahut.co) - Cloudflare engineering blog, on JA4 fingerprints and inter-request signals: [https://blog.cloudflare.com/ja4-signals/](https://blog.cloudflare.com/ja4-signals/?ref=blog.datahut.co) - curl\_cffi, supported browser impersonate targets: [https://curl-cffi.readthedocs.io/en/latest/impersonate/targets.html](https://curl-cffi.readthedocs.io/en/latest/impersonate/targets.html?ref=blog.datahut.co) For background: JA3 was invented at Salesforce in 2017, though the project is no longer actively maintained there; its original creator, John Althouse, now maintains the newer fingerprinting work at FoxIO. JA4 TLS client fingerprinting is open-source under the BSD 3-Clause licence, the same as JA3, with no patent claims, so any tool using JA3 can move to JA4 freely. ### Scaling Web Scraping Projects from Prototype to Production URL: https://www.blog.datahut.co/post/scaling-web-scraping-from-prototype-to-production-challenges-explained/ Last updated: 2026-09-07T09:44:08.000Z ## Why Scaling Web Scraping Is Harder Than You Think Web scraping is the automated process of extracting structured data from websites. It plays a vital role in data collection, market intelligence, competitive analysis, and AI-powered business strategies. Building a prototype scraper is relatively simple—most developers can set one up with Python and BeautifulSoup in a day. But scaling that prototype to handle millions of pages across multiple geographies introduces serious challenges in data volume, compliance, infrastructure, and performance. This guide explores the journey “from prototype to scale” and why building a scalable web scraping system is more difficult than it looks. ![scraping at prototype vs production scale](https://www.blog.datahut.co/content/images/2026/07/img-450.png.webp) Related read: \[[Web Scraping vs API: Which is Best for Data Extraction?](https://www.blog.datahut.co/post/web-scraping-vs-api/)\] ## What Is the Prototype Phase in Web Scraping? The prototype phase of a web scraping project typically involves: - Targeting a single or small number of web pages - Using simple tools like Python + BeautifulSoup or Puppeteer - Handling minimal edge cases in data extraction - Manually managing cookies, headers, and proxies At this stage, scraping feels simple and efficient. But what’s missing are: - Scalable architecture - Error handling & retries - Data pipelines - Sustainable data quality management ### Popular Prototype Tools - BeautifulSoup - Requests - Scrapy - Selenium ➡ The simplicity at this stage often hides the complexity of scaling. ## Challenges in Scaling Web Scraping Systems Scaling a scraper from prototype to production introduces multiple hurdles: ### 1\. Managing Data Volume and Speed Handling millions of pages requires distributed crawlers, rotating IPs, and cloud infrastructure. Without them, scrapers crash or slow down drastically. ### 2\. Overcoming Anti-Bot Mechanisms Websites deploy defenses like: - Rate limiting - Browser fingerprinting - CAPTCHAs You’ll need rotating proxy services, headless browsers, and resilient frameworks to avoid detection. ### 3\. Handling Performance Bottlenecks Without performance tuning, scrapers get blocked, delayed, or fail entirely.Related read: \[[Challenges of Large-Scale Web Scraping](https://www.blog.datahut.co/post/web-scraping-at-large-data-extraction-challenges-you-must-know/)\] ### 4\. Error Handling & Recovery at Scale When HTML structures change or 403 errors appear, scalable scrapers must use: - Retry logic - Adaptive parsers - Fault-tolerant workflows ### 5\. Infrastructure and Cost Challenges Managing distributed scrapers requires container orchestration tools such as: - Docker - Kubernetes - Apache Kafka ## Real-World Examples: Scaling With Datahut ### 1\. Retail Analytics Firm A global retail analytics firm needed to scrape over 2 million product listings across 4 countries to track competitor pricing and promotions. - Prototype Phase: Their Scrapy-based scraper worked for a few thousand pages. - Problem: It broke under dynamic content, IP bans, and slow page loads. - Solution: They shifted to Datahut’s Data-as-a-Service model, gaining:Distributed scraping infrastructureAutomated proxy rotationClean, validated data pipelines Result: The firm scaled seamlessly, reduced infrastructure costs, and redirected resources to insights instead of scraper maintenance. ### 2\. Financial Data Provider A fintech company needed stock price movements, filings, and market sentiment data from multiple sources. - Prototype Phase: Internal scripts using Puppetteer. - Problem: Scripts failed on frequent site structure changes and couldn’t meet regulatory compliance standards. - Solution: With Datahut, they gained:Enterprise-level monitoring and error recoveryGDPR/CCPA-compliant scraping infrastructureStructured datasets delivered via API Result: The fintech firm scaled globally with legally compliant, high-quality datasets, boosting their financial models’ accuracy. Takeaway: Scaling web scraping isn’t just about tech—it’s about having the right infrastructure, compliance practices, and data pipelines. Datahut helps companies achieve all three. ## Best Tools and Techniques for Scalable Web Scraping ### Tools to Consider - Scrapy + Splash (JS rendering at scale) - Playwright & Selenium (interactive scraping) - Multiple proxy vendors (to reduce IP bans) - Apache Kafka, Airflow, Redis (workflow orchestration) ### Techniques for Optimization - Parallel scraping with Celery or asyncio - Caching and request deduplication - Monitoring latency and uptime - Using headless browsers only when required Explore more: \[[Web Scraping Best Practices](https://www.blog.datahut.co/post/web-scraping-best-practices-tips/)\] ## Scalable Web Scraping Architecture: Key Components A reliable large-scale scraping system typically includes: - Scheduler → Assigns scraping jobs - URL Queue → Redis / Kafka - Scraping Workers → Dockerized containers - Proxy Manager → Manages IP rotation - Parser Modules → Extract structured data - Database → MongoDB / PostgreSQL - Monitoring Systems → Prometheus & Grafana ## Legal and Ethical Considerations in Large-Scale Scraping Scaling scraping requires compliance with legal frameworks: - Respect robots.txt and site terms - Follow GDPR, CCPA rules - Avoid scraping personal or copyrighted data ## Data Management and Quality Assurance ### Data Management Strategies - Use cloud data management solutions - Apply schema validation and deduplication - Timestamp & version scraped datasets ### Quality Assurance - Run automated QA pipelines - Audit random samples manually - Monitor failed URLs & re-scrape intelligently ## Conclusion: Building Reliable and Compliant Scrapers Scaling web scraping is not just adding more servers or faster crawlers. It requires: - Robust architecture (modular, distributed, fault-tolerant) - Legal compliance (GDPR, CCPA, ToS) - Data quality management (clean, deduplicated, structured outputs) Key Takeaways: - Prototypes are simple, but scaling demands enterprise-grade systems - Compliance and ethics must be built into strategy - Monitoring and data pipelines are critical for reliability Want to scale your scraping project without hitting roadblocks? Talk to Datahut , [datahut.co](https://www.datahut.co/?ref=blog.datahut.co) for enterprise-grade web scraping solutions. ## Frequently Asked Questions 1\. What are the main challenges in scaling web scraping? Scaling introduces issues with data volume, anti-bot mechanisms, infrastructure costs, and maintaining data quality. 2\. Which tools help optimize large-scale web scraping? Tools like Scrapy, Playwright, Splash, Docker, and Apache Kafka make scrapers more resilient at scale. 3\. How can I ensure legal compliance while scraping websites? Follow ethical scraping practices: respect robots.txt, avoid personal data, and comply with GDPR/CCPA. 4\. What is the difference between a prototype scraper and a scalable scraper? A prototype scraper targets small datasets, while a scalable scraper can handle millions of pages across multiple geographies. 5\. How does data management impact scraping efficiency? Good data management ensures cleaner datasets, faster processing, and higher accuracy for analytics or AI models. 6\. How much does it cost to scale web scraping? Costs depend on infrastructure, proxies, compliance, and monitoring. Outsourcing to a provider like Datahut often reduces overhead. 7\. Can AI improve large-scale web scraping? Yes—AI-driven scrapers help with adaptive parsing, anomaly detection, and automation of QA pipelines. ### Scraping Amazon US Product Data for Business Use URL: https://www.blog.datahut.co/post/how-to-scrape-product-data-from-amazon-us/ Last updated: 2026-07-23T07:48:31.000Z ## Introduction Ever tried shopping for vlogging equipment on Amazon? It's overwhelming. You've got thousands of microphones, cameras, and tripods to choose from, and manually comparing them all would take forever. That's exactly why I built this web scraping system - to automatically collect and organize all that product data so you can actually make informed decisions. This project shows you how to build a complete two-phase scraping system that systematically extracts vlogging equipment data from Amazon. We're talking about transforming scattered product information into a clean, structured database that you can actually analyze. The first phase quickly collects all the product URLs we need, while the second phase dives deep into each product page to extract detailed information like prices, ratings, and specifications. I chose this two-phase approach because it's more reliable and efficient than trying to do everything at once. If something goes wrong during the detailed extraction, you haven't lost all your URL collection work. Plus, it's much easier to handle Amazon's complex JavaScript-heavy pages when you break the process into focused chunks. We'll be using modern Python tools that can handle real web applications - Playwright for browser automation, BeautifulSoup for parsing HTML, and SQLite for data storage. The end result is a professional-grade scraping system that respects Amazon's servers while giving you the data you need. Whether you're researching equipment purchases, analyzing market trends, or just learning advanced scraping techniques, this guide will show you exactly how to build something that actually works in the real world. ## Phase 1: Collecting Amazon Product URLs for Vlogging Gadgets Welcome to the first phase of our Amazon web scraping adventure! This script is all about gathering product URLs from Amazon search results, specifically targeting vlogging equipment like microphones, cameras, and tripods. Think of this as the preliminary mission before the main event. We're systematically collecting every product URL we can find across multiple categories of vlogging gear. It's like walking through Amazon's virtual aisles and writing down the location of every interesting product we see. The beauty of this approach is that once we have all the URLs stored in our database, we can take our time with the detailed scraping in phase two. We're building a solid foundation that will make the next phase much more efficient and organized. ### Setting Up Our Tools Let's start with the imports and basic setup. Every web scraping project needs its essential tools, and ours is no different. ``` import sqlite3 import random from bs4 import BeautifulSoup from urllib.parse import urljoin from playwright.sync_api import sync_playwright # Categories and search URLs categories = { "microphone": "https://www.amazon.com/s?k=microphones+for+vlogging", "camera": "https://www.amazon.com/s?k=camera+for+vlogging", "tripod": "https://www.amazon.com/s?k=tripod+for+vlogging" } ``` We're using several libraries here, each with a specific purpose. SQLite handles our database storage - it's like having a filing cabinet where we can organize all our collected URLs. BeautifulSoup is our HTML parser, helping us extract information from web pages. The urllib.parse module helps us work with URLs properly. Playwright is our browser automation tool. Unlike simple HTTP requests, Playwright controls an actual browser, which means it can handle JavaScript and behave more like a real person browsing Amazon. This is crucial because Amazon's pages rely heavily on JavaScript. The categories dictionary defines what we're looking for. Each category has a search URL that takes us directly to Amazon's search results for that type of vlogging equipment. ### Loading User Agents for Stealth Web scraping requires being respectful and avoiding detection. One way to do this is by rotating user agents - the identifier that tells websites what browser and device you're using. ``` def load_user_agents(file_path="user_agents.txt"): """ Load user agent strings from a text file for browser rotation. This function reads a text file containing user agent strings (one per line) and returns them as a list. User agent rotation helps avoid detection by making requests appear to come from different browsers/devices. Args: file_path (str, optional): Path to the text file containing user agents. Defaults to "user_agents.txt". Returns: list: A list of user agent strings with empty lines filtered out. Raises: FileNotFoundError: If the specified file doesn't exist. IOError: If there's an error reading the file. """ with open(file_path, "r") as f: return [line.strip() for line in f if line.strip()] user_agents = load_user_agents() ``` This function reads a text file containing different user agent strings. Each line represents a different browser or device. When we make requests, we randomly pick one of these identifiers, making our scraper appear to come from different sources. It's like changing your disguise each time you visit a store. The function filters out empty lines and returns a clean list of user agents. We load these once at the start of our script and use them throughout the scraping process. ### Database Setup Before we start collecting URLs, we need somewhere to store them. We're using SQLite, which is perfect for this kind of project because it doesn't require a separate server. ``` def setup_database(): """ Initialize SQLite database and create the products_url table. Creates a SQLite database file named "amazon_vlogging.db" and sets up the products_url table to store scraped product information. The table has an auto-incrementing ID, category field, and URL field with unique constraint to prevent duplicates. Returns: tuple: A tuple containing (connection, cursor) objects for database operations. - connection (sqlite3.Connection): Database connection object - cursor (sqlite3.Cursor): Database cursor for executing queries Table Schema: - id: INTEGER PRIMARY KEY AUTOINCREMENT - category: TEXT (product category like 'microphone', 'camera', etc.) - url: TEXT UNIQUE (product page URL, duplicates ignored) """ conn = sqlite3.connect("amazon_vlogging.db") cur = conn.cursor() cur.execute(''' CREATE TABLE IF NOT EXISTS products_url ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT UNIQUE ) ''') conn.commit() return conn, cur ``` This function creates our database file and sets up a table called products\_url. Think of this table as a spreadsheet with three columns: an ID number that increases automatically, the category (like "microphone" or "camera"), and the actual URL. The UNIQUE constraint on the URL column is important - it prevents us from storing the same product URL twice. If we try to insert a duplicate URL, SQLite will simply ignore it, which saves us from having to check for duplicates manually. ### Extracting Product URLs from Search Results Now comes the heart of our URL collection process. Amazon's search results pages contain links to individual products, and we need to extract all of them. ``` def extract_product_urls(html): """ Extract product page URLs from Amazon search results HTML. Parses the HTML content of an Amazon search results page and extracts all product page URLs. Uses BeautifulSoup to find product image links which contain the product page URLs as href attributes. Args: html (str): Raw HTML content of an Amazon search results page. Returns: list: A list of unique, fully-qualified product page URLs. Duplicates are automatically removed. Technical Details: - Targets: span[data-component-type="s-product-image"] > a elements - Converts relative URLs to absolute URLs using Amazon's base URL - Removes duplicate URLs using set() conversion """ soup = BeautifulSoup(html, "html.parser") links = [] for tag in soup.select('span[data-component-type="s-product-image"] > a'): partial = tag.get("href") if partial: full_url = urljoin("https://www.amazon.com", partial) links.append(full_url) return list(set(links)) # remove duplicates ``` This function takes the raw HTML of a search results page and finds all the product links. We're specifically looking for links inside product image spans - these are the clickable product images that take you to individual product pages. Amazon often uses relative URLs (like "/dp/B08ABC123"), so we use urljoin to convert them into complete URLs. The function returns a list of unique URLs by converting to a set and back to a list, which automatically removes any duplicates we might have found on the same page. ### Setting ZIP Code for Consistent Results Amazon shows different prices and availability based on your location. To get consistent results, we need to set a specific ZIP code before scraping. ``` def apply_zip(page, zip_code="56901"): """ Set delivery ZIP code on Amazon using browser automation. Navigates to Amazon's homepage and programmatically sets the delivery ZIP code through the location selector interface. This affects product availability, pricing, and shipping options in search results. Args: page (playwright.sync_api.Page): Playwright page object for browser automation. zip_code (str, optional): ZIP code to set for delivery location. Defaults to "56901". Returns: None Raises: Exception: Catches and prints any errors during ZIP code setting process. Script continues execution even if ZIP setting fails. Process Flow: 1. Navigate to Amazon homepage 2. Click on location selector (glow-ingress-line2) 3. Fill ZIP code input field 4. Submit ZIP code update 5. Handle optional confirmation dialog """ try: page.goto("https://www.amazon.com", timeout=60000) page.wait_for_selector("span#glow-ingress-line2", timeout=10000) page.click("span#glow-ingress-line2") page.wait_for_selector("input#GLUXZipUpdateInput", timeout=10000) page.fill("input#GLUXZipUpdateInput", zip_code) page.click("#GLUXZipUpdate > span > input") page.wait_for_timeout(3000) try: page.click("span.a-button-inner > input[name='glowDoneButton']") except: pass print(f"📍 ZIP code set to {zip_code}") except Exception as e: print("❌ Failed to set ZIP:", e) ``` This function navigates to Amazon's homepage and simulates clicking on the location selector. It then fills in the ZIP code field and submits the form. The process mimics what you'd do manually when changing your delivery location on Amazon. We use try-except blocks because the ZIP code setting process can vary slightly depending on your account status or Amazon's interface changes. If something goes wrong, we print an error but continue with the scraping process. ### The Main Scraping Function Now we bring everything together in our main scraping function. This is where the magic happens - we systematically go through each category and collect all the product URLs. ``` def scrape_urls_with_zip(): """ Main scraping function that collects product URLs from Amazon search results. This is the primary orchestration function that coordinates the entire scraping process. It sets up the database, configures the browser with ZIP code settings, and systematically scrapes product URLs from multiple categories across all available pages. Returns: None: Results are saved directly to the SQLite database. Process Overview: 1. Initialize database connection and cursor 2. Launch Playwright browser with random user agent 3. Set ZIP code for consistent pricing/availability 4. Iterate through each product category 5. For each category, scrape all pages of search results 6. Extract and store product URLs in database 7. Handle pagination automatically 8. Clean up browser resources Features: - Automatic pagination handling - Random delays between requests (3-5 seconds) - User agent rotation for anti-detection - Cookie persistence across category scraping - Duplicate URL prevention via database constraints - Comprehensive error handling and logging Database Operations: - Creates products_url table if not exists - Inserts URLs with category labels - Uses INSERT OR IGNORE to prevent duplicates - Commits after each page to prevent data loss Anti-Detection Measures: - Random user agent selection - Variable delays between requests - Cookie preservation - Realistic browsing patterns Raises: Exception: Various exceptions may occur during scraping (network issues, page structure changes, etc.). Most are caught and logged without stopping the entire process. """ conn, cur = setup_database() with sync_playwright() as p: browser = p.chromium.launch(headless=False) user_agent = random.choice(user_agents) context = browser.new_context(user_agent=user_agent, locale="en-US") page = context.new_page() apply_zip(page, zip_code="56901") cookies = context.cookies() page.close() context.close() ``` We start by setting up our database connection and launching a browser. The headless=False parameter means we can see the browser window as it works - helpful for debugging and understanding what's happening. After setting the ZIP code, we save the browser cookies. These cookies contain our location preference and session information, which we'll reuse for each category to maintain consistency. ``` for category, url in categories.items(): print(f"\n🔍 Scraping category: {category}") context = browser.new_context(user_agent=user_agent, locale="en-US") context.add_cookies(cookies) page = context.new_page() page.goto(url, timeout=60000) page.wait_for_timeout(random.uniform(3000, 5000)) all_urls = set() ``` For each category, we create a fresh browser context but add our saved cookies. This gives us a clean slate while maintaining our location settings. We navigate to the search URL and wait a random amount of time between 3-5 seconds. This random delay makes our scraper behave more like a human user. ``` while True: html = page.content() product_urls = extract_product_urls(html) print(f"✅ Found {len(product_urls)} product URLs on this page") all_urls.update(product_urls) for link in product_urls: try: cur.execute("INSERT OR IGNORE INTO products_url (category, url) VALUES (?, ?)", (category, link)) except Exception as e: print("⚠️ DB Insert Error:", e) conn.commit() ``` The main scraping loop processes each page of search results. We extract the product URLs, add them to our running set of URLs, and save them to the database. The INSERT OR IGNORE statement means duplicate URLs won't cause errors - they'll simply be skipped. ``` try: next_button = page.query_selector("a.s-pagination-next") if next_button: next_href = next_button.get_attribute("href") if next_href: next_url = urljoin("https://www.amazon.com", next_href) print("➡️ Going to next page...") page.goto(next_url) page.wait_for_timeout(random.uniform(3000, 5000)) continue except: pass print(f"🔚 No more pages in category: {category}") break print(f"📦 Total unique URLs collected for {category}: {len(all_urls)}") page.close() context.close() browser.close() conn.close() print("✅ Done! Product URLs saved to products_url table.") ``` To handle pagination, we look for the "Next" button on each page. If we find it, we extract its URL and navigate to the next page. If there's no next button or we encounter an error, we assume we've reached the end of the search results for that category. ### Running the Scraper The final piece ties everything together with a simple entry point: ``` if __name__ == "__main__": """ Entry point for the Amazon product URL scraper. Executes the main scraping function when the script is run directly. This allows the script to be imported as a module without automatically starting the scraping process. Usage: python amazon_scraper.py Prerequisites: - user_agents.txt file with user agent strings - Required packages: playwright, beautifulsoup4, sqlite3 - Playwright browser binaries installed """ scrape_urls_with_zip() ``` This is a Python convention that ensures our scraping function only runs when the script is executed directly, not when it's imported as a module. It's like having a main function that kicks off the entire process. ### What We've Built This script creates a systematic approach to collecting product URLs from Amazon. It handles the complexity of browser automation, manages database storage, and respects Amazon's systems by using realistic delays and user agent rotation. The end result is a SQLite database filled with product URLs, organized by category. Each URL represents a potential vlogging gadget that we can analyze further in the second phase of our project. The scraper is designed to be robust, handling errors gracefully and providing clear feedback about its progress. It's also respectful of Amazon's servers, using appropriate delays and behaving like a real user would. This foundation sets us up perfectly for the next phase, where we'll visit each collected URL and extract detailed product information like prices, ratings, and descriptions. ## Phase 2: Extracting Detailed Product Information Now that we have all our product URLs safely stored in the database, it's time for the main event - extracting detailed information from each product page. This phase takes our collection of URLs and transforms them into a rich dataset of product details. Think of phase one as creating a map of all the stores we want to visit, and phase two as actually going into each store and carefully examining the products. We'll gather prices, ratings, product specifications, and everything else that makes each product unique. ### Setting Up for Detail Extraction Our second script starts with familiar territory - the same imports and user agent loading we used before, but now we're focusing on data extraction rather than URL collection. ``` import sqlite3 import time import random from bs4 import BeautifulSoup from playwright.sync_api import sync_playwright # Load user agents def load_user_agents(file_path="user_agents.txt"): """ Load user agent strings from a text file for browser rotation. This function reads a text file containing user agent strings (one per line) and returns them as a list. User agent rotation helps avoid detection by making requests appear to come from different browsers/devices during scraping. Args: file_path (str, optional): Path to the text file containing user agents. Defaults to "user_agents.txt". Returns: list: A list of user agent strings with empty lines filtered out. Raises: FileNotFoundError: If the specified file doesn't exist. IOError: If there's an error reading the file. """ with open(file_path, "r") as f: return [line.strip() for line in f if line.strip()] user_agents = load_user_agents() ``` Everything remains the same - we still need BeautifulSoup for parsing HTML, Playwright for browser automation, and our user agents for stealth. ### Expanding Our Database Schema This phase requires a more advanced database setup. We need to track our scraping progress and store much more detailed information about each product. ``` def setup_database(): """ Initialize SQLite database and create/modify tables for product detail scraping. Sets up the database schema for storing detailed product information scraped from Amazon product pages. Creates new tables and modifies existing ones as needed. Adds a 'scraped' column to the existing products_url table to track scraping progress. Returns: tuple: A tuple containing (connection, cursor) objects for database operations. - connection (sqlite3.Connection): Database connection object - cursor (sqlite3.Cursor): Database cursor for executing queries Database Schema Created/Modified: products_url table: - Adds 'scraped' column (INTEGER DEFAULT 0) to track processing status product_details table: - id: INTEGER PRIMARY KEY AUTOINCREMENT - category: TEXT (product category) - url: TEXT (product page URL) - title: TEXT (product title/name) - price: TEXT (current price) - original_price: TEXT (original/list price before discount) - discount: TEXT (discount percentage or amount) - details: TEXT (JSON string of product specifications) - rating: TEXT (customer rating) error_urls table: - id: INTEGER PRIMARY KEY AUTOINCREMENT - url: TEXT UNIQUE (URLs that failed to scrape) - error: TEXT (error message description) """ conn = sqlite3.connect("amazon_vlogging.db") cur = conn.cursor() cur.execute("PRAGMA table_info(products_url)") if "scraped" not in [col[1] for col in cur.fetchall()]: print("🔧 Adding 'scraped' column...") cur.execute("ALTER TABLE products_url ADD COLUMN scraped INTEGER DEFAULT 0") ``` First, we add a "scraped" column to our existing products\_url table. This acts like a checklist - we can mark each URL as processed so we don't waste time scraping the same product twice. The PRAGMA command lets us check what columns already exist before trying to add new ones. ```     cur.execute(''' CREATE TABLE IF NOT EXISTS product_details ( id INTEGER PRIMARY KEY AUTOINCREMENT, category TEXT, url TEXT, title TEXT, price TEXT, original_price TEXT, discount TEXT, details TEXT, rating TEXT ) ''') cur.execute(''' CREATE TABLE IF NOT EXISTS error_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, error TEXT ) ''') ``` We create two new tables. The product\_details table stores all the valuable information we'll extract from each product page - titles, prices, ratings, and specifications. The error\_urls table keeps track of any URLs that fail to scrape, along with the error messages. This helps us debug problems and retry failed URLs later. ### Building Our HTML Parsing Arsenal The heart of phase two is a collection of specialized parsing functions. Each function knows how to extract one specific piece of information from Amazon's product pages. ``` def parse_title(soup): """ Extract product title from Amazon product page HTML. Searches for the main product title element using Amazon's standard product title selector. The title is typically displayed prominently at the top of the product page. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML. Returns: str or None: Product title text with whitespace stripped, or None if not found. Technical Details: - Target selector: "span#productTitle" - Uses get_text(strip=True) to clean whitespace """ tag = soup.select_one("span#productTitle") return tag.get_text(strip=True) if tag else None def parse_price(soup): """ Extract current price from Amazon product page HTML. Searches for the current/sale price element in Amazon's pricing display. This is typically the prominently displayed price that customers see. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML. Returns: str or None: Current price text (e.g., "$99.99"), or None if not found. Technical Details: - Target selector: Complex selector for core price display - Looks for screen reader accessible price text - May include currency symbols and formatting """ tag = soup.select_one("#corePriceDisplay_desktop_feature_div > div.a-section.a-spacing-none.aok-align-center.aok-relative > span.aok-offscreen") return tag.get_text(strip=True) if tag else None def parse_original_price(soup): """ Extract original/list price from Amazon product page HTML. Searches for the original price (list price) before any discounts. This price is typically crossed out or shown smaller when there's a sale. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML. Returns: str or None: Original price text (e.g., "$129.99"), or None if not found. Technical Details: - Target selector: Complex nested selector for basis price - Often displayed as strikethrough text - Only visible when product is on sale """ tag = soup.select_one("#corePriceDisplay_desktop_feature_div > div.a-section.a-spacing-small.aok-align-center > span > span.aok-relative > span.a-size-small.a-color-secondary.aok-align-center.basisPrice > span > span.a-offscreen") return tag.get_text(strip=True) if tag else None ``` These functions use CSS selectors to pinpoint exactly where Amazon displays different pieces of information. The selectors look complex, but they're just very specific addresses that tell us exactly where to find each piece of data on the page. Amazon often hides the actual price text in elements with the "aok-offscreen" class - these are invisible to users but accessible to screen readers and our scrapers. This is why we target these specific elements rather than the visible price displays. ``` def parse_discount(soup): """ Extract discount percentage from Amazon product page HTML. Searches for the discount percentage or savings amount displayed when a product is on sale. Usually shown as a percentage off. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML. Returns: str or None: Discount text (e.g., "-31%"), or None if not found. Technical Details: - Target selector: Price savings percentage element - Typically displays percentage or dollar amount saved - Only visible when product has active discount """ tag = soup.select_one("#corePriceDisplay_desktop_feature_div > div.a-section.a-spacing-none.aok-align-center.aok-relative > span.a-size-large.a-color-price.savingPriceOverride.aok-align-center.reinventPriceSavingsPercentageMargin.savingsPercentage") return tag.get_text(strip=True) if tag else None def parse_rating(soup): """ Extract customer rating from Amazon product page HTML. Searches for the average customer rating typically displayed as stars or numerical rating near the product title. Args: soup (BeautifulSoup): BeautifulSoup object containing parsed HTML. Returns: str or None: Rating text (e.g., "4.5 out of 5 stars"), or None if not found. Technical Details: - Target selector: "#acrPopover > span.a-declarative > a > span" - May include star rating and text description - Located in the product overview section """ tag = soup.select_one("#acrPopover > span.a-declarative > a > span") return tag.get_text(strip=True) if tag else None ``` The discount and rating functions work the same way - they look for specific elements that Amazon uses to display this information. Not every product has a discount or rating, so we return None when these elements don't exist. ### Extracting Product Specifications One of the most valuable parts of product data is the specifications table that Amazon displays for most products. This contains detailed information like brand, model, dimensions, and features. ``` def extract_table_data(html): """ Extract structured product specifications from Amazon's product info table. Parses Amazon's standard product information table to extract key-value pairs of product specifications such as brand, model, dimensions, features, etc. Handles truncated values by preferring full text when available. Args: html (str): Raw HTML content containing the product specifications table. Returns: List[Dict[str, str]]: A list of dictionaries where each dictionary contains a single key-value pair representing a product specification. Technical Details: - Targets table with class 'a-normal a-spacing-micro' - Extracts key from 'a-span3' class cells - Extracts value from 'a-span9' class cells - Prioritizes full text from 'a-truncate-full' spans over truncated text - Handles cases where full value is hidden behind "Show more" functionality Table Structure: - Keys are typically: Brand, Model, Color, Dimensions, Weight, etc. - Values can be text, measurements, or feature descriptions - Some values may be truncated with expandable "Show more" options Returns: Empty list if no table found or parsing fails. """ soup = BeautifulSoup(html, 'html.parser') table = soup.find('table', class_='a-normal a-spacing-micro') data = [] if not table: return data rows = table.find_all('tr') for row in rows: key_td = row.find('td', class_='a-span3') value_td = row.find('td', class_='a-span9') if key_td and value_td: key = key_td.get_text(strip=True) # Prefer full hidden value if available full_value_span = value_td.find('span', class_='a-truncate-full') if full_value_span: value = full_value_span.get_text(strip=True) else: value = value_td.get_text(strip=True) data.append({key: value}) return data ``` This function tackles Amazon's product specifications table, which has a standard structure but some tricky aspects. Amazon sometimes truncates long specification values with a "Show more" button. We handle this by looking for the hidden full text first, then falling back to the visible truncated text if needed. The function returns a list of dictionaries, where each dictionary contains one specification. For example, we might get \[{"Brand": "Sony"}, {"Model": "XYZ-123"}, {"Weight": "2.5 pounds"}\]. ### Orchestrating the Data Extraction The parse\_product\_info function brings all our individual parsing functions together into one comprehensive extraction process. ``` def parse_product_info(html): """ Comprehensive product information parser for Amazon product pages. Orchestrates the extraction of all relevant product information from an Amazon product page HTML by calling individual parsing functions and combining the results into a structured dictionary. Args: html (str): Complete HTML content of an Amazon product page. Returns: dict: Dictionary containing all extracted product information with keys: - title (str or None): Product title - price (str or None): Current price - original_price (str or None): Original/list price before discount - discount (str or None): Discount percentage or amount - rating (str or None): Customer rating - details (List[Dict]): List of product specification dictionaries Data Processing: - Creates BeautifulSoup object for HTML parsing - Calls individual parsing functions for each data element - Combines results into unified data structure - Handles missing elements gracefully (returns None for missing data """ soup = BeautifulSoup(html, "html.parser") return { "title": parse_title(soup), "price": parse_price(soup), "original_price": parse_original_price(soup), "discount": parse_discount(soup), "rating": parse_rating(soup), "details" : extract_table_data(html) } ``` This function takes the raw HTML of a product page and returns a dictionary containing all the information we could extract. It creates one BeautifulSoup object and passes it to all the parsing functions, making the process efficient and organized. ### Setting ZIP Code for Consistent Results Amazon shows different prices and availability based on your location. To get consistent results, we need to set a specific ZIP code before scraping. So we used same function here which used in phase 1. ``` def apply_zip(page, zip_code="56901"): """ Set delivery ZIP code on Amazon using browser automation. Navigates to Amazon's homepage and programmatically sets the delivery ZIP code through the location selector interface. This affects product availability, pricing, and shipping options throughout the session. The ZIP code setting persists through cookies for subsequent requests. Args: page (playwright.sync_api.Page): Playwright page object for browser automation. zip_code (str, optional): ZIP code to set for delivery location. Defaults to "56901". Returns: None Raises: Exception: Catches and prints any errors during ZIP code setting process. Script continues execution even if ZIP setting fails. Process Flow: 1. Navigate to Amazon homepage with 60-second timeout 2. Wait for and click location selector (glow-ingress-line2) 3. Wait for ZIP code input field to appear 4. Fill ZIP code input field with provided value 5. Submit ZIP code update form 6. Wait 3 seconds for processing 7. Handle optional "Continue" confirmation dialog 8. Print success/failure message Side Effects: - Sets location cookies that affect pricing and availability - Changes default shipping location for the browser session - May trigger location-based content personalization Note: The ZIP code affects product availability, pricing, tax calculations, and shipping options. Using a consistent ZIP code across scraping sessions ensures data consistency. """ try: page.goto("https://www.amazon.com", timeout=60000) page.wait_for_selector("span#glow-ingress-line2", timeout=10000) page.click("span#glow-ingress-line2") page.wait_for_selector("input#GLUXZipUpdateInput", timeout=10000) page.fill("input#GLUXZipUpdateInput", zip_code) page.click("#GLUXZipUpdate > span > input") page.wait_for_timeout(3000) # If "Continue" button appears after ZIP, click it try: page.click("span.a-button-inner > input[name='glowDoneButton']") except: pass print(f"📍 ZIP code set to {zip_code}") except Exception as e: print("❌ Failed to set ZIP:", e) ``` ### The Main Scraping Orchestration Now we reach the conductor of our data extraction orchestra - the main scraping function that coordinates everything. ``` def scrape_with_zip_zipcode(): """ Main orchestration function for scraping detailed product information from Amazon. This is the primary function that coordinates the entire product detail scraping process. It retrieves unscraped product URLs from the database, sets up browser automation with location-specific settings, and systematically scrapes detailed product information from each URL. Returns: None: Results are saved directly to the SQLite database tables. Process Overview: 1. Initialize database connection and check for unscraped URLs 2. Launch Playwright browser with random user agent 3. Set ZIP code (56901) for consistent location-based pricing 4. Save cookies for reuse across product page visits 5. Iterate through each unscraped product URL 6. Extract comprehensive product information 7. Store results in product_details table 8. Mark URLs as scraped to prevent reprocessing 9. Handle and log errors for failed URLs 10. Clean up browser resources Database Operations: - Queries products_url table for unscraped entries (scraped = 0) - Inserts detailed product data into product_details table - Logs failed URLs and error messages in error_urls table - Updates scraped flag to 1 for processed URLs - Commits after each product to prevent data loss Error Handling: - Catches exceptions during page loading and parsing - Logs errors with URL and exception details - Continues processing remaining URLs after failures - Stores failed URLs for later investigation Anti-Detection Measures: - Random user agent selection for each session - Cookie persistence to maintain session state - Random delays (3-5 seconds) between requests - Consistent ZIP code for location-based consistency - Realistic browsing patterns with proper timeouts Data Extracted Per Product: - Product title and description - Current price and original price - Discount information - Customer ratings - Detailed product specifications table - Category classification - Product page URL Performance Considerations: - Processes URLs sequentially to avoid overwhelming Amazon's servers - Uses browser context reuse with cookie persistence - Implements proper timeouts for page loading - Includes random delays for realistic browsing simulation Prerequisites: - Existing products_url table with URLs to scrape - user_agents.txt file with user agent strings - Playwright browser binaries installed - Stable internet connection for Amazon access Raises: Various exceptions may occur during scraping: - Network timeouts and connection errors - Page structure changes breaking selectors - Database operation errors - Browser automation failures Most exceptions are caught and logged without stopping the process. """ conn, cur = setup_database() cur.execute("SELECT id, category, url FROM products_url WHERE scraped = 0") products = cur.fetchall() print(f"🔄 Found {len(products)} unscraped URLs") ``` We start by connecting to our database and finding all the URLs that haven't been scraped yet. This is where our "scraped" column comes in handy - we can easily resume scraping from where we left off if the process gets interrupted. ``` with sync_playwright() as p: browser = p.chromium.launch(headless=False) user_agent = random.choice(user_agents) context = browser.new_context( user_agent=user_agent, locale="en-US" ) page = context.new_page() apply_zip(page, zip_code="56901") # Save cookies to reuse for each product cookies = context.cookies() page.close() context.close() ``` Just like in phase one, we set up our browser with a random user agent and configure the ZIP code. The key difference is that we save the cookies after setting the ZIP code and then close the initial browser context. We'll reuse these cookies for each product page, maintaining our location settings without having to reset the ZIP code every time. ``` for pid, category, url in products: print(f"🔍 Scraping: {url}") context = browser.new_context( user_agent=user_agent, locale="en-US" ) context.add_cookies(cookies) page = context.new_page() try: page.goto(url, timeout=60000) page.wait_for_timeout(3000) html = page.content() product = parse_product_info(html) if not product["title"]: raise Exception("Missing title") ``` For each product URL, we create a fresh browser context but add our saved cookies. This gives us the benefits of a clean slate while maintaining our ZIP code settings. We navigate to the product page, wait a moment for everything to load, and then extract all the product information. The check for a missing title is our quality control - if we can't even find a product title, something is probably wrong with the page, and we should treat it as an error. ``` cur.execute(''' INSERT INTO product_details (category, url, title, price, original_price, discount, details, rating) VALUES (?, ?, ?, ?, ?, ?,?,?) ''', (category, url, product["title"], product["price"], product["original_price"],product["discount"],str(product["details"]),product["rating"])) print(f"✅ {product['title'][:60]}") except Exception as e: print(f"❌ Error scraping {url}: {e}") cur.execute("INSERT OR IGNORE INTO error_urls (url, error) VALUES (?, ?)", (url, str(e))) cur.execute("UPDATE products_url SET scraped = 1 WHERE id = ?", (pid,)) conn.commit() page.close() context.close() time.sleep(random.uniform(3, 5)) browser.close() conn.close() print("✅ Finished scraping all with ZIP 56901") ``` When extraction succeeds, we save all the product details to our database. The specifications are stored as a string representation of the list of dictionaries - we can parse this back into structured data when we need to analyze it later. If something goes wrong, we log the error and the URL that failed. Either way, we mark the URL as scraped so we don't try to process it again. We commit the database changes immediately to avoid losing progress if the script crashes. Finally, we close the browser context and wait a random amount of time before moving to the next product. This random delay makes our scraping pattern less detectable and more respectful of Amazon's servers. ### Running the Complete System The entry point ties everything together with a simple execution guard: ``` if __name__ == "__main__": """ Entry point for the Amazon product detail scraper. Executes the main scraping function when the script is run directly. This allows the script to be imported as a module without automatically starting the scraping process. Usage: python amazon_detail_scraper.py Prerequisites: - Existing amazon_vlogging.db with products_url table populated - user_agents.txt file with user agent strings - Required packages: playwright, beautifulsoup4, sqlite3 - Playwright browser binaries installed (playwright install) Expected Workflow: 1. Run URL collection script first to populate products_url table 2. Run this script to scrape detailed product information 3. Check product_details table for results 4. Review error_urls table for any failed scraping attempts """ scrape_with_zip_zipcode() ``` This ensures that our scraping function only runs when we execute the script directly, not when it's imported as a module. ### The Complete Picture By the end of this phase, we have transformed our collection of product URLs into a comprehensive database of product information. Each product now has detailed pricing, ratings, specifications, and categorization - everything we need for meaningful analysis of the vlogging equipment market. The two-phase approach gives us flexibility and robustness. We can collect URLs in bulk when Amazon's servers are responsive, then take our time with the detailed extraction. If something goes wrong during detail extraction, we haven't lost our URL collection work. Our database now contains a rich dataset ready for analysis, visualization, or any other insights we want to extract from Amazon's vlogging equipment marketplace. The structured approach makes it easy to extend the system with additional data points or to adapt it for different product categories. ## Conclusion We've built something pretty impressive here. What started as a simple idea to make vlogging equipment research easier turned into a complete data extraction system that can handle thousands of products across multiple categories. Our two-phase approach proved its worth - we can collect URLs quickly and then take our time with the detailed extraction, handling errors gracefully along the way. The database we created contains everything you'd want to know about vlogging equipment: prices, discounts, ratings, detailed specifications, and proper categorization. But the real value here isn't just the data - it's the system we built to get it. This architecture can be adapted for any kind of product research, price monitoring, or market analysis you might need. Throughout this project, we maintained ethical scraping practices with proper delays and respectful request patterns. The code is modular and well-documented, making it easy to extend or modify for different use cases. We've demonstrated how to handle complex e-commerce sites, manage large-scale data extraction, and build fault-tolerant systems that can recover from interruptions. The techniques you've learned here - browser automation, robust error handling, database design for scraping, and anti-detection strategies - are valuable skills that apply to countless other data extraction projects. Whether you're tracking competitor prices, monitoring product availability, or building recommendation systems, this foundation gives you everything you need to succeed. What's next? You could set up scheduled runs to track price changes over time, build visualization dashboards to spot trends, or even expand the system to scrape multiple e-commerce sites for comprehensive market coverage. The structured data you now have opens up endless possibilities for analysis and insights. The best part about this project is that it solves a real problem while teaching professional-grade techniques. You're not just extracting data - you're building sustainable systems that provide lasting value. That's the difference between amateur scraping and the kind of work that actually matters in the real world. AUTHOR I’m Shahana, a Data Engineer at Datahut, where I specialize in building smart, scalable data pipelines that transform messy web data into structured, usable formats—especially in domains like retail, e-commerce, and competitive intelligence. At Datahut, we help businesses across industries gather valuable insights by automating data collection from websites, even those that rely on JavaScript and complex navigation. In this blog, I’ve walked you through a real-world project where we created a robust web scraping workflow to collect product information efficiently using Playwright, BeautifulSoup, and SQLite. Our goal was to design a system that handles dynamic pages, pagination, and data storage—while staying lightweight, reliable, and beginner-friendly. If your team is exploring ways to extract structured product or pricing data at scale—or if you're just curious how web scraping can support smarter decisions—feel free to connect with us using the chat widget on the right. We’re always excited to share ideas and build custom solutions around your data needs. ### FAQs 1\. What is Amazon product data scraping? Amazon product data scraping is the process of automatically extracting product-related information such as titles, prices, reviews, ratings, sellers, and ASINs from Amazon’s product listings. It helps businesses analyze trends, monitor competitors, and make data-driven decisions. 2\. Is it legal to scrape product data from Amazon US? Scraping Amazon data for personal or research purposes is generally acceptable, but using automated bots to access data without Amazon’s permission may violate their terms of service. It’s best to use ethical and compliant scraping practices, such as scraping publicly available data responsibly or using Amazon’s official APIs. 3\. What tools can I use to scrape Amazon product data? Popular tools and libraries include Python’s BeautifulSoup, Scrapy, Selenium, and Playwright. These tools can automate data extraction from web pages and handle dynamic content effectively. 4\. What kind of data can I extract from Amazon US? You can extract data points such as: - Product title and description - Price and discounts - Ratings and reviews - Seller information - ASIN and category details - Availability and shipping options 5\. Why should businesses scrape Amazon product data? Businesses use Amazon product data scraping to: - Track competitor pricing and discounts - Identify top-performing products - Monitor customer sentiment - Analyze market demand - Optimize product listings and pricing strategies ### Scraping Amazon Ceiling Fan Data for Analysis URL: https://www.blog.datahut.co/post/how-to-scrape-ceiling-fans-data-from-amazon/ Last updated: 2026-07-23T07:48:31.000Z Web scraping can seem overwhelming at first, but it's really just about teaching your computer to visit websites and collect information automatically. Today, we'll walk through a project that scrapes ceiling fan data from Amazon India. We'll break this down into simple, manageable steps that anyone can follow. This project works in two phases. First, we collect all the product page links from Amazon's search results. Then, we visit each of those links to gather detailed information about each ceiling fan. Let's start with phase one. ## Url Collection ### Setting Up Our Tools Before we can start collecting data, we need to import the right tools for the job. Think of this like gathering all your materials before starting a craft project. ``` import sqlite3 import asyncio from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` We're using four main tools here. SQLite helps us store our data in an organised database. Asyncio lets our program handle multiple tasks efficiently. Playwright acts like a remote control for web browsers, letting us navigate websites automatically. Beautiful Soup helps us read and understand the structure of web pages. We also set up a simple variable to name our database file. This keeps everything organised and makes it easy to find our data later. ``` # SQLite database setup DB_NAME = "amazon_products.db" ``` ### Creating Our Data Storage Every good project needs a safe place to store the information we collect. We create a database that works like a digital filing cabinet. ``` def setup_database(): """ Create a SQLite database and table if it does not exist. This function initializes the database by creating a table named `product_urls` with two columns: - `id`: An auto-incrementing primary key. - `url`: A unique text field to store product URLs. If the table already exists, this function does nothing. """ conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT NOT NULL UNIQUE ) """) conn.commit() conn.close() ``` This function creates a simple table with two columns. The first column gives each entry a unique number automatically. The second column stores the actual web addresses of the products we find. The database only accepts each URL once, which prevents us from accidentally storing duplicates. ### Saving URLs to Our Database Once we find product links, we need a way to save them safely. This function handles that task for us. ``` def save_urls_to_db(urls): """ Save a list of product URLs to the SQLite database. Args: urls (list): A list of product URLs to be saved. This function inserts URLs into the `product_urls` table. Duplicate URLs are ignored using the `INSERT OR IGNORE` statement to ensure uniqueness. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() for url in urls: cursor.execute("INSERT OR IGNORE INTO product_urls (url) VALUES (?)", (url,)) conn.commit() except sqlite3.Error as e: print(f"Database error: {e}") finally: conn.close() ``` This function takes a list of URLs and adds each one to our database. The "INSERT OR IGNORE" command means that if we try to add a URL that already exists, the database will simply skip it instead of causing an error. We wrap everything in a try-except block to handle any problems gracefully. ### Collecting Product URLs from Amazon Now comes the main event. This function visits Amazon's search results and collects links to individual product pages. ``` async def scrape_product_urls(): """ Scrape product URLs from Amazon India and handle pagination. This function uses Playwright to navigate through Amazon India's search results pages for ceiling fans. It extracts product URLs using Beautiful Soup and saves them to the SQLite database. Pagination is handled by iterating through a range of pages (1 to 401). The URLs are dynamically constructed using the `base_url` template. Steps: 1. Navigate to each page. 2. Wait for the page content to load. 3. Parse the page content using Beautiful Soup. 4. Extract product URLs from the page. 5. Save the URLs to the database. Error handling ensures graceful recovery in case of scraping or database issues. """ base_url = "https://www.amazon.in/s?k=ceiling+fans&i=kitchen&page={page}&crid=28ITASDE5GSK7&qid=1752732819&sprefix=ceiling+fans+%2Ckitchen%2C229&xpid=j43vzzw1MmBiO&ref=sr_pg_{page}" product_urls = [] ``` We start by creating a template URL that we can modify for different pages. Amazon shows search results across many pages, so we need to visit each page systematically. The curly braces in the URL act like blanks that we'll fill in with page numbers. ``` async with async_playwright() as playwright: try: browser = await playwright.chromium.launch(headless=False) context = await browser.new_context() page = await context.new_page() ``` Here we launch a browser that our program can control. We set "headless=False" which means you can actually watch the browser work. This is helpful when you're learning because you can see exactly what's happening. The real work happens in a loop that visits each page of search results. ``` for page_number in range(1, 401): # Iterate through pages 1 to 401 current_url = base_url.format(page=page_number) print(f"Scraping: {current_url}") await page.goto(current_url, timeout=60000) # Wait for page content to load await page.wait_for_selector("div[role='listitem']", timeout=60000) content = await page.content() ``` For each page number from 1 to 400, we create the full URL by inserting the page number into our template. Then we navigate to that page and wait for it to load completely. The wait\_for\_selector line ensures that the product listings have appeared before we try to extract information from them. Once the page loads, we use Beautiful Soup to find the product links. ``` # Parse the page content with Beautiful Soup soup = BeautifulSoup(content, "html.parser") product_elements = soup.select("span.rush-component > a.a-link-normal.s-no-outline") # Extract product URLs page_urls = [] for element in product_elements: href = element.get("href") if href and href.startswith("/"): page_urls.append(f"https://www.amazon.in{href}") ``` Beautiful Soup reads the page like a structured document and finds all the links that match Amazon's pattern for product pages. We look for specific CSS selectors that Amazon uses for product links. Each link we find gets added to our collection, but we need to add Amazon's domain to the beginning since the links are relative. After collecting URLs from each page, we add them to our main list and show our progress. Then save the urls using save\_urls\_to\_db() function mentioned before. ``` # Print the number of URLs scraped from the current page print(f"Number of URLs scraped from page {page_number}: {len(page_urls)}") # Add the URLs from the current page to the main list product_urls.extend(page_urls) await browser.close() except Exception as e: print(f"Error during scraping: {e}") # Save URLs to the database save_urls_to_db(product_urls) print(f"Scraped {len(product_urls)} product URLs.") ``` ### Bringing It All Together The final part of our script coordinates everything we've built. ``` if __name__ == "__main__": """ Entry point of the script. This script performs the following tasks: 1. Sets up the SQLite database. 2. Initiates the scraping process to extract product URLs. """ setup_database() asyncio.run(scrape_product_urls()) ``` When we run the script, it first sets up our database table. Then it starts the URL collection process. The asyncio.run command handles all the complex timing that makes our web scraping work smoothly. This completes phase one of our project. We now have a database filled with links to individual ceiling fan product pages on Amazon India. In phase two, we would visit each of these URLs to collect detailed information about each product, like prices, ratings, and specifications. The beauty of this approach is that we've separated the two tasks. We can run this script once to collect all the URLs, then run a different script to gather the detailed information. This makes our code more organised and easier to debug if something goes wrong. ## Data Collection Now that we have all our product URLs safely stored, it's time to visit each page and gather the specific details we want. This phase is like having a list of addresses and then visiting each house to collect information about what's inside. ### Setting Up for Detailed Data Collection Phase two starts with some additional tools that we'll need for handling the more complex data we're about to collect. ``` import sqlite3 import asyncio from playwright.async_api import async_playwright from bs4 import BeautifulSoup import json # Add this import at the top of the file import random # Add this import at the top of the file # SQLite database setup DB_NAME = "amazon_products.db" ``` We add JSON support because product specifications and details come in complex formats that need special handling. The random module helps us add small delays between requests, which is important for being respectful to Amazon's servers. ### Creating Storage for Product Details Just like we needed a place to store URLs in phase one, we need a more sophisticated storage system for all the product details we're about to collect. ``` def setup_product_table(): """ Create a SQLite table named `product_data` if it does not exist. This table stores detailed product information scraped from Amazon. Columns: - `id`: Auto-incrementing primary key. - `title`: Product title. - `price`: Product price. - `rating`: Product rating. - `reviews`: Number of reviews. - `discount`: Discount percentage. - `original_price`: Original price. - `color`: Product color. - `url`: Product URL (unique). Also ensures the `scraped` column exists in the `product_urls` table. """ conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, title TEXT, rating TEXT, reviews TEXT, price TEXT, discount TEXT, original_price TEXT, color TEXT, specifications TEXT, extra_details TEXT, about TEXT, url TEXT NOT NULL UNIQUE ) """) ``` This new table has columns for all the different pieces of information we want to collect from each product page. Some columns like specifications, extra\_details, and about will store complex information as JSON text, which lets us keep lists and detailed information organised. We also need to modify our original URL table to track which pages we've already processed. ``` # Check if the `scraped` column exists in the `product_urls` table cursor.execute("PRAGMA table_info(product_urls)") columns = [row[1] for row in cursor.fetchall()] if "scraped" not in columns: cursor.execute(""" ALTER TABLE product_urls ADD COLUMN scraped INTEGER DEFAULT 0 """) conn.commit() conn.close() ``` This code checks if our URL table already has a "scraped" column. If not, it adds one. This column acts like a checkbox system, marking each URL as either processed (1) or not yet processed (0). ### Managing Our Scraping Progress We need helper functions to keep track of which URLs we still need to process and to mark them as complete when we're done. ``` def fetch_unscraped_urls(): """ Fetch all URLs from the `product_urls` table where `scraped` is 0. Returns: list: A list of unscraped product URLs. """ conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("SELECT url FROM product_urls WHERE scraped = 0") urls = [row[0] for row in cursor.fetchall()] conn.close() return urls ``` This function looks at our URL table and returns only the URLs we haven't processed yet. It's like having a to-do list that automatically shows us what work is left. ``` def mark_url_as_scraped(url): """ Mark a URL as scraped in the `product_urls` table. Args: url (str): The URL to mark as scraped. """ conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("UPDATE product_urls SET scraped = 1 WHERE url = ?", (url,)) conn.commit() conn.close() ``` After we successfully collect data from a product page, this function marks that URL as complete. If our scraping gets interrupted, we can restart and pick up exactly where we left off. ### Saving Complex Product Data The product information we collect is much more detailed than simple URLs, so we need a more sophisticated saving function. ``` def save_product_data(data): """ Save scraped product data to the `product_data` table. Args: data (dict): A dictionary containing product details. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() # Convert the specifications list to a JSON string data["specifications"] = json.dumps(data["specifications"]) data["extra_details"] = json.dumps(data["extra_details"]) data["about"] = json.dumps(data["about"]) cursor.execute(""" INSERT OR IGNORE INTO product_data (title, rating, reviews, price, discount, original_price, color, specifications, extra_details, about, url) VALUES (:title, :rating, :reviews, :price, :discount, :original_price, :color, :specifications, :extra_details, :about, :url) """, data) conn.commit() except sqlite3.Error as e: print(f"Database error: {e}") finally: conn.close() ``` Before saving, we convert complex data like specifications into JSON format. This lets us store lists and detailed information in a way that we can easily read back later. The function uses named parameters, which makes the code clearer and safer. ### Extracting Information from Product Pages Now comes the detailed work of finding and extracting specific pieces of information from each product page. Amazon's pages have consistent patterns, so we can write functions that know exactly where to look for each piece of data. ``` async def parse_title(soup): """ Parse the product title from the HTML content. Args: soup (BeautifulSoup): Parsed HTML content. Returns: str: Product title or None if not found. """ try: return soup.select_one("h1#title > span#productTitle").get_text(strip=True) except AttributeError: return None ``` Each parsing function follows the same pattern. We use CSS selectors to find the exact location of the information we want, then extract the text content. If something goes wrong or the information isn't there, we return None instead of crashing. The CSS selectors might look complex, but they're just precise addresses that tell Beautiful Soup exactly where to find each piece of information on the page. Each function handles one specific piece of data, like title, price, or rating. ``` async def parse_product_specifications(soup): """ Parse the product specifications from the HTML content. Args: soup (BeautifulSoup): Parsed HTML content. Returns: list: A list of dictionaries containing key-value pairs for product specifications. """ specifications = [] try: # Find all rows in the specifications table rows = soup.select("table.a-normal.a-spacing-micro tr") for row in rows: # Extract the key (left column) and value (right column) key = row.select_one("td.a-span3 span.a-size-base.a-text-bold") value = row.select_one("td.a-span9 span.a-size-base.po-break-word") if key and value: specifications.append({ "key": key.get_text(strip=True), "value": value.get_text(strip=True) }) except AttributeError: pass # Handle cases where the structure is not found return specifications ``` Some functions, like this one for specifications, are more complex because they extract multiple pieces of related information. This function finds Amazon's specifications table and converts each row into a key-value pair, creating a list of all the technical details about the product. ### Coordinating the Data Extraction We need a main function that uses all our individual parsing functions to extract complete product information from each page. ``` async def parse_product_page(content, url): """ Parse the product page content and extract details. Args: content (str): HTML content of the product page. url (str): URL of the product page. Returns: dict: A dictionary containing product details. """ soup = BeautifulSoup(content, "html.parser") title = await parse_title(soup) price = await parse_price(soup) rating = await parse_rating(soup) reviews = await parse_number_of_reviews(soup) discount = await parse_discount_percentage(soup) original_price = await parse_original_price(soup) color = await parse_color(soup) specifications = await parse_product_specifications(soup) extra_details = await parse_extra_details(soup) about = await parse_about(soup) return { "title": title, "price": price, "rating": rating, "reviews": reviews, "discount": discount, "original_price": original_price, "color": color, "url": url, "specifications": specifications, "extra_details": extra_details, "about": about } ``` This function takes the raw HTML content of a product page and runs it through all our parsing functions. The result is a clean dictionary containing all the information we could extract from that page. ### Running the Complete Data Collection The main scraping function brings everything together, processing each URL in our database systematically. ``` async def scrape_product_data(): """ Scrape product data by iterating through URLs in the `product_urls` table. This function uses Playwright to navigate to each product URL, extracts data, and saves it to the `product_data` table. It also marks URLs as scraped. """ urls = fetch_unscraped_urls() if not urls: print("No unscraped URLs found.") return async with async_playwright() as playwright: try: browser = await playwright.chromium.launch(headless=False) context = await browser.new_context() page = await context.new_page() for url in urls: try: print(f"Scraping: {url}") await page.goto(url, timeout=60000) # Wait for the page content to load await page.wait_for_selector("#productTitle", timeout=60000) content = await page.content() # Parse the product page product_data = await parse_product_page(content, url) # Save the product data to the database save_product_data(product_data) # Mark the URL as scraped mark_url_as_scraped(url) # Add a random delay between 2 and 3 seconds delay = random.uniform(2, 3) print(f"Delaying for {delay:.2f} seconds...") await asyncio.sleep(delay) except Exception as e: print(f"Error scraping {url}: {e}") await browser.close() except Exception as e: print(f"Error initializing Playwright: {e}") ``` The function starts by getting all unprocessed URLs. For each URL, it navigates to the page, waits for the main content to load, extracts all the product information, saves it to our database, and marks the URL as complete. The random delay between requests is important. It makes our scraping more natural and respectful to Amazon's servers. Each delay is between 2 and 3 seconds, which gives the server time to breathe between our requests. ### Completing the Project The final coordination happens when we run the script, just like in phase one. ``` if __name__ == "__main__": """ Entry point of the script. This script performs the following tasks: 1. Sets up the SQLite database and tables. 2. Scrapes product data from URLs stored in the `product_urls` table. """ setup_product_table() asyncio.run(scrape_product_data()) ``` When we run phase two, it sets up the new database table and then processes all the URLs we collected in phase one. The beauty of this two-phase approach is that each part has a clear job. Phase one focuses on finding all the products, while phase two focuses on collecting detailed information about each one. By the end of this process, we have a complete database of ceiling fan information from Amazon India. We can analyze prices, compare ratings, study specifications, and gain insights into the ceiling fan market. The structured data we've collected opens up possibilities for analysis, comparison shopping tools, or market research that would be impossible to do manually. ## Wrapping Up You've successfully built a complete web scraping system that collects product data from Amazon India. By separating URL collection from data extraction, you created a robust and maintainable solution that can handle thousands of products efficiently. The skills you've learned here go beyond just this project. You now understand how to navigate websites programmatically, extract structured data from HTML, manage databases, and implement respectful scraping practices. These techniques can be adapted to collect data from other websites and build different types of analysis tools. Your database is now filled with valuable product information that would have taken weeks to collect manually. Whether you use this data for market research, price comparison, or trend analysis, you have the foundation to turn web data into actionable insights. FAQ Section 1) What data should I collect when scraping ceiling fans on Amazon? Collect product title, ASIN, price, list price (MRP), discounts, product images, bullet points/specs, average rating, number of reviews, seller name, shipping info, availability, dimensions/weight (if shown), category breadcrumbs, and product URL. Also capture scrape timestamp and source page HTML or snapshot for auditing. 2) Which tools & approach work best for scraping Amazon listings? Start with requests + BeautifulSoup for static pages. Use Playwright or Selenium (headless) for JS-rendered content or infinite scroll. For scale or robust anti-bot handling, use Playwright with rotating proxies, randomized user-agents, and request throttling. Prefer using Amazon Product Advertising API if you have access — it’s official and safer. 3) How do I find the right HTML elements (selectors) for ceiling fan info? Open a product page in a browser, right-click → Inspect. Look for: title (#productTitle), price (#priceblock\_ourprice or [#priceblock\_dealprice](https://www.blog.datahut.co/blog/hashtags/priceblock%5Fdealprice)), images (#imgTagWrapperId img or data-a-dynamic-image), rating (.a-icon-alt), reviews count (#acrCustomerReviewText), and specs in the “Product details” or “Technical Details” table. Use CSS selectors or XPath to extract those nodes. Always test selectors across multiple product pages—Amazon templates vary. 4) What anti-blocking and legal/ethical practices should I follow? Rate-limit your requests (random delays), use backoff on failures, rotate User-Agent strings, and rotate IPs/proxies if scraping at scale. Respect robots.txt and Amazon’s Terms of Service; prefer their official API when possible. Never scrape or store personal data (buyer info). Monitor for CAPTCHA / 503 responses and stop if the site detects you. Log your activity and add caching to reduce load on Amazon. 5) How should I store, clean, and deduplicate scraped fan data? Store raw HTML or JSON output plus parsed fields and a timestamp. Use ASIN as the unique key to deduplicate. Normalize prices (store numeric + currency), strip whitespace from titles/specs, parse numeric values (ratings, review counts), and standardize units (e.g., dimensions). Keep a version or history table if you want price/availability time series. AUTHOR I’m Shahana, a Data Engineer at Datahut, where I specialize in building smart, scalable data pipelines that transform messy web data into structured, usable formats—especially in domains like retail, e-commerce, and competitive intelligence. At Datahut, we help businesses across industries gather valuable insights by automating data collection from websites, even those that rely on JavaScript and complex navigation. In this blog, I’ve walked you through a real-world project where we created a robust web scraping workflow to collect product information efficiently using Playwright, BeautifulSoup, and SQLite. Our goal was to design a system that handles dynamic pages, pagination, and data storage—while staying lightweight, reliable, and beginner-friendly. If your team is exploring ways to extract structured product or pricing data at scale—or if you're just curious how web scraping can support smarter decisions—feel free to connect with us using the chat widget on the right. We’re always excited to share ideas and build custom solutions around your data needs. ### Scraping Amazon Dog Food Data Using Python URL: https://www.blog.datahut.co/post/how-to-scrape-amazon-dog-food-using-python-libraries/ Last updated: 2026-07-23T07:48:31.000Z Have you ever wondered how major e-commerce platforms manage thousands of product listings across categories like electronics, fashion, or even pet food? One of the biggest names in this space, [Amazon](https://www.amazon.com/?ref=blog.datahut.co), holds an enormous inventory that spans nearly every product imaginable—earning it the title “The Everything Store.” Founded in 1994 by Jeff Bezos, Amazon began as a humble online bookstore and grew into one of the world’s most influential tech giants. Beyond e-commerce, its reach extends into cloud computing (AWS), digital streaming (Prime Video, Audible), consumer electronics (Kindle, Fire TV), and AI innovations. With over 4000+ products listed in just one category—dog food—it becomes clear why many businesses turn to web scraping to collect and analyze such vast information efficiently. This blog walks through a beginner-friendly journey into web scraping, focusing on the dog food section of Amazon US. If you're a Data Analyst, or Product Researcher, and you're curious about how to gather large-scale data to improve decision-making or stay ahead of the competition, you're in the right place. Whether you're just exploring or actively planning a data-driven strategy, you’ll learn how structured scraping methods can transform unmanageable volumes of online data into usable, insightful formats for data analysis, pricing strategies, trend monitoring, and much more. Curious how data can sharpen your edge in the market? Let’s dive into the scraping process and unlock the real value behind e-commerce product data. Want to supercharge your business decisions with real-time insights? Reach out today and let’s make data work for you. ## Smart and Automated Data Gathering What’s the best way to gather information when a website has thousands of product listings spread across multiple pages? That’s where web scraping becomes a game-changer. The process begins by collecting all the product URLs from a chosen category—like dog food on Amazon US—and then visiting each of those links to extract useful details such as product names, prices, ratings, and more. This structured data collection method helps turn scattered online content into organized datasets ready for data analysis, empowering e-commerce teams, product managers, and analysts to make smarter business decisions. ### Step 1: Collecting Product URLs Have you ever noticed how product listings on Amazon don’t appear all at once? Instead, they’re spread across multiple pages, especially in large categories. To collect all the product links, the first step was to change the delivery location from India to a U.S. zip code—10001, New York—so that only relevant U.S. listings would appear. Once the location was set, the scraper navigated through each paginated page of the dog food section, capturing every product URL displayed. A browser automation tool like Playwright played a key role here. It behaved like a real user—opening pages, waiting for content to load, and clicking the “Next” button to move through each page. Along the way, it handled pop-ups and ensured that only clean, working links were gathered. These product URLs—over 4,000 in total—were saved neatly into a SQLite database, making it easier to manage, avoid duplicates, and prepare for the next step: collecting detailed product information from each link. This structured setup ensures the data remains organized and ready for smooth analysis later. ### Step 2: Collecting Information from Product URLs So what happens after collecting thousands of product URLs? The next step is to visit each link and carefully extract detailed information about every item listed. For accurate results, the delivery location on Amazon was first set to New York, zip code 10001, so that only U.S.-relevant product details would appear. Then, each page was opened one by one using Playwright, which behaves like a real user—waiting for content to load, handling dynamic elements, and scrolling where needed. This approach helped gather key product attributes like the title, price, brand, number of reviews, availability, description, and more. All of this information was directly saved into the same database used earlier for the URLs, keeping everything neatly organized in one place. This setup not only simplifies data management but also ensures there’s a clear link between each product and its details. With structured and reliable data in hand, the foundation is now set for deeper data analysis, such as comparing prices, identifying trends, or tracking brand performance. ### Step 3: Cleaning the Extracted Amazon Dog Food Data Have you ever opened a spreadsheet full of messy, inconsistent data and wondered how anyone makes sense of it? That’s a common situation after collecting product information from large websites like Amazon. Even when every page is visited carefully and data is stored correctly, the result can include unwanted duplicates, inconsistent formatting, and symbols like the dollar sign cluttering up price fields. For example, the same product might appear under slightly different URLs, leading to multiple identical entries that need to be removed. To clean and prepare the data, tools like OpenRefine are incredibly useful. It’s like a smarter, more powerful version of Excel, letting you spot duplicates, fill in missing values with “N/A,” and fix typos or inconsistent brand names with just a few clicks. And for more advanced cleanup, especially when dealing with large datasets or hidden formatting issues, Python’s pandas library is an excellent choice. It works behind the scenes to strip out extra symbols, tidy up text, and format numbers so they’re easy to analyze. Cleaning might not feel as exciting as data collection, but it’s a critical step—because well-prepared data is what turns raw information into reliable insights. ## Comprehensive Tool-kits for Efficient Data Extraction What’s the secret to collecting data from thousands of product pages without losing speed or structure? The answer lies in combining the right tools and libraries. When it comes to web scraping at scale, efficiency and reliability go hand in hand—and that’s only possible with a well-chosen tech stack. This setup uses a group of powerful Python libraries that automate browsing, extract content, store data, and handle background tasks—all working together like parts of a well-oiled machine. At the heart of this system is asyncio, which allows multiple operations to run at the same time. Instead of waiting for each product page to load and finish one by one, asyncio keeps things moving—helping to process hundreds of pages without slowing down. Alongside this is Playwright’s async API, a modern browser automation library that opens web pages, scrolls, clicks, and extracts content as if a human were doing it. Combined with Playwright Stealth, the tool becomes even more powerful by masking automation patterns to avoid detection on websites with anti-bot mechanisms. Once the product data is collected, it needs to be stored in a clean and organized way. For this, sqlite3 is used to manage URLs and structured information in a lightweight database. It’s simple, fast, and doesn’t need a separate server—making it ideal for scraping workflows. For storing detailed product data with flexible structure, MongoDB is a better fit, and can easily handle complex, unstructured content like descriptions, reviews, and specifications using pymongo. Supporting tools like BeautifulSoup help extract text from HTML, while modules such as random, logging, and pathlib play vital roles behind the scenes—mimicking natural delays, managing error logs, and organizing data files. Together, this combination of libraries builds a resilient and scalable scraping system that can collect, process, and store thousands of data points with precision. With these tools working in harmony, even large-scale data extraction—from dynamic e-commerce sites to paginated product catalogs—becomes not only possible, but efficient and well-managed. ## Step 1: Scraping Product URLs from the Dog Food Section on Amazon ### Importing Libraries ``` import asyncio import random import sqlite3 import logging from pathlib import Path from bs4 import BeautifulSoup from urllib.parse import urljoin from playwright.async_api import async_playwright ``` Ever wondered how a script knows how to browse a website, click buttons, or collect data like a human would? It all starts with the right set of Python libraries—prebuilt tool-kits that save time and make automation easier. In this project, several libraries work together to make the process of collecting product URLs from Amazon’s dog food section smooth and efficient. The most important tool here is playwright.async\_api, a browser automation library that simulates real user behavior. It can open Amazon pages, scroll, wait for content to load, and even handle dynamic elements—all without manual clicks. This makes it ideal for interacting with paginated product listings that require multiple actions to reveal all items. To keep track of all the URLs being collected, sqlite3 is used as a lightweight database. It stores each product link in an organized format, making it easy to manage and access later. This database lives right on the local system, making it both portable and efficient for quick lookup. Meanwhile, logging is set up to create detailed records of what the script is doing—whether it’s successfully collecting data or running into issues. Each log entry is timestamped, helping developers understand what happened and when. Additional tools like BeautifulSoup help parse HTML content, urljoin builds complete URLs, and random introduces slight delays between actions, mimicking human browsing and reducing the chance of getting blocked. Together, these libraries form the backbone of a smart, stable web scraping setup. With each one playing a specific role, they make the data collection process faster, safer, and more reliable—especially when dealing with large volumes of dynamic content. ### Getting Started: Where to Save, What to Log, and Where Scraping Begins ``` # Config paths USER_AGENTS_PATH = "/home/anusha/Desktop/DATAHUT/Macys_clothing/user_agents.txt" DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon/Data/US/dogfood_us1.db" LOG_PATH = "/home/anusha/Desktop/DATAHUT/Amazon/Data/US/dogfood_scraper1.log" # Amazon URLs BASE_URL = "https://www.amazon.com" START_URL = "https://www.amazon.com/gp/browse.html?node=2975359011&ref_=nav_em__sd_df_0_2_21_4" """ Configuration Constants This section defines important configuration variables used throughout the script. 1. USER_AGENTS_PATH:    - Path to a text file containing a list of user agent strings.    - A user agent string simulates a specific browser/device when making requests to the website.    - Helps reduce the chances of getting blocked by rotating user agents.    - Example entry in file: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/... 2. DB_PATH:    - Full path to the SQLite database file where scraped product URLs will be saved.    - The database ensures persistent and duplicate-free storage of links. 3. LOG_PATH:    - Path to the file where all logs (info, warnings, and errors) will be written.    - Useful for debugging and tracking the scraping process. 4. BASE_URL:    - The root domain of the website being scraped, in this case, Amazon US.    - Used to build absolute URLs when only relative links are found on the page. 5. START_URL:    - The first URL to begin scraping from.    - This is the Amazon category page for dog food.    - The scraper will start from this URL and then follow pagination to scrape further pages. """ ``` How does a scraper know where to start, where to save data, or how to act like a real browser? The answer lies in its configuration settings—a set of predefined paths and constants that give structure to the entire process. These settings are like the map and tools needed before setting off on a data collection journey. They guide the script on what page to begin with, where to save its progress, and how to reduce the chances of being blocked. To begin with, a user agent file is provided at a specific path. This text file contains different browser signatures that the scraper can rotate through. Every time the scraper makes a request, it can pretend to be a different browser or device—like Chrome on Windows or Safari on iPhone—helping it blend in and stay under the radar. This is especially useful for websites like Amazon, which are quick to detect automated behavior. Next is the database path, pointing to a SQLite file that acts as a storage unit for all the collected product URLs and data. Unlike temporary memory, this file keeps everything safe even if the process is interrupted. It also helps avoid duplicates by tracking what’s already been saved. Then there’s the log file path, which records every step the script takes—successes, warnings, errors, and even timestamps. These logs act like a black box, making it easier to troubleshoot issues or review scraping performance. Lastly, the base URL and start URL set the foundation for navigation. The base URL defines the main site—https://www.amazon.com—while the start URL leads directly to the dog food category. From there, the scraper knows where to begin and how to explore further using pagination. Together, these configuration paths keep the process well-structured, transparent, and ready for large-scale data extraction. ### Setting Up Amazon Cookies to Act Like a Real User ``` # Predefined cookies to simulate a session  cookies = [    {"name": "i18n-prefs", "value": "USD", "domain": ".amazon.com", "path": "/"},    {"name": "lc-main", "value": "en_US", "domain": ".amazon.com", "path": "/"},    {"name": "session-id", "value": "131-4818556-2161121", "domain": ".amazon.com", "path": "/"},    {"name": "session-id-time", "value": "2082787201l", "domain": ".amazon.com", "path": "/"},    {"name": "ubid-main", "value": "134-2297602-3092101", "domain": ".amazon.com", "path": "/"} ] """ Amazon uses cookies to manage sessions, regional preferences, and user-specific settings. By setting these cookies manually in the browser context, we simulate a session that: 1. Prevents Amazon from redirecting to a different country site. 2. Loads pages with English language and USD currency preferences. 3. Appears more like a real user session to avoid bot detection. Each cookie is a dictionary with the following keys: - "name": The name of the cookie (e.g., "i18n-prefs"). - "value": The value associated with that cookie. - "domain": The domain to which the cookie applies ("amazon.com"). - "path": The path within the domain where the cookie is valid (usually "/"). """ ``` When accessing websites like Amazon, simply sending automated requests often isn’t enough. That’s because Amazon, like many large e-commerce platforms, closely monitors browsing behavior to detect bots. It uses cookies—tiny pieces of data stored by your browser—to track things like region, currency, and session details. If these aren’t set properly, the website may behave differently or even block access. That’s why configuring predefined cookies is an important step in building a scraper that mimics real user behavior. By manually setting cookies, the scraper can simulate a realistic browsing session. For example, specifying cookies like "i18n-prefs" and "lc-main" ensures that pages load in English and display prices in USD—which is important when targeting the U.S. Amazon site. Other cookies, such as "session-id" and "ubid-main", give the scraper a consistent session identity, helping it appear more like a regular shopper instead of a script. These details reduce the risk of being redirected to a different country site or triggering anti-bot defenses. Each cookie is defined as a small dictionary that includes the name, value, domain, and path. Together, these pieces act like an invisible user passport—helping the scraper blend in, stay on the correct site version, and avoid unnecessary blocks. For anyone building reliable web scraping workflows, especially for data analysis and product tracking on Amazon, managing cookies is a quiet but powerful way to keep the process smooth and uninterrupted. ### Logging Configuration ``` # Logging Configuration logging.basicConfig(filename=LOG_PATH, level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") """ Logging Configuration This setup initializes a basic logging system that records events during the scraping process. It helps track the scraper’s progress, debug issues, and maintain a record of what happened during execution. Configuration parameters: - filename: The full path to the log file where messages will be saved (LOG_PATH). - level: The minimum severity level of messages to log. 'INFO' logs info, warnings, and errors. - format: The layout of each log message. It includes:    - %(asctime)s: Timestamp of the log entry.    - %(levelname)s: The severity of the log message (INFO, ERROR, etc.).    - %(message)s: The actual log message. """ ``` When collecting data from hundreds or thousands of web pages, how do you keep track of what your script is doing behind the scenes? This is where logging becomes essential. A well-configured logging system acts like a live journal for your scraper—recording everything from routine progress to unexpected errors. In this setup, Python’s logging.basicConfig() is used to capture important events in a clean and readable format. The log file stores entries with a timestamp, the type of message (like INFO or ERROR), and a short description of what happened. By setting the log level to INFO, the system records not only problems but also routine actions—making it easier to understand where the scraper is in its workflow. This becomes incredibly useful when reviewing large-scale data extraction tasks, debugging failed requests, or simply ensuring that each product page was processed as expected. With proper logging, you gain visibility and control—two things every e-commerce analyst or data engineer needs for reliable automation. ### Database Management ``` # SQLite Database Initialization def init_db():    """    Initializes the SQLite database by creating the `product_urls` table if it does not exist.    This table stores unique product URLs scraped from Amazon.    """    with sqlite3.connect(DB_PATH) as conn:        cursor = conn.cursor()        cursor.execute("""            CREATE TABLE IF NOT EXISTS product_urls (                id INTEGER PRIMARY KEY AUTOINCREMENT,                url TEXT UNIQUE            )        """)        conn.commit() ``` When working with large-scale web scraping projects, managing the growing list of collected URLs becomes just as important as extracting the data itself. To keep things organized and avoid duplication, a SQLite database is often used. It’s lightweight, easy to set up, and doesn’t require any server—making it perfect for e-commerce data collection tasks. In this setup, a function is used to initialize the database by creating a table named product\_urls if it doesn’t already exist. Each URL is stored uniquely, thanks to a constraint that prevents the same link from being saved more than once. This clean structure ensures that no product page is visited twice, which saves both time and resources. For data analysts, developers, or product teams, it’s a practical way to maintain control over a constantly growing dataset. ### Setting the ZIP Code for Accurate Amazon Data ``` # Change ZIP Code Function async def change_zip_code(page):    """    Change the delivery ZIP code on Amazon to ensure region-specific product availability.    Purpose:    Amazon displays different products based on the delivery ZIP code (location). This function:    - Simulates a user manually changing the delivery address to a specific ZIP code (10001 - New York).    - Ensures that product listings are consistent and not filtered out due to regional shipping restrictions.    - Helps retrieve more complete product listings during scraping.    How it works:    1. Click the location link in the top navigation bar.    2. Waits for the ZIP code input modal to appear.    3. Inputs the ZIP code `10001` and submits the change.    4. Waits for confirmation and checks if the ZIP code was successfully applied.    It logs each step, and handles errors gracefully in case any interaction fails.    - A wait timeout is added for each step to ensure elements have time to load.    - If the ZIP code change is successful, a success log is generated. If not, a warning is logged.    - If any error occurs during the process, it is caught and logged as an error.    """    logging.info("Changing delivery zip code to 10001 (New York)")    try:        await page.wait_for_selector("#nav-global-location-popover-link", timeout=15000)        await page.click("#nav-global-location-popover-link")        logging.info("Clicked delivery location button")        await page.wait_for_selector("#GLUXZipUpdateInput", timeout=15000)        await page.fill("#GLUXZipUpdateInput", "10001")        logging.info("Entered zip code 10001")        await page.wait_for_selector("#GLUXZipUpdate span input[type='submit']", timeout=10000)        await page.click("#GLUXZipUpdate span input[type='submit']")        logging.info("Clicked Apply button")        await page.wait_for_selector("button[name='glowDoneButton']", timeout=10000)        await page.click("button[name='glowDoneButton']")        logging.info("Clicked Done button")        await page.wait_for_selector("#glow-ingress-line2", timeout=10000)        delivery_text = await page.inner_text("#glow-ingress-line2")        if "10001" in delivery_text:            logging.info(f"Delivery location successfully set to: {delivery_text.strip()}")        else:            logging.warning(f"Delivery location not updated as expected. Current: {delivery_text.strip()}")        await page.wait_for_timeout(2000)    except Exception as e:        logging.error(f"Failed to change zip code: {e}") ``` Have you ever noticed how Amazon shows different products depending on where you’re located? That’s because availability, pricing, and even product listings often depend on your delivery ZIP code. So, if a scraper runs without setting a U.S. ZIP code, it might miss out on items that are only shown to American customers. To avoid this, a ZIP code update function can be used to simulate a user manually setting their location—specifically to New York’s 10001 ZIP code. This is more than just a minor tweak. Automating the ZIP code change ensures the scraper collects region-specific results, reflecting what a real user in New York would see. The function interacts with Amazon’s page by clicking the location button, entering the new ZIP, applying it, and confirming the update. Each step includes a small wait time to match human-like behavior, reducing the risk of getting blocked or receiving incomplete data. By logging every action—from opening the ZIP input to confirming the location—this function ensures the scraping session starts on solid ground. Whether you’re a data analyst trying to understand market availability or a product manager tracking competitors, this simple yet powerful detail helps ensure consistency and accuracy in your data collection workflow. ### Smart Amazon Scraper: Human-Like Browsing for Accurate Product Data ``` # Scrape Amazon Function async def scrape_amazon():    """    Main Scraping Function for Extracting Amazon Product URLs (Dog Food Category)    Overview:    ---------    This is the main coroutine responsible for the entire scraping process. It uses Playwright (headless browser automation)    and BeautifulSoup (HTML parsing) to extract product URLs from the Amazon dog food category.    Key Steps:    ----------    1. Database Initialization:       - Creates a local SQLite database (if not already created) and a table to store product URLs uniquely.    2. User-Agent Rotation       - Loads a list of user-agent strings from a local file and selects one at random.       - This helps simulate different browsers and reduces the chances of getting blocked.    3. Launch Browser Using Playwright       - Starts a Firefox browser (in visible mode for debugging, `headless=False`).       - Opens a new context and page for interaction.    4. Set Cookies and Headers       - Injects predefined session cookies to simulate a real user session.       - Sets a random user-agent header for browser requests.    5. Navigate to Start URL       - Opens the Amazon dog food category page.       - Calls `change_zip_code()` to set the delivery location to ZIP code `10001` (New York).    6. Scraping Loop (Pagination)       - Continues visiting each product listing page until no "Next" button is found.       - On each page:         a. Waits a random delay (to mimic human behavior).         b. Loads the HTML content and parses it with BeautifulSoup.         c. Searches for product anchor tags using multiple CSS selectors.         d. Cleans and normalizes product URLs (ensuring absolute URLs).         e. De-duplicates and inserts product URLs into the SQLite database using `INSERT OR IGNORE`.    7. Pagination Logic       - Checks for the "Next" page using common selectors.       - If the selector fails, use JavaScript evaluation as a fallback.       - Continues scraping until no more next pages are found.    8. Cleanup       - After scraping is complete, close the browser session gracefully.    """    init_db()    # Load user agents    with open(USER_AGENTS_PATH, "r") as f:        user_agents = [line.strip() for line in f.readlines() if line.strip()]    async with async_playwright() as p:        browser = await p.firefox.launch(headless=False)        context = await browser.new_context()        page = await context.new_page()        # Set cookies for session continuity        await context.add_cookies(cookies)        # Apply random user agent to reduce blocking risk        ua = random.choice(user_agents)        await page.set_extra_http_headers({"User-Agent": ua})        # Connect to DB        with sqlite3.connect(DB_PATH) as conn:            cursor = conn.cursor()            # Start scraping            next_page = START_URL            await page.goto(next_page, timeout=60000)            await change_zip_code(page)            while next_page:                logging.info(f"Scraping page: {next_page}")                await page.goto(next_page, timeout=120000)                await page.wait_for_timeout(random.randint(3000, 6000))                html = await page.content()                soup = BeautifulSoup(html, "html.parser")                product_links = []                # Extract product URLs (main + sponsored)                selectors = [                    "div.a-section.a-spacing-none.a-spacing-top-small.s-title-instructions-style a.a-link-normal",                    "a.a-link-normal.s-line-clamp-3.s-link-style.a-text-normal"                ]                for selector in selectors:                    for a_tag in soup.select(selector):                        href = a_tag.get("href")                        if href:                            # Clean URL: remove multiple BASE_URL prefixes                            if href.startswith("https://"):                                cleaned_url = href                                if cleaned_url.startswith(BASE_URL + "/https://"):                                    cleaned_url = cleaned_url.replace(BASE_URL + "/", "")                            else:                                cleaned_url = urljoin(BASE_URL, href)                            product_links.append(cleaned_url)                # De-duplicate and insert                logging.info(f"Found {len(product_links)} product URLs on page")                for url in set(product_links):                    try:                        cursor.execute("INSERT OR IGNORE INTO product_urls (url) VALUES (?)", (url,))                        conn.commit()                    except Exception as e:                        logging.error(f"Failed to insert URL: {url} | Error: {e}")                                 # Attempt to find the 'Next' button for pagination                try:                    next_page_tag = await page.query_selector("a.s-pagination-next")                    if not next_page_tag:                        next_page_tag = await page.query_selector("a.s-pagination-item.s-pagination-next.s-pagination-button.s-pagination-button-accessibility.s-pagination-separator")                                       if next_page_tag:                        href = await next_page_tag.get_attribute("href")                        if href:                            next_page = urljoin(BASE_URL, href)                            logging.info(f"Next page found: {next_page}")                        else:                            next_page = None                    else:                        # Fallback using JS evaluation                        next_href = await page.evaluate("""() => {                            const next = document.querySelector('a.s-pagination-next');                            return next ? next.href : null;                        }""")                        if next_href:                            next_page = next_href                            logging.info(f"Next page found via JS evaluation: {next_page}")                        else:                            next_page = None                            logging.info("No more pages found. Scraping completed.")                except Exception as e:                    logging.error(f"Pagination extraction failed: {e}")                    next_page = None                # Add wait to mimic human-like behavior before next page                await page.wait_for_timeout(random.randint(4000, 8000))        await browser.close() ``` When collecting product data from massive e-commerce platforms, precision and planning are key. A simple request to load a page and extract links often isn't enough—Amazon constantly changes how data is presented, uses region-based filtering, and includes anti-bot mechanisms. That’s where an advanced scraper, driven by tools like Playwright and BeautifulSoup, comes into play. This setup isn’t just about grabbing data; it’s about thinking like a browser, acting like a human, and adapting like a smart assistant. At the heart of the process is a function that carefully simulates a user’s journey through the Amazon dog food category. It begins by preparing a local SQLite database to store product URLs without duplication. To behave more naturally online, the scraper randomly chooses a user-agent string—this makes it look like the request is coming from a real browser on someone’s laptop or phone. It also sets session cookies to avoid redirection and triggers a custom ZIP code update to 10001 (New York), ensuring the page shows the intended regional results. Once on the category page, the scraper patiently browses through each section. Using BeautifulSoup, it parses the HTML, finds product links using multiple selectors, and then standardizes them into clean, usable URLs. These URLs are checked to avoid duplicates and stored into the database with care. The scraping loop continues as long as a "Next" button is found—either directly via page elements or through a smart fallback using JavaScript evaluation. What makes this scraper efficient isn't just how much it collects, but how gracefully it handles the flow. Logging is used throughout to track activity, errors are caught with minimal disruption, and random delays are introduced between actions to mimic human-like behavior. Once there are no more pages left to visit, the browser session is closed neatly. For anyone managing e-commerce analytics or conducting competitor research, this kind of setup offers a reliable and structured way to access valuable product data—accurately, safely, and at scale. ### Script Entry Point ``` # Entry Point of the Scraper Script if name == "__main__":    asyncio.run(scrape_amazon()) """ Entry Point of the Scraper Script This block ensures that the asynchronous `scrape_amazon()` function is executed only when the script is run directly, not when it is imported as a module. Key Concepts: ------------- - `__name__ == "__main__"`:    This condition checks if the current script is being run as the main program.    If true, it executes the scraping function.    If the script is imported into another Python script, this block will be skipped. - `asyncio.run(scrape_amazon())`:    This starts the asynchronous scraping coroutine.    It handles setting up the event loop, running the task, and closing the loop when done. This will begin scraping product URLs from the Amazon dog food category and store them in a SQLite database. """ ``` When building a web scraper using Python, one small but important line often determines whether the script actually runs or just sits idle. That line is: if name == "\_\_main\_\_":. At first glance, it might look a bit mysterious. But in simple terms, this line acts like the front door of the program—it checks if someone is running the script directly or just peeking inside it from another file. Only when the script is run directly will the scraping process begin. Under that condition, the function asyncio.run(scrape\_amazon()) is called. This is what launches the main engine of the scraper. Since the scraping function is asynchronous (meaning it performs multiple actions efficiently without waiting around), it needs a special loop to run properly. That’s exactly what asyncio.run() provides. It opens the loop, runs the scrape\_amazon() task, and then closes the loop neatly when everything is finished. This setup is especially useful when organizing code into multiple files or modules. If the script is imported elsewhere, maybe for testing or for use in a larger pipeline, the scraping won’t automatically run. It will only activate when you execute the file directly. This simple structure makes the scraper more reusable, cleaner to manage, and easier to extend later. ## Step 2: Extracting Complete Product Details from Each URL ### Importing Libraries ``` import asyncio import json import random import sqlite3 import logging from pathlib import Path from bs4 import BeautifulSoup from playwright.async_api import async_playwright ``` The script begins, as expected, by importing essential built-in and third-party Python modules—such as sqlite3, asyncio, playwright, and BeautifulSoup—which provide the core functionality for database handling, asynchronous operations, browser automation, and HTML parsing. ### Essential Paths ``` # Database and file paths DB_PATH = "/home/anusha/Desktop/DATAHUT/Amazon/Data/US/dogfood_us1.db" USER_AGENTS_PATH = "/home/anusha/Desktop/DATAHUT/Macys_clothing/user_agents.txt" OUTPUT_JSON = "/home/anusha/Desktop/DATAHUT/Amazon/Data/US/product_data.json" LOG_PATH = "/home/anusha/Desktop/DATAHUT/Amazon/Log/US/amazon_data_scraper.log" """ These constants define paths and filenames used throughout the scraper: - DB_PATH: Location of SQLite database storing URLs and scraped product data. - USER_AGENTS_PATH: Text file with a list of user agent strings. - OUTPUT_JSON: Final output file to store structured product data. - LOG_PATH: File path for storing logs. """ ``` Every well-structured scraping project begins with a solid configuration. Think of it as setting the foundation before constructing a building. In this case, a few constants keep everything in order: the DB\_PATH points to where all scraped product URLs and details are stored safely in a SQLite database. The USER\_AGENTS\_PATH holds a list of browser identities to rotate during scraping, helping avoid detection. Once data is gathered, it’s saved in a clean and structured format using OUTPUT\_JSON, making it easy to analyze later. And to track the entire scraping journey—every step, error, or success—LOG\_PATH ensures everything is recorded in a neat log file. These paths may seem like small details, but they’re the backbone of a scraper that runs smoothly and reliably. ### Applying Amazon Cookies to Mimic Human Browsing ``` # Amazon Session Cookies cookies = [    {"name": "i18n-prefs", "value": "USD", "domain": ".amazon.com", "path": "/"},    {"name": "lc-main", "value": "en_US", "domain": ".amazon.com", "path": "/"},    {"name": "session-id", "value": "131-4818556-2161121", "domain": ".amazon.com", "path": "/"},    {"name": "session-id-time", "value": "2082787201l", "domain": ".amazon.com", "path": "/"},    {"name": "ubid-main", "value": "134-2297602-3092101", "domain": ".amazon.com", "path": "/"} ] """ Amazon uses cookies to store session and localization information. These cookies simulate a persistent session for scraping with consistent regional settings. """ ``` To make sure the scraper views consistent results every time, predefined session cookies are used. These cookies help simulate a real user's experience—maintaining USD currency, English language, and stable session behavior across pages. This approach improves accuracy and reduces the chances of getting blocked during web scraping. ### Mimicking Real Browsers with Custom HTTP Headers ``` # HTTP Headers headers_template = {    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9",    "Accept-Language": "en-US,en;q=0.9",    "Connection": "keep-alive",    "Upgrade-Insecure-Requests": "1", } """ HTTP Request Headers Template This dictionary defines custom HTTP headers that mimic those sent by real web browsers (like Chrome or Firefox). These headers are passed to the Playwright browser context to make automated requests appear more like genuine human browsing, reducing the risk of bot detection by Amazon. Headers Explained: ------------------ 1. "Accept":   - Tells the server what content types the client (browser) can process.   - The values here indicate support for standard HTML, XHTML, XML, and modern image formats like WebP and AVIF.   - Helps ensure the server responds with a fully formatted product page. 2. "Accept-Language":   - Indicates the preferred languages for content.   - "en-US,en;q=0.9" means US English is preferred, and any English dialect is acceptable as a fallback. 3. "Connection":   - "keep-alive" allows the connection to stay open for multiple requests, improving speed and mimicking real browser behavior. 4. "Upgrade-Insecure-Requests":   - Set to "1" to signal that the browser supports secure HTTPS connections and prefers them over HTTP. Why These Headers Are Important: -------------------------------- - Amazon and other large websites often use bot detection systems that analyze request headers. - Sending requests without proper headers (or with default ones) can quickly result in CAPTCHA challenges or IP blocking. - These headers help bypass basic bot detection by mimicking real users' network behavior. """ ``` In web scraping, blending in like a real user is half the battle—and that starts with how requests are made. When browsers visit a website like Amazon, they send HTTP headers that carry important context about the request. These headers tell the server what kind of content is expected, which language is preferred, how the connection should behave, and more. By manually setting headers like "Accept", "Accept-Language", and "Connection", the scraper behaves more like an actual browser. This makes the requests look natural, reducing the chance of being flagged or blocked. These small but powerful details help ensure smoother scraping sessions, especially on platforms known for strong bot protection. Want to avoid detection while gathering data? Setting the right headers is one of the smartest first steps. ### Logging Configuration ``` # Logging Configuration logging.basicConfig(filename=LOG_PATH, level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') """ Logs messages to a file for debugging and auditing. Logs include timestamps, severity levels, and messages. """ ``` Keeping track of what happens during a scraping session is essential, especially when working with complex websites like Amazon. That’s where logging comes in. By configuring a logging system with timestamps and message levels (like INFO or ERROR), it's easier to monitor progress, catch issues early, and understand exactly when something went wrong. All logs are saved to a file, creating a clear and searchable record of the scraping process for later review. ### Rotating User Agents to Avoid Detection ``` # Helper Function to Get Random User Agent def get_random_user_agent():    """    Returns a random user agent string from the user_agents.txt file.    Used to simulate requests from different browsers/devices.    """    with open(USER_AGENTS_PATH) as f:        user_agents = f.read().splitlines()    return random.choice(user_agents) ``` Websites often identify visitors by their browser’s identity, known as a user agent. This tiny piece of information tells the site whether you're using Chrome on Windows, Safari on iPhone, or something else entirely. But for web scraping, sending the same user agent again and again is like knocking on a door wearing the same outfit every time—eventually, someone notices. To solve this, a helper function is used to randomly choose a user agent from a list stored in a text file. Every time the script runs, it picks a different one—like rotating disguises—to avoid detection. This not only helps the scraper look more like a human visitor but also makes it more resilient against basic anti-bot filters. It’s a small move with a big impact, especially on websites that closely watch for suspicious patterns. Need your scraper to fly under the radar? Rotating user agents is a must-have tactic. ### Smart Text Extraction with Error-Free Handling ``` # Helper Function to Extract Text or Return None def extract_text_or_none(soup, selector):    """    Extracts text from the first HTML element matching the given CSS selector.    Returns None if the element is not found.    """    tag = soup.select_one(selector)    return tag.get_text(strip=True) if tag else None ``` Another useful trick in the scraper’s toolkit is handling missing or unpredictable data. Web pages often have inconsistent layouts—sometimes an element is there, and sometimes it's not. To deal with this gracefully, a simple helper function is used to extract text from a web element only if it exists. Instead of the script breaking when a product title or price isn’t found, this function quietly returns None, allowing the scraper to move on without interruption. It works by searching for the first HTML tag that matches a given CSS selector, and if found, it neatly pulls out the text. If not, it returns nothing—avoiding crashes and saving time on debugging. This little safeguard helps ensure the scraping process remains stable, even when the web page doesn't always behave as expected. For any team working with dynamic content extraction, this type of function is essential for building resilient, production-ready data pipelines. ### Extracting List Price Without Noise ``` # Strict List Price Extraction def extract_list_price_strict(soup):    """    Extracts the list price only from the target div containing List Price text.    Returns None if not found.    """    div = soup.select_one("div.a-section.a-spacing-small.aok-align-center")    if div:        span = div.select_one("span.a-price.a-text-price span.a-offscreen")        if span:            return span.get_text(strip=True)    return None ``` To go a step further in extracting product pricing, another helper function focuses specifically on capturing the price—that’s the original price shown before any discount is applied. On many e-commerce sites like Amazon, this value is nested deep inside styled HTML blocks, which can be tricky to navigate. This function carefully targets the exact HTML structure where Amazon places the list price, ensuring we don’t accidentally pick up discounted or promotional prices instead. It first looks for a div container that usually holds price information, then drills down into the nested span tags where the actual dollar value is displayed. If the element isn't present—like in cases where no original price is listed—it simply returns None without causing any disruption. This targeted extraction adds another layer of precision to the scraping process, helping analysts capture price comparisons accurately and making data analysis far more reliable. ### Extracting Complete Product Details from an Amazon Page ``` # Product Scraper Function async def scrape_product(page, url):    """    Scrapes detailed product information from a single Amazon product page.    Overview:    ---------    This function automates visiting a given Amazon product URL using a headless browser    (via Playwright), interacts with the page to set a delivery ZIP code (ensuring availability    data is accurate), and extracts structured information using BeautifulSoup.    Purpose:    --------    - Mimic human behavior to avoid Amazon's anti-bot detection.    - Ensure delivery location is set to ZIP code `10001` (New York) for consistent product visibility.    - Extract product-related data from the rendered HTML, even if dynamic content is involved.    - Return clean and structured data in dictionary form for further storage or analysis.    Function Workflow:    ------------------    1. Visit the Product Page:        - Navigates to the provided URL with a timeout to avoid infinite waits.    2. Set Delivery Location to ZIP 10001:        - Simulates the user clicking on the delivery address area.        - Fills in ZIP code `10001`, applies the change, and closes the popup.        - Ensures that location-dependent product data is displayed correctly.    3. Wait for the Page to Load Fully:        - Uses a hard wait (5 seconds) to give the dynamic content time to render.    4. Parse the Page Content with BeautifulSoup:        - Grabs the pages HTML and parses it using BeautifulSoup for easier selector-based extraction.    5. Extract Product Fields:        - Uses helper functions (`extract_text_or_none`, `extract_list_price_strict`) to grab each field safely.        - Handles missing elements gracefully (returns None).        - Combines feature bullet points into a single string.    6. Log and Return Extracted Data:        - Logs the complete dictionary of extracted fields.        - Returns the dictionary for database insertion or JSON saving.    """    logging.info(f"Scraping URL: {url}")    await page.goto(url, timeout=60000)    await page.wait_for_timeout(5000)    # Change delivery location to 10001 (New York)    try:        await page.click('#nav-global-location-data-modal-action', timeout=10000)        await page.wait_for_timeout(2000)        await page.fill('#GLUXZipUpdateInput', '10001')        await page.click('#GLUXZipUpdate')        await page.wait_for_timeout(4000)        await page.click('button[name="glowDoneButton"]')        logging.info(f"Delivery location changed to 10001 for {url}")    except Exception as e:        logging.warning(f"Failed to change delivery location for {url}: {e}")    await page.wait_for_timeout(5000)    content = await page.content()    soup = BeautifulSoup(content, "html.parser")    # Extract required fields with None if not present    data = {        "product_url": url,        "name": extract_text_or_none(soup, "#productTitle"),        "image_url": soup.select_one("#landingImage")['src'] if soup.select_one("#landingImage") else None,        "brand": extract_text_or_none(soup, "tr.po-brand span.po-break-word"),        "size": extract_text_or_none(soup, "#inline-twister-expanded-dimension-text-size_name"),        "flavor": extract_text_or_none(soup, "tr.po-flavor span.po-break-word"),        "age_range": extract_text_or_none(soup, "tr.po-age_range_description span.po-break-word"),        "description": " | ".join([li.get_text(strip=True) for li in soup.select("#feature-bullets ul li")]) if soup.select("#feature-bullets ul li") else None,        "item_category": extract_text_or_none(soup, "tr.po-item_form span.po-break-word"),        "total_purchased_count": extract_text_or_none(soup, "div.social-proofing-faceout-title span.a-text-bold"),        "selling_price": extract_text_or_none(soup, "span.a-price span.a-offscreen"),        "discount": extract_text_or_none(soup, "span.savingPriceOverride"),        "original_price": extract_list_price_strict(soup),  # STRICT list price extraction        "rating": extract_text_or_none(soup, "span.reviewCountTextLinkedHistogram span.a-size-base"),        "rating_count": extract_text_or_none(soup, "#acrCustomerReviewText")    }    logging.info(f"Scraped data for {url}: {data}")    return data ``` In the world of e-commerce, where product availability and pricing shift constantly, having up-to-date and structured product data is essential. Whether you're a data analyst tracking market trends, a product manager monitoring competitor listings, or part of an e-commerce intelligence team, automated web scraping can unlock valuable insights—especially when done with care and precision. This blog walks through a Playwright-based product scraper for Amazon, designed specifically for the dog food category, though the logic can be extended to other segments. It combines dynamic browser automation with HTML parsing to collect consistent, structured product data while navigating real-world challenges like dynamic content, location-specific availability, and anti-bot mechanisms. A critical part of this process is a function that extracts detailed product information from an Amazon product page. Once the browser reaches the given product URL, it begins by setting the delivery location to ZIP code 10001 (New York). This step ensures consistency—Amazon often tailors availability and pricing based on location, and skipping this can lead to missing or incorrect data. After confirming the ZIP code update, the script waits for the page to fully load to avoid capturing incomplete content. BeautifulSoup is then used to parse the HTML, making it easier to navigate the DOM and extract necessary fields. The function pulls a wide range of data: product title, brand, image URL, size, flavor, age range, key features, product form (like dry or wet food), purchase count, price details, discounts, and customer ratings. To ensure reliability, it uses helper functions to fetch each field, handling missing elements without breaking the flow. For example, it checks if an image or rating exists before trying to read its value. Feature bullets are combined into a clean string, and even pricing is handled carefully to differentiate between original and discounted rates. In the end, this function returns a well-structured dictionary of all relevant product information. It’s clean, readable, and ready for storage in databases or transformation into a JSON file for further analysis. By logging each step—from page visit to ZIP code application to final data capture—it becomes easier to monitor the scraping process, debug issues, and scale up confidently. ### Main Function to Run the Scraper ``` # Main Function to Run the Scraper async def main():    """    Main function that coordinates scraping all unprocessed product URLs.    Workflow:    - Connects to SQLite database.    - Creates tables and adds missing columns if needed.    - Loads URLs that haven't been scraped (scraped=0).    - Opens browser using Playwright with random user agent and cookies.    - Loops over each product URL:        - Extracts product data        - Saves it to database and JSON        - Marks the URL as scraped    - Saves final results to a JSON file.    """    conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    # Create scraped column if not exists    try:        cursor.execute("ALTER TABLE product_urls ADD COLUMN scraped INTEGER DEFAULT 0")        conn.commit()    except:        pass    # Create product_data table if not exists    cursor.execute("""        CREATE TABLE IF NOT EXISTS product_data (            id INTEGER PRIMARY KEY AUTOINCREMENT,            product_url TEXT UNIQUE,            name TEXT,            image_url TEXT,            brand TEXT,            size TEXT,            flavor TEXT,            age_range TEXT,            description TEXT,            item_category TEXT,            total_purchased_count TEXT,            selling_price TEXT,            discount TEXT,            original_price TEXT,            rating TEXT,            rating_count TEXT        )    """)    conn.commit()   # Load URLs that haven't been scraped yet    cursor.execute("SELECT id, url FROM product_urls WHERE scraped=0")    urls = cursor.fetchall()       # Pick a random user agent    user_agent = get_random_user_agent()      # Launch browser and start scraping    async with async_playwright() as p:        browser = await p.chromium.launch(headless=False)        context = await browser.new_context(            user_agent=user_agent,            extra_http_headers=headers_template        )        await context.add_cookies(cookies)        page = await context.new_page()        all_results = []        for row in urls:            id_, url = row            try:                data = await scrape_product(page, url)                all_results.append(data)                # Insert into product_data table                cursor.execute("""                    INSERT OR REPLACE INTO product_data (                        product_url, name, image_url, brand, size, flavor, age_range,                        description, item_category, total_purchased_count,                        selling_price, discount, original_price, rating, rating_count                    ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)                """, (                    data["product_url"], data["name"], data["image_url"], data["brand"], data["size"],                    data["flavor"], data["age_range"], data["description"], data["item_category"],                    data["total_purchased_count"], data["selling_price"], data["discount"],                    data["original_price"], data["rating"], data["rating_count"]                ))                conn.commit()                # Mark as scraped                cursor.execute("UPDATE product_urls SET scraped=1 WHERE id=?", (id_,))                conn.commit()                logging.info(f"Data saved for {url}")            except Exception as e:                logging.error(f"Error scraping {url}: {e}")        # Save all results to output JSON file        with open(OUTPUT_JSON, "w") as f:            json.dump(all_results, f, indent=4)        await browser.close()    conn.close() ``` Managing the end-to-end workflow of a web scraping project requires more than just visiting pages and pulling data—it also involves organizing that process systematically. The main() function acts as the central controller that brings together all the moving parts of the scraping system. Think of it like the manager of a busy kitchen, coordinating ingredients (URLs), tools (browser and headers), and helpers (functions) to produce a well-structured data dish. It starts by connecting to an SQLite database and making sure everything is in place: from creating required tables to ensuring a scraped column exists to track which URLs have already been processed. This prevents duplication and lets the scraper resume smoothly even if it was stopped midway. Next, it selects all unprocessed product URLs from the database—those with a scraped status of 0. Once the data is ready, the function launches a browser session using Playwright. To avoid detection by anti-bot systems, it sets a random user agent and predefined headers, then loads cookies to simulate a real user session. For every product URL, it visits the page, extracts structured product details using the scrape\_product() function, and saves the results both in a product\_data database table and into a local JSON file. After each product is successfully scraped, its status is updated in the database so that it's not processed again. If an error occurs, it is logged without stopping the entire process. Finally, all gathered data is saved into an output JSON file for easy access and analysis. This well-organized loop not only makes the scraping process scalable and efficient but also ensures that results are stored reliably for future use. Whether you're a data analyst preparing for insights or a product manager monitoring listings, this kind of structured workflow ensures you always have clean, up-to-date data in hand. ### Execution Flow ``` # Entry Point if name == "__main__":    asyncio.run(main()) """ This block ensures the script runs only when executed directly (not imported as a module). It uses asyncio to run the `main()` function asynchronously, which starts the scraping process. """ ``` Every Python script needs a clear starting point—just like a train waits for the signal before it moves. In this case, the line if name == "\_\_main\_\_": acts as that signal. It tells Python to begin the execution of the script only if it’s run directly, not when it’s imported elsewhere. This setup ensures that the scraper launches at the right time. Inside this block, asyncio.run(main()) kicks off the entire asynchronous scraping process. It’s like pressing the “Go” button, allowing all the carefully written logic to unfold—from loading pages to saving data. This simple structure keeps everything organized and ensures that automation begins exactly where and when it should. ## Conclusion In today’s data-driven world, staying ahead in e-commerce often comes down to having the right information at the right time. This blog demonstrated how web scraping—when done strategically using tools like Playwright, SQLite, and BeautifulSoup—can transform thousands of Amazon product listings into clean, structured, and insightful data. By automating the collection of real-time product details from the dog food category, we’ve shown how even a complex site like Amazon can be decoded with the right approach. Whether you're a data analyst, researcher, or entrepreneur, this scraping workflow lays the foundation for smarter decision-making, competitive analysis, and scalable data solutions—without the guesswork. ## Libraries and Versions Name: playwright Version: 1.48.0 Name: beautifulsoup4 Version: 4.13.3 AUTHOR I’m Anusha P O, Data Science Intern at Datahut. I specialize in building smart scraping systems that automate large-scale data collection from complex e-commerce sites like Amazon. In this blog, I walk you through how we extracted and structured thousands of product listings from Amazon’s dog food section using Playwright, SQLite, and asynchronous Python workflows—turning vast amounts of raw HTML into clean, analysis-ready datasets. At Datahut, we help businesses unlock the full potential of web data by designing robust, scalable scraping solutions tailored for competitive intelligence, pricing analysis, and product visibility tracking. If you’re exploring data-driven strategies for e-commerce or product research, reach out via the chat widget on the right. Let’s work together to transform your data needs into actionable insights. ### FAQs 1\. Is it legal to scrape Amazon dog-food product data using Python libraries? Scraping Amazon must be done carefully, as Amazon’s Terms of Service restrict automated scraping. To stay compliant, focus on publicly available data, respect robots.txt, and avoid overloading servers. For commercial use, consider Amazon’s official APIs. 2\. Which Python libraries are best for scraping Amazon dog-food product listings? Popular libraries include Requests for sending HTTP requests, BeautifulSoup for parsing HTML, Scrapy for large-scale crawling, and Selenium or Playwright for handling dynamic content like JavaScript-rendered pages. 3\. Can I scrape product details like price, reviews, and ratings for Amazon dog-food items? Yes, you can extract details such as product title, price, ratings, number of reviews, brand, and ingredients. However, prices and stock change frequently, so it’s important to schedule your scraper to run at regular intervals for updated data. 4\. How do I avoid getting blocked while scraping Amazon? Use techniques like rotating user agents, proxy servers, and adding random delays between requests. Also, keep your scraping rate slow to mimic human browsing and reduce the risk of CAPTCHAs or IP bans. 5\. What are the real-world applications of scraping Amazon dog-food data? Scraped data can help with price comparison, market trend analysis, competitor monitoring, customer sentiment analysis through reviews, and inventory optimization for pet product retailers and e-commerce businesses. ### Building Fast and Resilient Web Scrapers for Dynamic Sites URL: https://www.blog.datahut.co/post/how-to-build-smart-fast-resilient-web-scrapers-for-dynamic-websites/ Last updated: 2026-09-07T09:44:09.000Z Scraping [dynamic websites](https://www.blog.datahut.co/post/scrape-a-dynamic-website-using-python/) can be tricky, content often loads via JavaScript after the initial page render, making traditional HTML parsing useless. In this guide, you’ll learn exactly how to build smart, fast, and bot-resistant web scrapers for dynamic websites, with real examples from Datahut’s scraping projects. When I first started [web scraping](https://www.blog.datahut.co/post/web-scraping-at-large-data-extraction-challenges-you-must-know/), I thought it would be simple — send a request, get the HTML, and extract what I need. But then I came across dynamic websites. These sites didn’t give me all the data in one go. Some content, like product reviews or ratings, would only load after scrolling or clicking. That’s when I realized scraping dynamic websites is a whole different game. ## What is a Dynamic Website? Dynamic websites are pages that load some or most of their content only after the initial page is loaded — usually through JavaScript. This means if you just check "View Page Source," the content won’t be there. Instead, JavaScript runs in the browser, fetches data in the background, and updates the page. To show this clearly, I took an example from [Blinkit’s](https://blinkit.com/cn/fresh-vegetables/cid/1487/1489?ref=blog.datahut.co). When JavaScript is enabled, product listings appear and load as you scroll down the page. But when JavaScript is disabled, the page looks almost empty — no product data is visible, and scrolling doesn't load any new items. Here, I’ve added screenshots showing the difference:One with JavaScript enabled (you can see the products) and other with JavaScript disabled (only the header and layout are visible, products are missing) ![Example page showing javascript enabled and disabled difference](https://www.blog.datahut.co/content/images/2026/07/img-57.jpg.webp) This kind of behaviour confirms that we’re dealing with a dynamic, scroll-triggered loading website — and we’ll need to use browser automation (like Playwright or Selenium) with smart scrolling strategies to handle it properly. Over time, after scraping a variety of dynamic websites, I’ve picked up a bunch of techniques and tricks that make the scraping process smoother, faster, and less error-prone. This isn’t a theoretical guide. These are all from real scraping experiences — what worked for me, and why I keep using them. Here's how I write smart and efficient scrapers for modern dynamic sites. ## Use the Right Technology for Dynamic Websites If you’re working with a dynamic website, one of the first and most important decisions is choosing the right tool. Why not use requests or BeautifulSoup directly? Because those tools only work with static HTML. Dynamic websites use JavaScript to load most of their content after the page initially opens — which requests and BeautifulSoup won’t see. ### What can you use instead? - Playwright (Recommended): Can control a headless or visible browser, wait for JavaScript to load, click buttons, scroll, and more. - Selenium: Another browser automation library, great for beginners and widely supported. - Puppeteer: Node.js-based automation library like Playwright. - Splash (for Scrapy users): Lightweight headless browser with a Lua scripting interface. But personally, I stick with Playwright for most of my dynamic scraping needs — it’s fast, reliable, and supports async code well Example Code: Playwright Setup ``` from playwright.async_api import async_playwright import asyncio async def main(): async with async_playwright() as p: browser = await p.chromium.launch(headless=True) # use False during debugging page = await browser.new_page() await page.goto("https://example-dynamic-site.com", wait_until="networkidle") await page.wait_for_selector(".target-element") # wait for specific content html = await page.content() # for parsing via BeautifulSoup await browser.close() asyncio.run(main()) ``` ## Splitting the Task into Two Phases This is one of the first things I learned after scraping a couple of dynamic websites. Trying to do everything — from collecting links to scraping details — in one go is messy and hard to fix if something goes wrong. That’s why I always split my work into two clear scripts: - Phase 1: A script just to collect all product or listing URLs. - Phase 2: A separate script to open each of those URLs and extract full product details. This structure helped me debug things faster and avoid unnecessary repetition. Also, if scraping gets interrupted, I don’t lose everything. ## Using Async/Await for Concurrency When I first started scraping dynamic websites, I opened one page, waited for it to load fully, scraped the content, and then moved to the next. But I noticed this took too long — especially when I had 100s of product pages to scrape. That’s when I learned about async and await. These two helped me open and scrape multiple pages at the same time, without waiting one-by-one. But let me be honest — this doesn't mean you can open 50 pages at once! That will crash your browser or slow down everything. From my experience, it’s best to limit the scraper to open only 3–4 pages at a time to keep things stable and safe. ### How I Run 3–4 Pages at a Time Safely I use asyncio.gather() to scrape multiple pages at once, but in small batches. That keeps memory usage low and prevents crashes. ``` from playwright.async_api import async_playwright import asyncio product_urls = [ "https://example.com/product1", "https://example.com/product2", "https://example.com/product3", "https://example.com/product4", "https://example.com/product5" ] async def scrape_page(playwright, url): browser = await playwright.chromium.launch(headless=True) page = await browser.new_page() await page.goto(url) await page.wait_for_selector("h1.product-title") title = await page.text_content("h1.product-title") print(f"Scraped: {title}") await browser.close() async def main(): async with async_playwright() as playwright: for i in range(0, len(product_urls), 3): # Run 3 at a time batch = product_urls[i:i+3] tasks = [scrape_page(playwright, url) for url in batch] await asyncio.gather(*tasks) asyncio.run(main()) ``` ### What await Really Does This is the part that clicked for me: await pauses that line until the browser finishes that task — like loading a page, finding an element, or clicking a button. It doesn't block the whole program, just that one function. So other parts can continue running in the background. That’s what makes it non-blocking and perfect for automation or scraping multiple pages smoothly.Everywhere I need to wait for something to finish — like: ``` await page.goto(url) await page.wait_for_selector("div.item") await page.text_content("h1.title") await page.screenshot(path="item.png") await browser.close() ``` Without await, the code will try to scrape content before it even appears — and then it fails or scrapes nothing.await is your best friend when scraping dynamic websites. ## Smart Scrolling (Full, Section, or Incremental) Many dynamic websites don’t load all products or content at once. Instead, they show more items only when we scroll down. So just visiting the page won’t be enough — we need to scroll like a real user would. At first, I didn’t know this. I used to go to the page and see only 10 items, but I knew there were more. When I manually scrolled the page in the browser, the rest loaded. That’s when I realized I had to add scrolling logic in my script. Based on how the site is built, I use different scrolling methods. Here’s how I decide what to use: ### 1\. Full Page Scroll (scroll to the bottom all at once) This is the simplest one. Some websites load everything once we scroll to the bottom. I just scroll to the bottom and wait for the content to load. ``` await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await asyncio.sleep(2) # Wait for items to load ``` This worked for sites that dump all content when you reach the bottom — no step-by-step needed. ### 2\. Section Scroll (scroll inside a particular box or div) Some websites use a scrollable box inside the page — not the full page scroll. This confused me at first. I scrolled the whole page and nothing happened. Then I noticed that the products were inside a container with its own scrollbar. In that case, I scroll that specific element. ``` await page.evaluate(''' document.querySelector("div.scrollable-section").scrollTop = document.querySelector("div.scrollable-section").scrollHeight; ''') await asyncio.sleep(2) ``` Always inspect the page and find the class or ID of the scrollable container. ### 3\. Incremental Scroll (scroll step-by-step) This one is super useful when the website loads more items only after small scrolls — like lazy loading. Instead of jumping to the bottom, I scroll down little by little in a loop, giving time for new content to load at each step. ``` for _ in range(10): await page.evaluate("window.scrollBy(0, 1000)") await asyncio.sleep(1) # Wait after each scroll ``` I use this when the content loads slowly or in parts. This one gives the most control. ## Waiting for Selectors Instead of Just Sleeping When I first started automating dynamic websites, I used to add time.sleep(5) thinking that the page would load in that time. But I soon realized — just sleeping doesn’t mean the page is ready. Sometimes 5 seconds is too much, sometimes too little. And the scraper would either crash or miss data. So instead of using a fixed time, I started using smart waits — where the script waits for a specific element to appear before moving forward. That’s how I know the content I need is actually loaded. ``` await page.wait_for_selector("div.review-block") ``` This tells the browser: "Wait here until the review section shows up." This is much better than guessing how long to wait. It only moves forward when that part of the page is really visible. ### Also: Wait Until Network Becomes Idle Sometimes I want to be extra sure the page has loaded — for that i use Playwright’s wait\_until="networkidle" option. It waits until there’s no more network activity for a while. ``` await page.goto(url, wait_until="networkidle") ``` This is like saying: "Go to this URL, but don’t continue until everything in the background is calm and finished." ## Anti-Bot Detection and How I Avoid It When I started scraping dynamic websites, I thought just loading the page and extracting data would be enough. But very soon, I began facing strange issues — sometimes data was missing, sometimes the site would load differently, and sometimes I’d get blocked completely. That’s when I understood that these websites actively try to detect bots. They look for small signs — like repeating patterns, strange browser behavior, or too many visits from the same IP — and if something feels off, they stop you right there. So, through trial and error, I figured out a bunch of methods that helped me avoid detection and make my scraper behave more like a real human user. Here’s what I do: ### User Agent Rotation At first, I didn’t think much about headers or how my scraper “looked” to the site. But after a few runs, I started seeing blocks, empty data, or strange redirects. That’s when I came to know about User-Agent strings. A User-Agent is just a line in the browser request that tells the site what kind of device or browser you're using. For example, it might say "I’m a Chrome browser on Windows" or "I’m Safari on an iPhone." Websites use this to understand who is visiting. If we keep using the same User-Agent for every request, it becomes obvious that it’s a bot doing the job. That’s risky because many sites block such patterns. So I created a simple text file called useragents.txt, where I pasted over 100 different User-Agent strings — taken from actual browsers, mobiles, and operating systems. During scraping, my code picks one randomly for each page visit. This randomness helps my scraper behave like different users — as if people are visiting from different devices. It reduces the chances of getting blocked or flagged. Here’s the short code I use: ``` import random with open("useragents.txt") as f: USER_AGENTS = f.read().splitlines() headers = {"User-Agent": random.choice(USER_AGENTS)} ``` This tiny change actually helped me a lot while scraping dynamic websites. It’s like a small disguise for my scraper — not foolproof, but definitely useful when combined with other smart methods. ### Adding Realistic Headers from DevTools Even after setting the User-Agent, some websites didn’t behave normally. Then I realized — the browser usually sends extra information in the background, called headers. These headers tell the website what kind of browser is being used, where the request is coming from, and more. That’s when I started checking the Network tab in Chrome DevTools or Postman, just to copy the same headers that a real browser sends. After I added them to my Playwright browser context or requests inside the script, it worked much better — the page loaded properly, fewer errors, and less blocking. This made a big difference. The website started responding like it would for a regular user. ``` headers = { "accept": "text/html,application/xhtml+xml", "user-agent": random.choice(USER_AGENTS), "referer": "https://example.com" } context = await browser.new_context(extra_http_headers=headers) ``` With this, Playwright acts more natural, just like a normal browser. ### Random Delay Between Requests In the beginning, I used to send requests one after the other — super fast. At first, it felt cool that my scraper could go through 50 pages in seconds. But soon, I started facing issues: some requests would return empty, or worse, the site would block me completely. Then I realized that humans don’t browse like that. No one clicks every link within milliseconds. So, it became clear — I had to slow my scraper down a bit and make it behave more like a real person. The solution? Add random delays between each page visit or scraping task. Not a fixed sleep like await asyncio.sleep(2) everywhere — that’s still a pattern. Instead, I use a random time gap, like 5 to 10 seconds, for every request. ``` import random, asyncio await asyncio.sleep(random.uniform(5,10)) ``` This makes each request take a slightly different amount of time. It’s not only safer but also more natural. From my experience, this simple trick reduced my chances of getting blocked and gave my scrapers a longer life. So now, I always include random delays as a part of my scraping routine — especially for dynamic websites that load content using JavaScript. ### Headless=False During Development or Detection Some websites behave differently when they know the browser is in headless mode — meaning it runs in the background with no window. At first, I didn’t understand why data was missing or why the page looked broken, even though it worked fine when I visited it manually. Later, I figured out that some websites detect if the browser is headless and either block the content or behave differently on purpose. So, when I face such issues, I just change headless=False in my code to open the browser in visible mode. This way, the browser behaves more like a real user and most of the problems go away. Using headless=False is also super helpful during development. I can actually see what the browser is doing, whether the page is loading properly, and where the scraper is clicking or scrolling. So now, especially while testing or when a site is blocking me in headless mode, I always switch to headless=False. It has helped me fix a lot of unexpected bugs during scraping. ### Geolocation Settings for Area-Specific Content Sometimes when I scrape websites like Blinkit or food delivery sites, I notice that they show different products or shops based on my location. This is because they use geolocation to decide what to display. If the browser doesn’t share a proper location, they may show errors or not load anything useful. So what I do is set a fake location (like Mumbai or Delhi) using Playwright’s built-in geolocation feature. This makes the browser think I'm visiting from that place, and the site loads data just like it would for a real user in that city. ``` context = await browser.new_context( geolocation={"longitude": 77.2090, "latitude": 28.6139}, permissions=["geolocation"] ) ``` In this example, I set the geolocation to Delhi’s coordinates. If I wanted to set it to Mumbai, I’d just change the latitude and longitude. This trick has helped me many times when a website refused to load or showed a "service not available in your area" message. Setting the location properly gives better control and makes the scraper behave more like a real local user. ### Proxies to Avoid IP Bans or Region Restrictions Even after doing everything else right, some sites block my IP address if I scrape too many pages or sometimes the first page itself. That’s where [proxies](https://www.blog.datahut.co/post/a-guide-to-using-proxies-for-web-scraping/) help. A proxy acts like a middleman. Instead of your real IP reaching the website, the proxy’s IP is used — so the website thinks it’s a different user. I mostly use rotating proxies, which means my scraper switches IP addresses automatically after each request or after a few. This gives each visit a new identity, reducing the chance of getting blocked. Some websites also show different content based on your region. For example, a product or service might only be available in Bangalore but not in Delhi. In that case, I use a region-based proxy, which lets me pretend I’m browsing from that specific location. Using proxies has saved me many times — especially when scraping large websites where too many visits from the same IP would result in blocks. ``` browser = await playwright.chromium.launch( proxy={"server": "http://your-proxy-ip:port"}, headless=False ) ``` With proxies, I can scrape more safely without worrying about IP bans. ### Shipping Pincode Set for Local Stores For websites like Blinkit, BigBasket, or any online grocery or delivery service, I noticed that the products, prices, and even whether something is available — all depend on the shipping pincode. If the pincode is not set properly, the website either shows limited content or nothing at all. So before scraping any data, I first check if there’s a popup asking for the pincode, or if there’s a small input box on the page where I can enter it. Using Playwright, I automate this step just like a normal user: type the pincode and press Enter or click the submit button. Once the location is set, the entire page refreshes and shows data that’s specific to that area — exactly what a real user from that region would see. Only after this step do I continue scraping the product or price details. Here’s a small example of how I do it in Playwright: ``` await page.goto("https://www.blinkit.com") await page.click("input[placeholder='Enter your PIN code']") # or use actual selector await page.fill("input[placeholder='Enter your PIN code']", "600001") # Example pincode await page.keyboard.press("Enter") await asyncio.sleep(3) # wait for page to reload ``` I always make sure the location is fully set before scraping — otherwise, I might collect wrong or incomplete data. This trick is especially useful when a client or project needs data from a specific city or area. ### Stealth Plugins to Hide Automation Even after all these steps, some sites still figured out that I was using automation. That’s because Playwright (and tools like it) leave small signs — like the navigator.webdriver flag — which websites can detect. To fix this, I started using stealth plugins like playwright-stealth. These remove those signs and tweak the browser to behave like a real human-controlled one. After using stealth settings, I saw a huge improvement. No more blank pages or errors due to automation detection. ## Some Additional Practices for Structuring and Writing Smart Scrapers - Use BeautifulSoup for Parsing: After the page loads using Playwright, I grab the HTML with page.content() and parse it using BeautifulSoup because it's simple and flexible. - Break the Code into Functions:When the script gets long, I split it into small functions like get\_price(), get\_title(), etc., which makes things cleaner and reusable. - Save One Product at a Time: Instead of saving all data at the end, I insert product info one by one as soon as it's scraped to avoid memory loss or crashes. - Use INSERT OR IGNORE for Duplicates:To avoid inserting the same product twice, I use INSERT OR IGNORE in my SQL queries — it keeps the database clean. - Add Full Error Handling:I wrap the whole scraping logic in try-except blocks and log or save the failed URLs separately so I can retry them later. - Enable Logging for Everything:I log each step — scraping success, skips, failures — so I can easily trace bugs without having to rerun everything. - Use a Scraped Status Column:In the database, I keep a scraped column and mark URLs as done after scraping, so I can restart from where I left off. - Reuse Browser Sessions and Tabs:Instead of opening a new browser each time, I open the browser once and reuse tabs to save both time and system resources. - Close Pages and Free Resources:After scraping each product, I close the page and clear resources — this avoids memory leaks and keeps things running smoothly. - Avoid Repeating Code — Reuse Everything:Whether it's selectors, headers, or parsing logic, I reuse everything by keeping them in helper functions or separate modules. ## Common Mistakes to Avoid When Scraping Dynamic Websites Even with all the techniques and best practices above, there are a few common pitfalls that many beginners (and sometimes even experienced scrapers) fall into. These mistakes can slow down your scraper, cause errors, or even get you blocked. Here’s what to watch out for: ### 1\. Overloading the Browser with Too Many Pages at Once One of the first mistakes I made when I started scraping dynamic sites was opening too many pages at the same time. At first, it seemed efficient — more pages, faster scraping, right? Not really. The browser slows down, memory usage spikes, and sometimes it crashes completely. Always stick to small batches (3–4 pages at a time) and use async/await to manage concurrency safely. ### 2\. Using Fixed sleep() Instead of Smart Waits Relying on time.sleep() or await asyncio.sleep() for fixed durations is tempting, but it’s unreliable. Pages load at different speeds depending on network conditions or server response times. This can cause missed elements or partial scraping. Instead, always wait for specific elements with await page.wait\_for\_selector() or use wait\_until="networkidle" to make sure the page is fully ready before extracting data. ### 3\. Ignoring Geolocation or Pincode Requirements Dynamic websites, especially e-commerce or delivery platforms, often show content based on your location. Skipping the geolocation or pincode step can result in missing data or incorrect product listings. Always check if the site requires a location, and automate that step to make sure you scrape what a real user in that area would see. ### 4\. Forgetting to Rotate User-Agents and Headers Using the same User-Agent for every request or not mimicking real browser headers makes your scraper easy to detect. Many sites block repeated requests or serve different content. Randomizing User-Agents and copying realistic headers from DevTools can save your scraper from unnecessary blocks. ### 5\. Not Handling Scrollable Sections or Lazy Loading Properly A lot of dynamic websites use infinite scroll or scrollable containers to load content incrementally. Simply opening the page won’t give you all the data. Missing this step is a common mistake. Inspect the page carefully, choose the right scroll method (full, section, or incremental), and give the page time to load new content between scrolls. ### 6\. Not Saving Data Incrementally Another mistake I often see is waiting to save all data at the end. If your scraper crashes midway, you lose everything. Save product data one by one, use INSERT OR IGNORE for duplicates, and keep track of scraped URLs with a status column. This makes your scraper more resilient and easy to resume. ### 7\. Overlooking Anti-Bot Measures Many scrapers ignore proxies, random delays, or stealth plugins. The result? Blocks, empty responses, or CAPTCHA challenges. Incorporate these preventive measures from the start — random delays, headless=False during testing, proxies, and stealth plugins help your scraper act like a real human user and keep it running longer. By keeping these common mistakes in mind, you can save yourself hours of frustration and make your dynamic scrapers more efficient, reliable, and resilient. ## Conclusion Scraping dynamic websites isn’t just about writing code — it’s about building scrapers that think and act like real users. From handling JavaScript rendering and infinite scrolling to rotating user-agents, setting geolocation, and adding smart delays, every detail makes the difference between a scraper that breaks quickly and one that runs smoothly for months. The key is to stay smart, fast, and resilient: - Smart by using the right tools (like Playwright) and strategies (async/await, modular design). - Fast by optimizing concurrency and avoiding unnecessary waits. - Resilient by preparing for anti-bot measures, website changes, and unexpected failures. If you follow these practices, you’ll be able to build scrapers that not only extract the data you need but also adapt to the constantly evolving web. At [Datahut](https://datahut.co/?ref=blog.datahut.co), we’ve applied these exact methods across hundreds of projects — proving that with the right approach, even the most complex dynamic websites can be scraped reliably. ### FAQs Q1\. What makes scraping dynamic websites more challenging than static ones? A: Unlike static websites where content is directly available in the HTML, dynamic websites load content using JavaScript, AJAX, or APIs. This requires techniques like headless browsing, JavaScript rendering, or API calls to reliably extract data. Q2\. How can I make my web scraper faster? A: You can optimize speed by using asynchronous requests, efficient libraries (like Scrapy or Playwright), rotating proxies, caching repeated requests, and reducing unnecessary page loads. Q3\. What strategies help make a scraper resilient against website changes? A: To handle frequent website updates, use robust CSS/XPath selectors, leverage APIs when available, build modular scraper code, monitor for changes, and implement fallback mechanisms. Q4\. How can I avoid IP blocking while scraping dynamic sites? A: Use techniques like IP rotation, proxy pools, user-agent rotation, request throttling, and respecting robots.txt to minimize the risk of being blocked. Q5\. Which tools or frameworks are best for scraping dynamic websites? A: Popular choices include Playwright, Selenium, Puppeteer for handling JavaScript-heavy pages, and Scrapy, Beautiful Soup, Requests for efficiency. The best tool depends on the complexity of the website and the scraping requirements. ### Using Competitor Data to Improve Category Management URL: https://www.blog.datahut.co/post/how-competitor-data-transforms-category-management-pro-tips-best-practices/ Last updated: 2026-07-23T07:48:31.000Z ## Introduction: Why Competitor Data is Critical for Category Managers Ever wondered how analyzing your competitors’ moves could supercharge your own category management strategy? If you’re a Category Manager, you already know how difficult it is to balance pricing, promotions, inventory, and assortment planning. But one secret weapon can make your job easier and more effective: competitor data. Imagine running your category without knowing what your competitors are doing: - Are they selling similar products at lower prices? - Are they adding trending items that you’ve overlooked? - Are they bundling products or offering aggressive promotions? Without competitor intelligence, it’s harder to make confident, data-driven decisions. Competitor data fills in these gaps, providing clarity, reducing mistakes, and revealing opportunities to grow market share. According to McKinsey, companies that leverage competitive intelligence in category management see 10–15% improvements in margin growth. This blog will explore how competitor data supports smarter decisions and share 10 best practices to help you use it effectively. ## How Competitor Data Supports Smarter Category Management Category management is more than analyzing your own KPIs — it’s about understanding the bigger market picture. Competitor data empowers managers across five key areas: ### 1\. Smarter Pricing Decisions Pricing strategy is the foundation of retail success. If competitors consistently undercut you, customers will switch without hesitation. Example: - Competitor sells a home appliance for ₹500 - You price it at ₹650 - Without competitor pricing data, you wouldn’t realize you’re overpriced On the flip side, underpricing can shrink margins unnecessarily. Competitor pricing insights help you: - Stay competitive while protecting profits - Spot opportunities to apply dynamic pricing - Identify loss-leader tactics used by rivals Pro Tip: Use competitor data combined with real-time updates from web scraping to keep pricing agile. ### 2\. Better Assortment Planning Assortment planning ensures you stock what customers actually want. Competitor data highlights catalog gaps and trending categories. - Fashion example: Competitors launch eco-friendly collections → missing this trend risks losing sustainability-driven shoppers. - Electronics example: Rivals add smart home hubs or earbuds → sticking to old SKUs makes your catalog outdated. Competitor assortment intelligence enables: - Smarter data extraction to identify missing products - Better forecasting for seasonal shifts - Inventory planning that improves user experience ### 3\. Stronger Promotions and Offers Promotions are often the deciding factor during peak shopping. But weak promotions get overshadowed. Example: You run a 15% discount, but competitors launch a flash “Buy One Get One” deal — their promotion wins. Competitor promotion tracking helps you: - Monitor real-time updates on offers - React faster with bundles or shipping perks - Differentiate through loyalty rewards instead of deeper discounts Stat: Retail Dive found that 64% of shoppers actively compare promotions across stores before purchasing. ### 4\. More Confident Supplier Negotiations Suppliers respect data. With competitor intelligence, you can: - Show competitor product prices during talks - Benchmark offers and promotions - Request equal or better terms This creates leverage, especially in industries where suppliers dominate margins. ### 5\. Benchmarking Your Performance Competitor data also provides a benchmarking framework for performance. Ask: - Are our product prices aligned with competitors? - Do we offer more variety? - How do customer reviews compare? Benchmarking ensures you don’t just celebrate internal growth while losing external market share. ## Best Practices for Using Competitor Data in Category Management Collecting competitor data is only step one. Real impact comes from how you use it. Here are 10 best practices: ### 1\. Keep It Ongoing, Not One-Time Markets change daily. Prices shift with HTTP requests and stock levels fluctuate constantly. Promotions appear and vanish quickly. Best Practice: Use web scrapers to track competitors daily or weekly — not just quarterly. ### 2\. Don’t Focus Only on Price To understand true competitor strategy, also monitor: - Stock availability (capture demand when rivals go out of stock) - Customer reviews (find weaknesses to exploit) - Delivery speed (shipping perks influence conversions) - New launches (anticipate demand shifts) Competitor insights give a holistic view of the customer experience. ### 3\. Use Web Scraping for Efficiency Manual monitoring doesn’t scale. Web scraping automates data extraction across competitor websites. You can scrape: - Product names and prices - Reviews & ratings - Promotions - Stock availability Pro Tip: Use Selenium WebDriver, headless browsers, and CSS selectors to scrape dynamic websites with infinite scrolling or complex page structures. Pair with proxy services, IP rotation, and CAPTCHA handling to avoid blocks. ### 4\. Turn Data into Actionable Steps Competitor data must drive action. Examples: - Price drop? Offer free shipping or bundle discounts. - New eco-friendly SKUs? Launch sustainability campaigns. - Heavy discounts? Focus on service and loyalty instead of a price war. ### 5\. Align Insights with Your Strategy Don’t copy blindly. Align competitor analysis with your brand’s positioning. - Premium brands → Compete on quality and trust - Value brands → Focus on affordability and promotions Competitor data should guide, not control your decisions. ### 6\. Benchmark Your Own Performance Compare yourself with competitors regularly. - Are your HTTP response times affecting performance? - Is your assortment broader or narrower? - Do your customer reviews highlight strengths? Continuous benchmarking ensures competitiveness. ### 7\. Group Competitors by Relevance Not all rivals matter equally. - Direct competitors: Same products, same customers - Indirect competitors: Different products, same need Example: Ceiling fans vs. portable air coolers. ### 8\. Understand the “Why” Behind Moves A sudden price cut could be inventory clearance, not strategy. By analyzing motives, you avoid random delays between responses and unnecessary reactions. ### 9\. Help Teams Interpret Data Competitor insights are useless if teams don’t understand them. - Use dashboards linked to Google Sheets modules - Automate alerts with APIs - Train staff on competitor benchmarking This improves project inception to execution speed. ### 10\. Stay Ethical and Legal Competitor data collection must be ethical. ✅ Okay: Scraping public sites, analyzing ads, monitoring product pages ❌ Avoid: Fake accounts, fake orders, violating site terms of service Ethical data practices build credibility and long-term success. ## Wrapping It Up: Smarter Category Management with Competitor Insights Today’s competitive retail environment demands more than instinct. Competitor data empowers Category Managers to: - Optimize pricing strategy - Strengthen promotional strategy - Improve assortment planning - Negotiate better with suppliers - Benchmark performance effectively Start small, track a few competitors, and expand gradually using [web scraping tools](https://www.datahut.co/?ref=blog.datahut.co) and data mining techniques. The goal isn’t to copy — it’s to understand and respond in ways that strengthen your user experience and keep your brand profitable. ## FAQs: Competitor Data in Category Management Q1: What is competitor data in category management? It’s information about competitor pricing, promotions, assortment, and reviews collected via web scraping, APIs, and data mining to guide strategic decisions. Q2: How often should you track competitor data? Daily or weekly for fast-moving categories (like electronics, fashion), monthly for slower ones (like furniture). Q3: What tools help collect competitor data? Web scrapers, Selenium WebDriver, APIs, and SaaS platforms. For complex dynamic websites, use headless browsers, proxy services, and IP rotation. Q4: How does competitor data help with supplier negotiations? It strengthens your bargaining position by showing suppliers competitor prices, promotions, and real-time updates from the market. Q5: Is collecting competitor data legal? Yes, if done ethically — collecting public data, respecting HTTP response codes, and avoiding deceptive practices. ### Amazon US Vlogging Gadgets Market Analysis Using Data URL: https://www.blog.datahut.co/post/inside-amazon-us-what-data-analysis-reveals-about-vlogging-gadgets/ Last updated: 2026-07-23T07:48:31.000Z Vlogging has rapidly grown from a niche hobby to a mainstream form of storytelling. Whether it’s travel diaries, product reviews, or daily life updates, millions of creators rely on gadgets to capture and share their experiences online. But what gadgets are vloggers actually buying? To answer this, we analyzed 1,177 vlogging products listed on Amazon US, covering details such as price, discount, brand, rating, and product type. The findings reveal fascinating trends about what makes vlogging gear popular, affordable, and well-rated among creators. Let’s dive into the highlights. ![Amazon US Full data](https://www.blog.datahut.co/content/images/2026/07/img-333.png.webp) ## Pricing: Accessibility Over Luxury When it comes to cost, vlogging gear is surprisingly budget-friendly. - 88% of products fall in the $0–500 range, showing that the market is largely geared toward affordability. - The mid-tier ($501–2000) accounts for just 7.2% combined, while high-end tools above $3000 make up barely 1%. This distribution highlights a democratized market—creators don’t need to break the bank to get started. Premium equipment exists but remains niche. ## Discounts: Rare but Strategic Everyone loves a good sale, but how often do vlogging gadgets actually get discounted? The data suggests: not much. - 61% of products have discounts below 10%, meaning many items are sold close to full price. - Only 2% of gadgets feature discounts above 50%. - The steepest discounts (80–90%) are almost nonexistent. For shoppers, this means flashy banners might exaggerate value. Big markdowns are rare, so choosing gear often comes down to balancing features and price rather than relying on discounts. ## Ratings: A Strongly Positive Market Vlogging gadgets on Amazon enjoy generally high customer satisfaction. - 81% of products are rated between 4–5 stars. - Only 1% dip below 3 stars. - Around 10% sit in the 0–1 star range, which often includes either poorly received items or new listings without reviews. This tells us most products available are either genuinely good or filtered by market demand. For creators, the odds of buying well-rated equipment are high. ## Leading Brands: Familiar Giants Meet Rising Players Some brands dominate the vlogging market with a wide range of products: - Fujifilm leads with 43 listings, followed by ULANZI (39) and Canon (34). - Sony and Movo tie at 24 each. - Budget-friendly names like NEEWER, UBeesize, and ORDRO hold strong ground with 21–22 listings. - DJI and SMALLRIG add diversity with stabilizers, drones, and rigs. The mix shows how both established camera giants and newer, affordable brands coexist, shaping a market that serves professionals and beginners alike. ## Price vs. Ratings: Value Doesn’t Always Mean Expensive Comparing average prices with ratings uncovered a key insight: affordability often wins on customer satisfaction. - Fujifilm tops pricing at $2226, but lacks sufficient rating data in the dataset. - Canon ($991, 3.72 rating) and Sony ($1570, 2.03 rating) show that high prices don’t guarantee better feedback. - Ulanzi ($28, 4.44 rating) and UBeesize ($31, 4.57 rating) shine as affordable yet highly rated choices. - DJI strikes a balance—$455 with a 4.48 rating. In short, budget brands like UBeesize and Ulanzi are not only affordable but also loved by customers, proving that creators don’t need to overspend for quality. ## Customer Favorites: High-Rated Niche Brands Interestingly, some lesser-known brands outperform household names in customer satisfaction. - Rythflo and Monitech lead with 4.68 ratings, followed by Fotopro (4.67). - FIFINE, UBeesize, and SMALLRIG all hover around 4.57–4.6, showing that niche players can win big with reliability and affordability. For creators, this means exploring beyond big names can often reveal better deals with equal—or higher—quality. ## Discounts by Brand: Who’s Most Generous? Some brands consistently offer deeper discounts than others: - PQRQP leads with 36.5% average discounts, also ranking high in customer satisfaction. - FJFJOPK and CAMBOFOTO follow at 28%+. - Movo, SelfieShow, and Saneen offer steady 23–25% discounts. The overlap of high discounts and high ratings (e.g., PQRQP) shows that shoppers can find rare “best of both worlds” deals. ## Product Types: The Essentials of Vlogging The dataset grouped gadgets into three major categories: - Tripods (39.3%) lead slightly, with countless variations (flexible, tabletop, stabilizers). - Cameras (37.3%) follow closely, from action cams to mirrorless models. - Microphones (23.4%) take a smaller slice but remain critical for quality content. This spread reveals a balanced ecosystem: visual stability and quality dominate, but sound equipment holds a specialized niche. ## Which Gear Costs the Most? A Surprising Twist Contrary to what many would expect, tripods—not cameras—had the highest average cost. - Tripods average $737, likely due to high-end stabilizers and pro rigs. - Cameras average $478, spanning budget-friendly to premium models. - Microphones average $287, making them the cheapest entry point for creators. This shows how “accessories” like tripods can carry premium value when designed for professional use. ## Category Leaders: Brand Specializations Looking at each category reveals distinct brand strengths: - Cameras: Canon and ORDRO lead alongside a significant number of unbranded products. - Microphones: Movo dominates with 20 listings, while BOYA also has a strong presence. - Tripods: Fujifilm (32) and Ulanzi (31) show their stronghold in stabilizing gear. This breakdown highlights how brands often specialize, building reputations in certain niches. ## Final Thoughts The analysis paints a clear picture of a diverse, competitive, and accessible vlogging gear market on Amazon US. - Budget-friendly gadgets dominate, but many still earn top ratings. - Discounts are modest overall, but certain brands consistently offer great value. - Tripods surprisingly outprice cameras, while niche brands often outperform giants in customer love. For creators—whether beginners or professionals—the good news is clear: today’s market offers reliable, affordable, and varied options to build the perfect vlogging setup. 👉 Want the full analysis with charts and data? [Download the complete documentation here.](https://tally.so/r/wAgArB?ref=blog.datahut.co) ### FAQ SECTION 1\. What is the most common price range for vlogging gadgets on Amazon? Most vlogging gadgets fall in the $0–500 range, making them highly accessible for beginners and hobbyists. 2\. Are vlogging gadgets on Amazon usually discounted? Not really. Over 60% of gadgets have discounts below 10%, and only about 2% get discounts above 50%, so deep deals are rare. 3\. Which brands dominate the vlogging gadget market? Top brands include Fujifilm, Ulanzi, Canon, Sony, and Movo, alongside rising budget-friendly names like UBeesize and ORDRO. 4\. Do expensive vlogging gadgets have better ratings? Not always. Brands like Ulanzi and UBeesize offer affordable gear under $50 but still earn 4.4+ star ratings, sometimes higher than premium brands. 5\. What type of vlogging gear is most popular? Tripods and cameras dominate Amazon listings, making up over 75% of all products, while microphones account for about 23% ### Scraping Blinkit Grocery Data for Market Insights URL: https://www.blog.datahut.co/post/how-to-scrape-blinkit-s-fruits-and-vegetables-data/ Last updated: 2026-07-23T07:48:31.000Z Have you ever wondered how businesses or analysts find out what groceries cost on websites like Blinkit? The trick is something called web scraping , a smart and automated way to gather information from websites. Think of web scraping like a helpful robot assistant. It goes through web pages, picks out the bits we care about like product names, prices, or weights and saves them neatly for us to use. It’s quick, reliable, and way faster than doing it by hand. ### Why Blinkit? Blinkit, which you might remember as Grofers, is one of India’s most popular apps for quick grocery deliveries. It has a wide selection of fresh fruits and vegetables. If we can collect this data regularly, we can learn a lot — like how prices vary, which items are in stock, or how things change with the seasons. That kind of information is super helpful for both businesses and researchers. ### How We Scrape Blinkit’s Fruits and Vegetables Data We break the task down into two easy steps: 1. First, we gather all the links to individual products listed under the fruits and vegetables section. 2. Then, we visit each of those links one by one and collect details like the product name, price, weight, and a short description. This step-by-step approach helps us stay organized and makes it easier to spot and fix any problems if something goes wrong. ## Links Collection In today’s world of e-commerce, data plays a huge role. Whether it's tracking prices or studying customer trends, good data is at the heart of smart decision-making. In our case, we’re in the early stage of the web scraping process — and this step focuses on collecting product links from Blinkit's fruits and vegetables section. These product links are important because they act like doorways. Once we have them, we can step through each one to gather more useful details like pricing, weight, and descriptions. So, collecting the links is our foundation — the first layer of data we need before diving deeper. Here’s how the process works behind the scenes:We point our scraper to Blinkit's website, scroll down the page so that all the fruits and veggies are loaded, then collect the links to each product. After that, we save these links in a database. This way, we can come back later and pull more detailed information from each link. Now that we understand the purpose, let’s take a closer look at how this actually works step by step. ### Setting Up the Environment ``` import sqlite3 import logging from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup import time from datetime import datetime # Configure logging logging.basicConfig(filename="scraper.log", level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") BASE_URL = "https://blinkit.com" ``` Let’s now look at the tools our scraper uses to do its job. Just like how a chef needs the right ingredients before cooking, our web scraping script needs a few key packages to run smoothly. We start by importing everything we need: - SQLite3 helps us connect to a small, lightweight database where we’ll store the product data we collect. Think of it like a notebook for saving all our results. - Logging keeps track of what our scraper is doing. If something goes wrong, these logs will help us figure out what happened. - Playwright is what lets us control a web browser automatically — like telling it where to go, what to click, or when to scroll. - BeautifulSoup is the tool that reads the webpage and pulls out the information we’re interested in. - We also bring in time and datetime to manage waiting periods and to keep track of when our scraper runs. Before we dive into scraping, we also set up our log file, called scraper.log. This file saves messages like when the scraper starts, what it's doing, or if any errors show up. Each message is time-stamped and labeled so we can understand the flow of the process easily. Lastly, we define a small but useful variable called BASE\_URL, which is just the Blinkit homepage link — [https://blinkit.com](https://blinkit.com/?ref=blog.datahut.co). We’ll use this later to build full product links as we collect them. ### Fetching Page Content ``` def fetch_page_content(url): """ Fetches the full HTML content of a webpage using Playwright with incremental scrolling. This function launches a Chromium browser, navigates to the specified URL, and simulates scrolling through the entire page to ensure dynamically loaded content is captured. It uses a bottom-to-top scrolling technique to trigger lazy loading. Parameters: url (str): The URL of the webpage to fetch. Returns: str or None: The complete HTML content of the page if successful, None otherwise. Raises: Exception: Any exceptions that occur during the browsing session are caught, logged, and None is returned. Notes: - Uses non-headless browser mode (visible) which may be changed to headless=True for production. - Includes a timeout of 60 seconds (60000ms) for page loading. - Logs success or failure information to the configured logger. """ try: with sync_playwright() as p: browser = p.chromium.launch(headless=False) page = browser.new_page() page.goto(url, timeout=60000) scroll_position = 0 # Start from the top while True: # Scroll down in increments scroll_position += 600 page.evaluate(f"window.scrollTo(0, {scroll_position})") time.sleep(2) # Wait for new content to load # Get new page height after scrolling new_height = page.evaluate("document.body.scrollHeight") # Stop if we can't scroll further if scroll_position >= new_height: break content = page.content() browser.close() logging.info(f"Successfully fetched content from {url}") return content except Exception as e: logging.error(f"Error fetching page content from {url}: {e}") return None ``` One of the key parts of our scraper is the fetch\_page\_content function. This function is in charge of opening a web page, scrolling through it, and collecting the full content — kind of like someone visiting a webpage and slowly scrolling down to see everything. Here’s how it works, step by step: When we call this function, it opens the link in a visible Chrome browser using Playwright. This visible mode helps us make sure everything loads properly, especially on pages that load content as you scroll. Starting from the top of the page, it scrolls down 600 pixels at a time. After each scroll, it pauses for 2 seconds. This small wait gives the page time to load more products — since many websites (like Blinkit) load items gradually as you move down. As it keeps scrolling, it checks whether it’s reached the bottom of the page. This is important because we want to make sure all products — even the ones that show up only at the end — are included. Once the full page has loaded, the function saves the entire HTML content, closes the browser, and sends the content back so we can extract the data we need. If something doesn’t work — maybe the page didn’t load or the internet dropped — the function catches the error, logs it in our scraper.log file, and safely returns None so the script doesn’t crash. ### Parsing Product Links ``` def parse_links(html_content): """ Parses HTML content and extracts all product links using BeautifulSoup. This function takes HTML content, particularly from Blinkit product listing pages, and extracts links to individual product pages by targeting specific CSS selectors that identify product elements. Parameters: html_content (str): The HTML content to parse. Returns: list: A list of complete product URLs (with BASE_URL prefixed). Raises: Exception: Any exceptions during parsing are caught, logged, and an empty list is returned. Notes: - Uses BeautifulSoup's CSS selector to target elements with 'plp-product' data-test-id. - Prefixes all relative links with BASE_URL to create absolute URLs. - Logs the number of extracted links for monitoring. """ try: soup = BeautifulSoup(html_content, 'html.parser') links = [BASE_URL + a['href'] for a in soup.select('.ProductsContainer__ProductListContainer-sc-1k8vkvc-0 > a[data-test-id="plp-product"]')] logging.info(f"Extracted {len(links)} links.") return links except Exception as e: logging.error(f"Error parsing links: {e}") return [] ``` Once we have the full content of the page, the next step is to pull out the product links — and that’s exactly what the parse\_links function does. This function uses BeautifulSoup to read through the HTML content, just like skimming through a web page’s behind-the-scenes code. It looks for a specific section of the page that holds all the product listings. This section is marked with a special class name: 'ProductsContainer\_\_ProductListContainer-sc-1k8vkvc-0'. Inside this section, it searches for all the anchor () tags that have an attribute called data-test-id="plp-product". These tags hold the actual links to each product. Each link it finds is just a part of the full address. So, to make it a complete and usable web link, the function adds the BASE\_URL (which we defined earlier) at the beginning. This makes sure every link points to the right page on the Blinkit site. Even if the layout of the site changes a bit — like headings move around — as long as the main product container stays the same, this function will still find the product links correctly. Once all the links are collected, the function logs how many were found and returns them as a list. And just like earlier steps, if something goes wrong, it writes an error to the log file and returns an empty list instead of crashing. ### Saving Links to Database ``` def save_links_to_db(links, category, db_name="scraped_links.db"): """ Saves extracted product links to an SQLite database with metadata. This function connects to an SQLite database (creates it if it doesn't exist), creates a table for storing links if needed, and inserts the links along with their category, current date, and a flag indicating they haven't been scraped yet. Parameters: links (list): List of URLs to save to the database. category (str): Category identifier for the links (e.g., "vegetables", "fruits"). db_name (str, optional): Name of the SQLite database file. Defaults to "scraped_links.db". Returns: None Raises: Exception: Any database-related exceptions are caught and logged. Notes: - Uses SQLite's UNIQUE constraint to prevent duplicate links. - Sets a default 'scraped' value of 0 to indicate the link hasn't been processed yet. - Logs warnings for duplicate links instead of failing. - The table schema includes: * id: Autoincrementing primary key * url: The product URL (must be unique) * category: Product category identifier * scraped_date: Date when the link was added to the database * scraped: Flag indicating whether detailed information has been scraped (0=no, 1=yes) """ try: conn = sqlite3.connect(db_name) cursor = conn.cursor() # Create table if not exists cursor.execute(""" CREATE TABLE IF NOT EXISTS links ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, category TEXT, scraped_date TEXT, scraped INTEGER DEFAULT 0 ) """) current_date = datetime.now().strftime("%Y-%m-%d") for link in links: try: cursor.execute("INSERT INTO links (url, category, scraped_date, scraped) VALUES (?, ?, ?, ?)", (link, category, current_date, 0)) except sqlite3.IntegrityError: logging.warning(f"Duplicate link ignored: {link}") conn.commit() conn.close() logging.info("Links saved successfully to the database.") except Exception as e: logging.error(f"Error saving links to database: {e}") ``` After collecting the product links, the next step is to save them somewhere safe — and that’s what the save\_links\_to\_db function takes care of. This function connects to a SQLite database, which is like a small local storage system. Once connected, it checks if there’s already a table named links. If the table doesn’t exist, it creates one. The links table is set up with a few useful columns: - An ID that’s created automatically for each entry - The product URL, which is marked as unique so the same link isn’t saved more than once - The product category, like "fruits and vegetables" - The date the link was saved - A scraped flag, which starts at 0 to show that the product hasn’t been scraped yet The function then goes through each link in the list. For each one, it tries to insert the product URL, category, current date, and the scraped flag (set to 0 for now). If the link already exists in the database (a duplicate), the function simply skips it and logs a warning — instead of stopping the whole process. This way, it avoids errors and keeps things running smoothly. Finally, it saves all changes, closes the connection to the database, and logs a message to confirm that everything was completed successfully. ### Orchestrating the Scraping Process ``` def main(): """ Main function that orchestrates the web scraping process. This function defines a dictionary of URLs to scrape, each with its associated category, and then processes each URL by: 1. Fetching the full page content 2. Parsing the content to extract product links 3. Saving those links to the database Returns: None Notes: - Currently configured to scrape vegetable and fruit categories from Blinkit. - Can be extended by adding more URL-category pairs to the urls dictionary. - Prints a success message when all operations are complete. """ urls = { "https://blinkit.com/cn/fresh-vegetables/cid/1487/1489": "vegetables", "https://blinkit.com/cn/fresh-fruits/cid/1487/1503": "fruits" } for url, category in urls.items(): html_content = fetch_page_content(url) if html_content: links = parse_links(html_content) if links: save_links_to_db(links, category) print("Links saved successfully.") if __name__ == "__main__": main() ``` The main function is like the manager of the whole scraping process. It decides which Blinkit pages we want to scrape — for example, the fruits section or the vegetables section — and keeps track of the category each URL belongs to. For every category and URL, it follows a clear three-step process: 1. It loads the full page 2. It extracts the product links 3. It saves those links into the database This setup is clean and flexible. So, if you ever want to add more categories later — like dairy or snacks — you can easily plug them in without changing much of the code. Now, when you run the script directly (without importing it into another program), the main function runs automatically. After everything is done, it prints a message to let you know that scraping was successful. ## Data Collection Now that we’ve collected all the product links from Blinkit’s fruits and vegetables section, it’s time to take things a step further. This part of the scraper focuses on gathering detailed information about each product — not just the links anymore. Think of this as the second phase of the project. Earlier, we built a list of doors (product links), and now we’re opening each door to see what’s inside. For every product link, the script visits the page and pulls out useful details like the product’s name, price, weight, and more. All of this information is then stored in our database, ready for analysis later. In this phase, we’re expanding our scraper’s reach — from just collecting URLs to capturing actual product data. Let’s walk through how this part works and look at the key steps that help keep the scraping process both reliable and efficient. ### Imports and Configuration ``` import sqlite3 import logging from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup import time import json import random import re from datetime import datetime # Configure logging logging.basicConfig(filename="blinkit_scraper.log", level=logging.INFO, format="%(asctime)s - %(levelname)s - %(message)s") DB_NAME = "scraped_links.db" ``` At the start of this script, we bring in all the important tools we’ll need to make everything work smoothly as we done before. We begin by importing SQLite3, which lets us connect to and interact with our database — where we’ll store all the product details we scrape. Next, we use Playwright to open a browser automatically and visit each product page. This makes the process hands-free and much faster than doing it manually. To read the contents of each webpage and pull out the information we want, we use BeautifulSoup. It helps us navigate through the HTML structure and find exactly what we’re looking for. We also include a few helpful modules like json, random, and re (short for regular expressions). These help us format the data properly, generate random choices when needed, and search for specific patterns in the data. Finally, we set up logging so everything the scraper does is recorded in a file called blinkit\_scraper.log. This is useful for spotting issues or understanding what happened if something goes wrong during scraping. ### User-Agent Rotation ``` # Load user-agents from file def load_user_agents(): """ Reads user-agents from a file and returns them as a list. This function opens and reads the user_agents.txt file, strips each line, and returns only non-empty lines as a list of user agent strings. Returns: list: A list of user agent strings to be used for request rotation. Raises: Exception: Any file-related exceptions are caught, logged, and an empty list is returned. Notes: - The file should contain one user agent string per line. - Empty lines are ignored. - If the file cannot be read, an error is logged and an empty list is returned. """ try: with open("user_agents.txt", "r") as file: return [line.strip() for line in file.readlines() if line.strip()] except Exception as e: logging.error(f"Error reading user_agents.txt: {e}") return [] USER_AGENTS = load_user_agents() ``` One smart feature in this script is something called user-agent rotation. This is handled by a function named load\_user\_agents, which reads from a text file filled with different user-agent strings. A user-agent is basically a small piece of information sent to a website that says, “Hey, I’m a browser on this device.” It helps the website understand what kind of user is visiting — whether it’s someone on a phone, laptop, or using a specific browser like Chrome or Firefox. Now, instead of using the same user-agent every time (which might make it obvious that we’re a bot), we randomly change the user-agent for each request. This makes it look like the visits are coming from different people using different devices. This simple trick helps reduce the chances of getting blocked by the website. It also supports responsible scraping — we want to collect data without putting too much strain on the site or drawing unnecessary attention. ### Retrieving Unscraped Links ``` def get_unscraped_links(): """ Fetch all unprocessed links from the database where scraped = 0. This function connects to the SQLite database and retrieves all records from the links table where the scraped flag is set to 0, indicating they haven't been processed yet. Returns: list: A list of tuples containing (id, url, category) for unscraped links. Each tuple contains: - id (int): The primary key of the link record - url (str): The product URL to be scraped - category (str): The category the product belongs to Raises: Exception: Any database-related exceptions are caught, logged, and an empty list is returned. Notes: - The number of fetched links is logged for monitoring. - If an error occurs, it is logged and an empty list is returned. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("SELECT id, url, category FROM links WHERE scraped = 0") links = cursor.fetchall() conn.close() logging.info(f"Fetched {len(links)} unscraped links from the database.") return links except Exception as e: logging.error(f"Error fetching unscraped links: {e}") return [] ``` The get\_unscraped\_links function helps us keep everything organized by acting like a to-do list for our scraper. Here’s how it works: it connects to our SQLite database and looks for all the product links where the scraped value is still set to 0\. This tells us that the product details haven’t been collected yet. By keeping track of which links are done and which ones are pending, this function helps us manage our progress. So, if the script stops in the middle — maybe due to a network issue or system crash — we can easily pick up right where we left off, without starting over. Each link returned by this function includes three things: the ID (from the database), the URL (where the product is), and the category (like fruits or vegetables). We’ll need all of this information in the next steps to scrape and save the product details properly. ### Content Fetching ``` def fetch_page_content(url): """ Fetches the full product page content using Playwright with a rotating User-Agent. This function launches a Chromium browser with a randomly selected user agent, navigates to the specified URL, waits to ensure content is loaded, and then captures the complete HTML content of the page. Parameters: url (str): The URL of the product page to fetch. Returns: str or None: The complete HTML content of the page if successful, None otherwise. Raises: Exception: Any exceptions during browsing are caught, logged, and None is returned. Notes: - Uses non-headless browser mode (visible) which may be changed for production. - Randomly selects a user agent from USER_AGENTS list for request fingerprint randomization. - Waits 10 seconds after page load to ensure dynamic content is rendered. - Includes a timeout of 60 seconds (60000ms) for page loading. - Contains commented code for saving HTML to file for debugging purposes. - Logs the user agent used for each request for tracking. """ try: user_agent = random.choice(USER_AGENTS) if USER_AGENTS else None # Pick a random User-Agent with sync_playwright() as p: browser = p.chromium.launch(headless=False) context = browser.new_context(user_agent=user_agent) page = context.new_page() page.goto(url, timeout=60000) time.sleep(10) content = page.content() browser.close() logging.info(f"Successfully fetched page content for {url} using User-Agent: {user_agent}") return content except Exception as e: logging.error(f"Error fetching page content from {url}: {e}") return None ``` The job of actually visiting each product page is handled by the fetch\_page\_content function. This function uses Playwright to open up a Chrome browser, go to the product link, and wait until the page is fully loaded. We run the browser in non-headless mode, which means it opens up visibly on the screen. This helps make sure that all parts of the page — especially the content loaded through JavaScript — have a chance to appear properly. To be extra safe, the function also waits 10 more seconds after the page says it’s done loading. This gives any slow-loading elements, like pop-ups or late-appearing product details, time to show up. This extra wait ensures we capture all the product information, even the parts that don’t appear immediately when the page first opens. ### Parsing Product Details The parsing functions are the most important part of our scraper. Each one is designed to extract a specific detail about the product. These functions work like tools, each picking out different information from the page. ``` def parse_product_name(soup): """ Extracts the product name from a Blinkit product page. This function uses a CSS selector to locate and extract the product name from the BeautifulSoup object representing the product page. Parameters: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The product name if found, "N/A" otherwise. Notes: - Uses a specific CSS selector targeting the product name element on Blinkit pages. - If the element is not found, logs a warning and returns "N/A". - The selector may need updating if Blinkit changes their page structure. """ try: return soup.select_one("#app > div > div > div:nth-child(3) > div > div.Product__ProductWrapper-sc-18z701o-3.eJdNpg > div.Product__ProductWrapperRightSection-sc-18z701o-5.hbMSwc > div.ProductInfoCard__ProductInfoWrapper-sc-113r60q-3.iwxjxo > h1").text.strip() except AttributeError: logging.warning("Product name not found.") return "N/A" ``` The parse\_product\_name function finds and collects the product name from the webpage. It uses a CSS selector to locate the title section of the product quickly and accurately. ``` def parse_net_quantity(soup): """ Extracts the net quantity information from a Blinkit product page. This function locates and extracts the net quantity (weight, volume, or units) from the BeautifulSoup object of the product page. Parameters: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: The net quantity information if found, "N/A" otherwise. Notes: - Uses a specific CSS selector targeting the quantity element on Blinkit pages. - If the element is not found, logs a warning and returns "N/A". - The selector may need updating if Blinkit changes their page structure. """ try: return soup.select_one("p.ProductVariants__VariantUnitText-sc-1unev4j-6.dhCxof").text.strip() except AttributeError: logging.warning("Net quantity not found.") return "N/A" ``` The parse\_net\_quantity function gets the product’s weight or size (like grams or kilograms). This helps us understand how much of the fruit or vegetable is being sold. ``` def parse_price_details(soup): """ Extracts the sale price and MRP (Maximum Retail Price) from a Blinkit product page. This function locates and extracts both the current sale price and the original MRP from the BeautifulSoup object of the product page. Parameters: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: tuple or None: A tuple containing (sale_price, mrp) if successful, None otherwise. - sale_price (str): The current selling price - mrp (str): The Maximum Retail Price (original price) Raises: Exception: Any exceptions during parsing are caught, logged, and None is returned. Notes: - Uses specific CSS selectors targeting the price elements on Blinkit pages. - If elements are not found, returns 'N/A' for the respective values. - The selectors may need updating if Blinkit changes their page structure. """ try: # Extract sale price sale_price_element = soup.select_one("div.ProductVariants__PriceContainer-sc-1unev4j-7.gGENtH") sale_price = sale_price_element.get_text(strip=True) if sale_price_element else 'N/A' # Extract MRP mrp_element = soup.select_one("span.ProductVariants__MRPText-sc-1unev4j-8.gNKjjk") mrp = mrp_element.get_text(strip=True) if mrp_element else 'N/A' # Return individual price details return sale_price, mrp except Exception as e: logging.error(f"Error parsing price details: {e}") return None ``` The parse\_price\_details function collects two key prices from the product page: - Sale Price – the current price the customer pays. - MRP (Maximum Retail Price) – the original price before any discounts. This helps us compare prices, track discounts, and analyze pricing trends over time. ``` def parse_details(soup): """ Extracts detailed product information from a Blinkit product page. This function locates and extracts all available product details such as description, ingredients, nutritional information, etc. from the product details section of the page. Parameters: soup (BeautifulSoup): The BeautifulSoup object of the product page. Returns: str: A JSON string containing key-value pairs of product details. Returns an empty JSON object string "{}" if no details are found. Raises: Exception: Any exceptions during parsing are caught, logged, and an empty JSON object string is returned. Notes: - Uses specific CSS selectors targeting the product details section. - Each key-value pair in the details section is extracted and added to a dictionary. - The dictionary is then converted to a formatted JSON string with indentation. - If an error occurs processing a specific detail, that detail is skipped. - The selector may need updating if Blinkit changes their page structure. """ try: details_section = soup.select("#app > div > div > div:nth-child(3) > div > div.Product__ProductWrapper-sc-18z701o-3.eJdNpg > div.Product__ProductWrapperLeftSection-sc-18z701o-4.fkQLTf > div:nth-child(3) > div > div.ProductDetails__RemoveMaxHeight-sc-z5f4ag-3.fOPLcr > div") if not details_section: logging.warning("No product details found in the given HTML.") return json.dumps({}) details = {} for div in details_section: try: key_element = div.find("p") value_element = div.find("div") if key_element and value_element: key = key_element.get_text(strip=True) value = value_element.get_text(strip=True) details[key] = value else: logging.warning("Missing key-value pair in a product highlight div.") except Exception as e: logging.error(f"Error processing a highlight div: {e}") continue return json.dumps(details, indent=4) except Exception as e: logging.error(f"Error parsing product highlights: {e}") return json.dumps({}) ``` The parse\_details function extracts extra product information that might differ from one item to another. It looks through the product details section, collects key-value pairs (like "Shelf Life: 3 days" or "Country of Origin: India"), and organizes them into a clear JSON format. This makes the data easy to store and analyze later. ``` def parse_product_details(html_content): """ Parses all product details from the HTML content of a Blinkit product page. This function serves as an orchestrator that calls individual parsing functions to extract different components of product information and combines them into a comprehensive product data dictionary. Parameters: html_content (str): The raw HTML content of the product page. Returns: dict or None: A dictionary containing all parsed product details if successful, None otherwise. The dictionary includes: - name (str): Product name - net_quantity (str): Product quantity/weight - sale_price (str): Current selling price - price (str): Original MRP (Maximum Retail Price) - product_details (str): JSON string of additional product details Raises: Exception: Any exceptions during parsing are caught, logged, and None is returned. Notes: - Creates a BeautifulSoup object from the HTML content for parsing. - Calls specialized parsing functions for each data point. - Logs the extracted product data for monitoring and debugging. """ try: soup = BeautifulSoup(html_content, 'html.parser') product_data = { "name": parse_product_name(soup), "net_quantity": parse_net_quantity(soup), "sale_price": parse_price_details(soup)[0], "price":parse_price_details(soup)[1], "product_details":parse_details(soup) } logging.info(f"Extracted product details: {product_data}") return product_data except Exception as e: logging.error(f"Error parsing product details: {e}") return None ``` The parse\_details function acts like the master coordinator. It brings together all the smaller parsing functions to collect complete product information in one go. It calls each of the specialized functions — the ones that extract the name, quantity, price, and other details — and then neatly combines everything into a single dictionary. This structured format makes it easy to work with the data later. This modular approach keeps the code clean and easy to update. If you ever want to extract more details from the page, you can simply add another parsing function and plug it into this one. ### Database Storage ``` def save_product_data(product_data, category, scraped_date, url): """ Saves product details into the database and returns success status. This function connects to the SQLite database, creates the products table if it doesn't exist, and inserts the product data along with category, scraping date, and URL information. Parameters: product_data (dict): Dictionary containing parsed product details. category (str): Category of the product (e.g., "vegetables", "fruits"). scraped_date (str): Date when the product was scraped (YYYY-MM-DD format). url (str): URL of the product page. Returns: bool: True if data was successfully saved, False otherwise. Raises: Exception: Any database-related exceptions are caught, logged, and False is returned. Notes: - Creates the products table if it doesn't exist. - The table schema includes: * url: URL of the product page * name: Product name * net_quantity: Weight/quantity of the product * sale_price: Current selling price * price: Original MRP (Maximum Retail Price) * category: Product category * product_details: JSON string of additional product details * scraped_date: Date when the product was scraped - Logs success or failure for debugging and monitoring. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS products ( url TEXT, name TEXT, net_quantity TEXT, sale_price TEXT, price TEXT, category TEXT, product_details TEXT, scraped_date TEXT ) """ ) cursor.execute(""" INSERT INTO products (url, name, net_quantity, sale_price, price, category, product_details, scraped_date) VALUES (?, ?, ?, ?, ?, ?, ?, ?) """, (url, product_data["name"], product_data["net_quantity"], product_data["sale_price"], product_data["price"], category, product_data["product_details"], scraped_date)) conn.commit() conn.close() logging.info(f"Product data saved successfully for {url}") return True except Exception as e: logging.error(f"Error saving product data to database: {e}") return False ``` The save\_product\_data function plays a key role in saving the scraped product details to our database. First, it checks if there’s a table called products in our SQLite database. If it doesn’t exist yet, the function creates one with all the necessary columns — like the product name, quantity, price, other specifications (saved as JSON), URL, category, scraping date, and more. Once the table is ready, the function takes the product information we’ve collected and inserts it into the database. This organized structure makes it easy to access and work with the data later — whether for analysis, reporting, or even visualizations using tools like Excel or Power BI. ``` def mark_link_as_scraped(link_id): """ Updates the `scraped` column in the `links` table to mark the link as processed. This function connects to the SQLite database and updates the scraped flag to 1 for the specified link ID, indicating that the link has been processed. Parameters: link_id (int): The ID of the link record to mark as scraped. Returns: None Raises: Exception: Any database-related exceptions are caught and logged. Notes: - Sets the scraped flag to 1 to prevent re-processing in future runs. - Logs the ID of the link being marked for tracking and debugging. """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute("UPDATE links SET scraped = 1 WHERE id = ?", (link_id,)) conn.commit() conn.close() logging.info(f"Marked link ID {link_id} as scraped.") except Exception as e: logging.error(f"Error updating scraped status for link ID {link_id}: {e}") ``` You’re absolutely right — the mark\_link\_as\_scraped function is a crucial checkpoint in the scraping process. After a product link has been successfully processed, this function updates its scraped value to 1 in the database. Why is this so important? - No Extra Work: Once a link is marked as scraped, the scraper knows not to visit it again. This avoids unnecessary repetition and saves time. - Smooth Recovery: If the script stops or crashes partway through, it can pick up exactly where it left off — starting from the next unprocessed link. - Clean Data: It helps keep the products table free of duplicates. Since each product is only scraped once, you won’t end up with the same product saved multiple times. This small function becomes especially valuable in larger or longer scraping jobs, where reliability and clean, accurate data really matter. ### The Scraping Workflow ``` def scrape_products(): """ Main function to scrape product details from all unscraped links in the database. This function orchestrates the complete scraping workflow: 1. Retrieves all unscraped links from the database 2. For each link, fetches the HTML content 3. Parses product details from the HTML 4. Saves the details to the database 5. Marks the link as scraped 6. Introduces a random delay between requests to avoid detection Returns: None Notes: - If no unscraped links are found, logs an informational message and exits. - Implements a random delay between 8 and 15 seconds between requests to avoid overwhelming the server and to mimic human browsing behavior. - Only marks a link as scraped if the product data was successfully saved. - Logs progress and completion for monitoring and debugging. """ unscraped_links = get_unscraped_links() if not unscraped_links: logging.info("No unscraped links found. Exiting scraper.") return for link_id, url, category in unscraped_links: logging.info(f"Processing {url} (Category: {category})") html_content = fetch_page_content(url) if html_content: product_data = parse_product_details(html_content) if product_data: scraped_date = datetime.now().strftime("%Y-%m-%d") if save_product_data(product_data, category, scraped_date, url): mark_link_as_scraped(link_id) logging.info(f"Successfully saved and marked {url} as scraped.") else: logging.warning(f"Skipping marking {url} as scraped due to save failure.") time.sleep(random.uniform(8, 15)) # Random delay between processing requests logging.info("Scraping completed for all available links.") if __name__ == "__main__": scrape_products() ``` The scrape\_products function is the heart of the script. It goes through all the product links that haven’t been scraped yet and processes them one by one. To act more like a real user and avoid putting too much pressure on Blinkit’s servers, the scraper waits for a random amount of time — between 8 and 15 seconds — before moving on to the next link. This small delay helps reduce the chances of being blocked and shows respect for the website’s limits. The script is designed to run this function only when called directly, so if you ever want to import this code into another project, it won’t start scraping automatically — which gives you more control. Overall, this scraper follows good scraping practices. It rotates user agents to avoid detection, adds delays between requests, collects detailed product information, and stores everything neatly in a database. The code is modular, meaning it’s easy to update or expand if Blinkit’s site changes or if you want to scrape other sections later. With clear logs showing each step, you can easily track how the scraper is doing. All in all, it’s a solid tool for gathering and analyzing data from Blinkit’s fruits and vegetables section. ## Conclusion In this project, we successfully built a browser automation and scraping workflow using Blinkit’s fruits and vegetables catalog as an example. By breaking the process into two clear steps — first collecting product links, then visiting each link to gather detailed information — we created a reliable system that collects useful data without overwhelming the website. To make sure our scraping was both ethical and stable, we followed standard best practices. We rotated user agents, added random wait times between requests, and kept detailed logs of everything the scraper did. Using SQLite as a lightweight database also gave us a smart way to track progress and recover easily if something interrupted the scraping. The code is structured in a way that’s easy to update or expand, whether you want to scrape more product details or add new categories later. AUTHOR I’m Shahana, a Data Engineer at Datahut, where I design robust, scalable data pipelines that turn messy web content into clean, usable datasets—especially in fast-moving industries like e-commerce, grocery delivery, and retail insights. At Datahut, we work with clients to automate data collection from modern, JavaScript-heavy websites using tools like Playwright and BeautifulSoup. In this blog, I shared a hands-on scraping project built around Blinkit’s Fruits and Vegetables catalog, demonstrating a two-step strategy for gathering product links and detailed data while keeping the process efficient, reliable, and ethical. If your team is looking to automate product data collection in the eyewear space or beyond, reach out to us through the chat widget on the right. We’d love to help you build a solution that fits your goals. FAQs 1\. What is web scraping and how does it work for Blinkit? Web scraping is the process of using automated tools to extract data from websites. In the case of Blinkit, we use a scraper to visit product listing pages, collect links to fruits and vegetables, and then gather details like name, price, and weight from each product page. 2\. Is it legal to scrape data from Blinkit? Web scraping legality depends on the site's terms of service and local laws. Always review Blinkit’s terms, avoid overloading their servers, and use the data responsibly for research or analysis. 3\. Why collect fruits and vegetables data from Blinkit? Blinkit offers a wide variety of fresh produce. By scraping this data regularly, businesses and researchers can track price changes, monitor stock availability, and identify seasonal trends. 4\. What tools are used for scraping Blinkit data? We use tools like Playwright to automate the browser, BeautifulSoup to parse HTML, and SQLite to store the scraped data. Logging helps track the scraper’s performance, and user-agent rotation reduces the risk of being blocked. 5\. Can this scraping method be used for other categories on Blinkit? Yes. By adjusting the target URLs and category labels in the script, you can scrape data for other Blinkit categories like dairy, snacks, or packaged goods. ### India Luxury Watch Market Insights from Ethos Data URL: https://www.blog.datahut.co/post/inside-ethos-what-data-analysis-reveals-about-india-s-luxury-watch-market/ Last updated: 2026-09-07T09:44:11.000Z How much does the average Ethos luxury watch cost? Which brands lead in value, variety, or horological innovation? Are Indian buyers still favoring traditional analog models over modern smartwatch alternatives? These aren’t just speculative questions. These are data-backed insights that paint a rich picture of how India embraces luxury in timekeeping — and why a strong, [data-driven retail strategy](https://www.blog.datahut.co/post/two-reasons-why-a-retailer-must-embrace-a-data-first-growth-strategy/) matters more than ever.” We analyzed thousands of listings from [Ethos Watches](https://www.ethoswatches.com/?ref=blog.datahut.co), India’s leading luxury watch retailer, to understand how price, design, heritage, and origin shape the preferences of Indian watch buyers. What emerged was more than just trends—it was a window into what luxury means in India today. ![Ethos Watches](https://www.blog.datahut.co/content/images/2026/07/img-232.png.webp) ## The Sweet Spot of Pricing: Between Prestige and Practicality India's luxury watch buyers may love heritage, but they also understand value. The data reveals that the majority of watches listed on Ethos fall in the ₹1.5 to ₹3 lakh price range. This sweet spot blends affordability (in the luxury context) with prestige. We’ve seen similar ‘value bands’ in our breakdowns of [H&M’s pricing strategy](https://www.blog.datahut.co/post/h-m-s-pricing-strategy-detailed-analysis-data/) and [Zara’s pricing strategy](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/) Yet, there’s room for opulence. Ethos carries pieces that go far beyond this mid-premium tier—some crossing the ₹1 crore mark. These ultra-high-end timepieces, often from boutique Swiss brands, cater to elite collectors and HNIs. Interestingly, the pricing distribution follows a right-skewed curve, meaning that while most watches are clustered around ₹1–5 lakh, a few astronomical prices stretch the upper limit—creating a balanced ecosystem of accessible luxury and rare indulgence. ## Analog vs. Smart: A Clear Winner in the Indian Luxury Space In a world shifting rapidly toward smart tech, Ethos seems to be holding its ground. An overwhelming 99.9% of Ethos' catalog consists of analog watches, signaling a strong consumer preference for traditional mechanical or quartz timepieces over digital or hybrid models. The luxury consumer in India is clear—they seek craftsmanship, not notifications. Ethos, by focusing almost exclusively on analog, caters to those who see watches as art, not gadgets. ## Brand Battles: Volume vs. Value In India’s Luxury Watch Market In terms of product listings, Omega takes the crown with over 450 models on the platform. Its dominance is followed by Rado, Longines, and Tissot, brands known for offering a wide range of Swiss-made options with varying price points. But when it comes to average price, it’s a different story. Boutique brands like Laurent Ferrier (₹1 crore+), Jacob & Co., Urwerk, and Bovet soar to the top. These ultra-luxury makers produce fewer pieces, each a technical marvel—turning every watch into a statement of taste and wealth. This dual dynamic reveals Ethos' twofold strategy: - Serve the premium-conscious buyer with variety from brands like Omega or Rado. - Satisfy connoisseurs and collectors with exclusive, ultra-high-end timepieces. ## Warranty, Durability, and Trust Warranty is a crucial differentiator in luxury. Omega again leads, offering most of its models with a 5-year warranty. Brands like Breitling and Titoni go a step further with warranties extending up to 8 years, offering peace of mind to high-investment buyers. For those considering not just beauty but long-term value and reliability, these warranties become important indicators of brand confidence and product longevity — and a perfect example of how [e-commerce product enrichment](https://www.blog.datahut.co/post/e-commerce-product-enrichment/) can surface the details buyers actually care about.. ## Design Meets Durability: Glass and Glow When it comes to design and materials, Sapphire Crystal glass dominates, especially in watches featuring luminous elements like glowing hands, hour markers, and bezels. This combination signifies Ethos’ emphasis on both aesthetics and durability. While Mineral Crystal plays a distant second, other materials like Plexiglass and Hesalite remain niche, often reserved for retro or very specific styles. The preference for Sapphire Crystal also reinforces the brand’s alignment with premium craftsmanship and scratch-resistant luxury. ## Color & Gender: The Subtleties of Style Color reveals more than fashion—it reflects identity. Silver straps reign across all categories, but the gender breakdown is telling: - Men gravitate toward black and brown, echoing professionalism and versatility. - Women lean into rose gold and white, reflecting elegance and refined style. - Unisex models, though fewer in number, stick to neutral tones—silver and black—catering to universal appeal. This stratification offers insights for both watchmakers and marketers to fine-tune gender-specific campaigns and product designs, and to [increase your eCommerce product’s visibility](https://www.blog.datahut.co/post/increase-your-ecommerce-product-visibility/) where it matters most ## Power Reserves: Watches That Outlast the Week Luxury isn’t just about looks—it’s about engineering too. BVLGARI leads with a jaw-dropping 652-hour average power reserve, nearly 27 days without a wind. For comparison, most other high-end brands sit in the 50–100 hour range, with Bovet, Omega, and Armin Strom exceeding 100 hours. Such performance points to technical mastery, positioning these brands not just as fashion statements but as mechanical marvels. ## Swiss (Still) Made: Country of Origin Breakdown When it comes to origin, Switzerland remains the heart of luxury watchmaking. Over 60 brands on Ethos are Swiss, far outpacing any other country. Germany, China, and Japan follow—but by a wide margin. This dominance reinforces Swiss reputation in horology and reflects Indian consumer trust in the legacy of Swiss craftsmanship. ## Why These Insights Matter Whether you’re a watch buyer, brand strategist, or just a curious enthusiast, the data from Ethos offers more than just numbers—it offers a lens into Indian luxury culture. - Buyers can spot value zones and plan their investments wisely. - Retailers and merchandisers gain clarity on product range, customer taste, and pricing sweet spots. - Brand managers get insights into positioning, warranty appeal, and material preferences. - Analysts and enthusiasts gain access to rare, structured views of luxury consumption patterns — a textbook example of how to [use big data to create value for customers](https://www.blog.datahut.co/post/use-big-data-create-value-customers/). India's luxury watch market is no longer just about the brand. It’s about design, engineering, tradition, and consumer identity—woven together through data. ![Ethos Watches Dataset ](https://www.blog.datahut.co/content/images/2026/07/img-233.png.webp) ## Powered by Datahut: Explore the Full Dataset This deep dive was made possible by Datahut, a leading provider of e-commerce data and web scraping solutions. Specializing in transforming raw web data into structured insights, Datahut empowers businesses to understand markets, competitors, and consumer behavior with clarity. Want the full picture? (Data scraped and processed by Datahut) ## FAQs 1. What is the average price of a luxury watch on Ethos?Most luxury watches on Ethos fall between ₹1.5–3 lakh, striking a balance between premium appeal and affordability. 2. Which brands are the most expensive on Ethos?Laurent Ferrier, Jacob & Co., and Urwerk top the chart with average prices reaching up to ₹1 crore. 3. Are analog watches still popular in India’s luxury segment?Yes, 99.9% of Ethos watches are analog, highlighting the strong preference for traditional craftsmanship among Indian buyers. 4. Which country dominates the luxury watch market on Ethos?Switzerland leads with over 60 brands listed, establishing its supremacy in India’s premium watch segment. 5. Who provided the data for this analysis?The data was scraped and processed by [Datahut](http://datahut.co/?ref=blog.datahut.co), a web scraping and data intelligence company. ### What Fashion Brands Can Learn from Zara Pricing Data URL: https://www.blog.datahut.co/post/what-fashion-retailers-can-learn-from-zara-s-pricing-strategy/ Last updated: 2026-09-07T09:44:13.000Z ## Introduction Zara’s pricing strategy is a benchmark in fast fashion retail. Known for combining trend responsiveness with smart pricing, Zara balances affordable fashion with premium appeal. Through data-driven pricing, the brand adjusts product costs by region, fabric type, and category—while maintaining brand value through minimal discounting. In this blog, we explore Zara’s pricing insights, backed by web scraping techniques and product-level data analysis, to uncover what fashion retailers can learn and implement in their own pricing strategies. 👉 Looking for an in-depth breakdown with detailed visuals and statistics?[ ](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/)[Check out our full analysis of Zara’s pricing strategy here](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/). ![Understanding Zara's Pricing Strategy: Analyzing Product Price Distribution](https://www.blog.datahut.co/content/images/2026/07/img-91.jpg.webp) ## Strategic Takeaways from Zara’s Pricing Strategy ### 1\. Affordable Meets Premium: A Balanced Price Range Zara's prices span from ₹490 to ₹15,590, but most items cluster under ₹6,000\. This lets the brand cater to both budget shoppers and aspirational buyers. Real-World Parallel: Like H&M’s ₹299–₹29,999 strategy, Zara uses a pricing pyramid: basics at the base, statement pieces at the top. Retailer Tip: Offer accessible price anchors with a few premium SKUs to elevate brand perception. ### 2\. Localized Pricing for Global Relevance Zara adjusts average product prices by region—e.g., higher in Myanmar and India, lower in Portugal. Retailer Insight: Uniqlo follows a similar path with region-specific pricing that factors in import costs and brand perception. Retailer Tip: Use web scraping and market data to optimize region-specific pricing ### 3\. Material and Quality Drive Pricing Sheep leather and silk items top Zara’s price list, while cotton and polyester dominate the mid-range. Real-World Scenario: H&M Conscious collection uses this model—eco fabrics cost more but cater to a niche, quality-focused audience. Retailer Tip: Price according to perceived fabric value. Upsell premium materials through storytelling ### 4\. Category-Based Pricing Anchors Coats and blazers are high-ticket items, while tops and accessories are entry points. Retailer Parallel: Mango and Massimo Dutti (Inditex siblings) use similar pricing anchors—luxury outerwear, accessible separates. Retailer Tip: Let premium outerwear offset margin losses on fast-moving essentials. ### 5\. Discounts Done Sparingly Only 3.1% of Zara products are discounted—most between 40–50%. It uses discounts to move inventory, not to drive regular traffic. Compare With: Brands like Forever 21 rely on constant markdowns, diluting brand value. Zara preserves its premium perception. Retailer Tip: Don’t over-discount. Use markdowns as strategic inventory levers—not customer expectations. 👉The pricing insights shared here were powered by structured web scraping and analysis. Check out our step-by-step guide to scraping Zara’s catalog:[ ](https://www.blog.datahut.co/post/decoding-fashion-a-beginner-s-guide-to-scraping-zara-s-online-catalog/)[Decoding Fashion – Beginner’s Guide](https://www.blog.datahut.co/post/decoding-fashion-a-beginner-s-guide-to-scraping-zara-s-online-catalog/) ![Decoding Fashion: A Beginner's Guide to Scraping Zara's Online Catalog](https://www.blog.datahut.co/content/images/2026/07/img-92.jpg.webp) ## Challenges & Opportunities in Fashion Pricing ### Challenges - Global Price Sensitivity: Uniform pricing can cause friction in price-sensitive markets. - Fast Trend Cycles: Overpricing rapidly outdated items can lead to unsold stock. - Supply Chain Volatility: Costs of premium materials (like silk/leather) can fluctuate with global events. ### Opportunities - Dynamic Pricing Models: Using AI to adjust prices in real time (Zara is already doing this). - Ethical Pricing Strategies: Premium sustainability-focused collections are increasingly acceptable at higher price points. - Data-Driven Markdown Planning: Limit discounts to low-performing items based on live data, preserving brand equity. ## Conclusion Zara proves that pricing isn’t just a finance decision—it’s a brand strategy. From category anchoring to disciplined discounting, its pricing helps drive both profit and perception. Retailers who embrace smart pricing, powered by real-time data, can stand out in a crowded market. To explore the complete dataset, box plot and histogram analyses, and global material trends,[ read the full Zara pricing analysis blog](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/). ## Want to replicate Zara’s pricing success using data? [Talk to Datahut](https://www.datahut.co/?ref=blog.datahut.co) — we help fashion retailers harness web scraping and pricing analytics to stay ahead of the curve. ## FAQ SECTION 1. What makes Zara’s pricing strategy so effective?Zara combines affordability with aspirational value by offering a wide price range, minimal discounts, and region-specific pricing—all backed by real-time data. Explore Zara’s full pricing insights! 2. How does Zara use data to determine prices?Zara leverages product-level data and market trends to adjust prices by category, fabric, and region. Read our Beginner’s Guide to Scraping Zara’s Catalog for a behind-the-scenes look. 3. Should fashion retailers always avoid discounts like Zara?Not necessarily. Zara uses discounts strategically to clear stock, while other brands may rely on heavy markdowns. Learn why strategic discounting helps maintain brand value. 4. How can other retailers implement Zara’s pricing model?Retailers can analyze competitors’ prices using web scraping tools, then apply dynamic pricing and premium positioning tactics. Talk to Datahut for custom data solutions. 5. Why is localized pricing important in fashion retail?Pricing based on regional demand and costs helps align with local markets—Zara practices this successfully. See how H&M and Uniqlo approach it too. ### Blinkit Fresh Basket Pricing and Inventory Data Analysis URL: https://www.blog.datahut.co/post/inside-blinkit-s-fresh-basket-a-3-day-data-driven-analysis-of-prices-stock-and-strategy/ Last updated: 2026-09-07T09:44:16.000Z ## Introduction In the hyper-competitive world of online groceries, where minutes matter and freshness drives loyalty, [Blinkit](https://blinkit.com/?ref=blog.datahut.co) has carved out a space by delivering essentials at lightning speed. But beyond the sleek UI and express delivery promise, what’s happening behind the curtain in terms of pricing, stock levels, and category strategy? To find out, we analyzed a compact yet powerful dataset: three consecutive days (March 17–19, 2025) of product data from Blinkit’s Fruits & Vegetables category. This time-series dataset includes over 1,100 data points, capturing fluctuations in sale prices, availability, and discounts across one of the most dynamic segments in grocery. While three days might sound like a short time, the data tells a compelling story about stock movement, price strategy, and category balance. Let’s peel back the layers of Blinkit’s fresh shelf. ## Section 1: Stock Availability – Blinkit’s Shelf Recovery in Action Stock availability is at the heart of every grocery platform’s success. A “Sorry, out of stock” message on your daily essentials can be a deal-breaker for customers. Blinkit seems to understand this well. ![Blinkit stock availability](https://www.blog.datahut.co/content/images/2026/07/img-383.png.webp) ### Key Takeaways: - March 17: 219 items were available, while 155 were out of stock. - March 18: Availability improved to 240, while out-of-stock items reduced to 124. - March 19: Highest availability at 263, with only 110 out-of-stock—nearly a 30% improvement in availability over 3 days. ### What This Suggests: - Blinkit rapidly recovered inventory, possibly through real-time restocking or backend optimization. - The increasing stock trend and decreasing out-of-stock count over just three days reflect a responsive supply chain. - March 19 marked the best day in terms of inventory health, highlighting the platform’s ability to adjust quickly to demand. For quick-commerce, even short outages can cause frustration. This data suggests Blinkit is not only aware of it but is actively countering inventory gaps with agile restocking processes. ## Section 2: Blinkit’s Price Movement – A Micro-Trend with Macro Implications Price fluctuations in fresh produce are expected. But even within a tight 3-day window, Blinkit displayed noticeable price volatility—especially in sale prices. ![Blinkit sale price ](https://www.blog.datahut.co/content/images/2026/07/img-384.png.webp) ### Key Takeaways: - On March 17, the average sale price was ₹125. - On March 18, prices surged to ₹145 — a 16% jump. - On March 19, prices corrected to ₹135, slightly above the starting point. ### Why This Matters: - A ₹20 increase in just 24 hours reflects supply constraints, demand surges, or the restocking of higher-value items. - The minor drop on Day 3 suggests that while availability improved, Blinkit maintained a higher pricing tier—perhaps signaling a shift in product mix (e.g., premium fruits or organic vegetables). - Despite restocking more items, Blinkit didn't drop prices aggressively, which might indicate limited price elasticity or confident pricing. This shows how Blinkit’s pricing algorithm may respond to short-term supply shifts, and that customers might want to time their purchases accordingly. ## Section 3: Discounts in Action – Transparent and Reliable How do listed prices compare to what shoppers actually pay? A look at Blinkit’s discount behavior across three days reveals a consistent markdown strategy. ![Blinkit Sale vs Price](https://www.blog.datahut.co/content/images/2026/07/img-385.png.webp) ### Key Takeaways: - Sale prices were consistently lower than original listed prices. - This discount trend was present across all product tiers, not just for premium SKUs. - While original prices had a wider spread (especially March 18 and 19), sale prices remained in a tighter band—signaling controlled and even discounts. ### What This Means: - Blinkit applies price drops broadly rather than selectively, helping create trust in the platform’s pricing. - The consistent markdown behavior over three days suggests there was no flash discounting or sudden price wars—just a steady, shopper-friendly strategy. - For price-sensitive buyers, this reliability matters. You don’t have to chase deals—they’re already baked in. It’s a well-balanced pricing model that rewards both volume and loyalty without relying on gimmicky discounts. ## Section 4: Fruits vs Vegetables – Inventory Imbalance and Strategic Prioritization Fruits and vegetables may share shelf space, but Blinkit doesn’t treat them equally—and neither does the supply chain. ![Blinkit Instock vegetables](https://www.blog.datahut.co/content/images/2026/07/img-386.png.webp) ### Key Takeaways: - Fruits consistently had higher in-stock numbers than vegetables. - On March 18, the gap widened to 160 fruits vs. 79 vegetables. - Vegetables dipped sharply on Day 2, recovering mildly on Day 3 (96 in stock). - Fruits showed continuous growth in availability, from 119 → 160 → 167. ### Interpretation: - Fruits may be prioritized due to higher shelf life, better procurement channels, or greater demand predictability. - Vegetables, likely sourced more locally and perishable, showed instability—possibly hinting at restocking or sourcing issues. - March 18’s vegetable dip could also be due to a sales spike rather than a supply issue. Blinkit’s apparent emphasis on keeping fruits consistently stocked may point to revenue optimization—especially when we compare pricing next. ## Section 5: Price Disparity – Fruits Bring Margin, Veggies Drive Volume Fresh produce isn’t just about freshness—it’s about margins and basket economics. And Blinkit’s data shows fruits dominate revenue potential. ![Blinkit average sale price ](https://www.blog.datahut.co/content/images/2026/07/img-387.png.webp) ### Key Takeaways: - Fruits were priced 4–5 times higher than vegetables across all three days.Fruits: ₹189–₹198 rangeVegetables: ₹38–₹45 range - March 18 was pivotal:Fruits hit their highest average price (₹198)Vegetables hit their lowest (₹38) ### Strategic Implications: - Fruits, likely including exotic and seasonal varieties, fetch better margins per unit. - Vegetables, though lower-priced, may be Blinkit’s footfall driver—offering affordable essentials to keep daily users engaged. - This pricing duality supports a “high-margin + high-volume” revenue model. For Blinkit, this means balancing stock and pricing carefully. For shoppers, it suggests that premium fruits are driving up your cart totals—while veggies help balance your budget. ## Conclusion: 3 Days, 1 Clear Story Despite the short window, Blinkit’s fresh produce data reveals sharp insights into the operational heartbeat of an online grocery app: - Stock availability improved visibly over 3 days. - Prices shifted dynamically, with a clear midweek peak. - Discounts were broad, consistent, and reliable. - Fruits led in availability and pricing, while vegetables lagged slightly in stock but offered affordability. - Fruits = Margin, Vegetables = Volume—a balanced playbook for grocery success. For consumers, the takeaway is clear: shopping on Blinkit means predictable savings, but basket totals can swing depending on the day and product mix.For businesses? It’s a case study in responsive inventory management, consistent discounting, and revenue optimization through strategic category focus. Data used for this analysis has been extracted using web scraping from the Blinkit website, cleaned and enriched using different techniques. Contact us to access the complete web scraped data from Blinkit. Connect with[ Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the information you need, hassle-free. ## Frequently Asked Questions (FAQs) ### 1\. How was the Blinkit data collected for this analysis? The data used in this analysis was collected using[ ](https://www.blog.datahut.co/post/introduction-to-web-scraping-with-python/)[web scraping techniques](https://www.blog.datahut.co/post/introduction-to-web-scraping-with-python/), specifically tailored for e-commerce platforms like Blinkit. We used a Python-based scraper that tracked fresh produce listings over a 3-day window. For more technical guidance on building your own scraper, check out our tutorial: 👉[ ](https://www.blog.datahut.co/post/how-to-scrape-blinkit-s-fruits-and-vegetables-data/)[How to Scrape Blinkit Product Data Using Python](https://www.blog.datahut.co/post/how-to-scrape-blinkit-s-fruits-and-vegetables-data/) ### 2\. Why does Blinkit show such sharp price fluctuations in fruits and vegetables? Blinkit operates in the quick-commerce space, which means inventory and prices can change rapidly due to supply chain dynamics, demand surges, and product mix changes (like premium fruits being stocked). This pricing behavior is similar to what we’ve seen in other dynamic retail platforms—like in our blog on 👉[ ](https://www.blog.datahut.co/post/what-fashion-retailers-can-learn-from-h-m-s-data-driven-pricing/)[H&M’s Data-Driven Pricing Strategy](https://www.blog.datahut.co/post/what-fashion-retailers-can-learn-from-h-m-s-data-driven-pricing/) and 👉[ ](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/)[Zara’s Product Price Distribution](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/). ### 3\. Are Blinkit's discounts applied consistently across all products? Yes, based on the violin plot visualization, our 3-day EDA shows that discounts are not exclusive to high-end items. Blinkit applies consistent markdowns across fruits and vegetables, making the sale price range narrower and more predictable. You can explore a similar trend in our 👉[ ](https://www.blog.datahut.co/post/decoding-office-depot-an-exploratory-analysis-of-pricing-discounts-and-printer-trends/)[Office Depot Pricing Analysis](https://www.blog.datahut.co/post/decoding-office-depot-an-exploratory-analysis-of-pricing-discounts-and-printer-trends/), where discounts were strategically applied across SKUs. ### 4\. What’s the difference between fruit and vegetable strategy on Blinkit? Fruits generally had higher availability and were priced 4–5x more than vegetables during the 3-day period. This suggests a margin-maximizing approach for fruits and a volume-driven approach for vegetables. This mirrors inventory strategies discussed in other product-focused analyses, such as 👉[ ](https://www.blog.datahut.co/post/lowe-s-exploratory-data-analysis-trends-and-insights-in-smart-home-technology/)[Our Smart Home Trends EDA at Lowe’s](https://www.blog.datahut.co/post/lowe-s-exploratory-data-analysis-trends-and-insights-in-smart-home-technology/), where high-margin and high-volume categories were handled differently. For more on category behavior, see the full Blinkit analysis above or check Blinkit’s product lineup on their[ ](https://blinkit.com/?ref=blog.datahut.co)[official website](https://blinkit.com/?ref=blog.datahut.co). ### 5\. How can businesses use this Blinkit EDA for competitive advantage? If you're running an online grocery, retail tech startup, or supply chain business, you can use short-window time-series data (like Blinkit’s) to make decisions on pricing optimization, discount models, and real-time inventory restocking. Learn more about scraping competitor prices in this guide: 👉[ ](https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/)[Free Web Scraping Workflow for Competitor Price Tracking](https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/) And if you're new to scraping altogether, start here: 👉[ ](https://www.blog.datahut.co/post/introduction-to-web-scraping-with-python/)[What Is Web Scraping & Why It Matters for E-commerce](https://www.blog.datahut.co/post/introduction-to-web-scraping-with-python/) ### Perfume Market Insights from Boutiqaat Data URL: https://www.blog.datahut.co/post/how-to-scrape-boutiqaat-for-actionable-perfume-market-insights/ Last updated: 2026-07-23T07:48:31.000Z Have you ever thought about how companies manage to track thousands of products, prices, and even trends in multiple online stores all at once? That's where web scraping comes in: the automated way of collecting data from websites, comparable to a personal assistant that works around the clock. Web scraping helps business aggregators retrieve product information such as details, prices, and reviews much faster than the time it would take doing it manually. Boutiqaat is one of the many companies transforming the beauty and lifestyle market in the Middle East. With Arabic Fragrances, Niche Perfumes, and Scents from almost every corner of the world, the list of products seems endless. However, using automation for intelligence solves consumers' pervasive problem of being unable to sort and analyze information. For this project, we utilized up-to-date technologies to scrape data from the website Boutiqaat, focusing on three categories of interest: Arabic Fragrances, International Fragrances, and Niche Fragrances. Automatically collecting perfume listings and storing them in a tidy, logical format boosts customers’ and businesses’ awareness. Customers can browse through the myriad of options available, and retailers can utilize accurate information to make better business decisions. Whether you're a curious techie, a data analyst, or a brand strategist, this guide breaks down how we built a reliable, efficient scraping system to collect perfume data from Boutiqaat—step by step, and in a way that’s easy to understand. ## About Boutiqaat Boutiqaat is not simply an eCommerce platform—it is a breathtaking beauty and lifestyle destination which has transformed shopping behavior for the Middle Eastern consumers. Initially started in Kuwait, Boutiqaat integrates the luxury goods of regional traders with the tremendous influence of celebrities and beauty influencers from the region. Ahead of its competitors, Boutiqaat incorporates an extensive selection of makeup, skincare, and fragrances that appeal both regionally and from international labels. Unlike ordinary online shops, Boutiqaat accounts for trust and authenticity, which transforms shopping into an experience unlike any other. In this project, we aimed at obtaining product data from three primary categories of fragrances within the women’s section: Arabic Fragrances, Niche Perfumes, and an all-encompassing Fragrances category. This data has great significance in terms of consumer and business insight. Customers receive a better understanding of the available options, trends, and pricing while businesses are able to re-calibrate their offerings and marketing approaches based on the data. By shelving the data, we this way remove the gap between the offered supply and the actual market demand. ## Automated and Intelligent Data Collection We developed Boutiqaat’s data in two phases:collecting product page links and preparing them for further analysis. We focused on three major fragrance categories from the women’s section—Arabic Fragrances, International Fragrances, and Niche Perfumes. ### Step 1: Gathering Product Links We began with making an outline of all the product links available under each subheading for the fragrance categories. Like other modern shopping websites, Boutiqaat appears to have an endless scrolling feature or auto-loads products. This means that not all products on the page are listed at once. This requires a more specialized approach. Fortunately, we developed an intelligent scraper with Playwright which enables us to control web browsers and scroll down the page as a human would. No advanced masking techniques were employed; it was all about timing with Playwright. Our agents have been programmed to slowly scroll down the page, wait for new items to load, and harvest all relevant product links. We also taught it to deal with annoying pop-up screens that obscure vision—our agent checks for these and closes them if they come into view. This makes the scraping process more streamlined and efficient. We saved each product link and the category it belongs to in a SQLite database rather than dumping everything into a single file. This facilitates organization and makes it simple to retrieve data at a later time. To keep things tidy and effective, we even made sure the scraper doesn't save duplicate links. We were prepared for the following stage of our data journey by the time this phase ended, having gathered a substantial collection of clear, functional product URLs from each of the three fragrance sections. ### Step 2: Data Extraction from Product Links When the product URLS are stored in the SQLite database, the detailed information on each product is yet to be scraped. Pages of each product need to be visited individually to extract particular information about them. Current browsers available like Playwright help mitigate issues around being noticed during scraping as they portray themselves as humans. This ensures that the pages are properly rendered before gathering information. Tools, for instance, gather the following product information; the item name, its current and budget price, reviews count, description, brand name, discount offered, specification(s) such as size and SKU, availability for purchase, etc. The above information is stored in a mongodb database. This facilitates easier access and subsequent analysis of the data. With script reliability improvements, every product is accompanied with a JSONL file containing the information in a stripped manner. Freezing, crashing, or halting the scraping procedures mid-way does not result in lost data. In conclusion, numerous approaches all in all aid facilitate the effective collecting, documenting, and storage of richly detailed product data. ### Data Cleaning The URLs of the products have already been added to the SQLite database and now they need the information from the specific websites for each product. The Playwright script has been implemented because it mimics a user behavior for traversing the website and helps in avoiding detection ensuring all the pages are fully loaded. The tool collects critical information for each individual product which consists of the product’s name, brand, current price, old price, discount, review count, description, characteristics like size or SKU, stock availability, and others. All of this is now stored in a MongoDB database which makes it easier to handle and analyze data later. In order to make the data more reliable, the script also creates a local product backup in a JSONL file for every product. When the computer crashes without notification or the scraping process halts mid-way, this approach guarantees data redundancy. Clearly, the data collection and storage described above helps in reliable storage during scrapping and capturing extensive product datasets. ## Powerful Tools and Libraries for Smarter Data Extraction The data extraction procedure applies an ingenious combo of Python libraries and tools, all of which serve different purposes— from gathering the information and automating the browsers to storing them and dealing with errors. We will try to explain the above as clearly and simply as possible without losing professionalism. sqlite3:When dealing with data, SQLite is considered a lite weight and file based database when you're running it within Python, as there is no additional setup or server required. It works as an elegantly organized digital notebook that keeps record of urls and tells you which links have been processed. And considering it saves all the information in a well structured manner, it prevents you from working over and over again on the same tasks especially when the web scrapes thousands of pages. logging:If there are people who prefer the quiet Kyrgyz Republic, the logging module can be considered its politest citizen, the first being on the scene but staying quiet and neutral. It assists the script silently and does each and every task including scraping a link for data, saving data, hitting an error and even tracebacking an already stored error. Instead of using print() statements, hassle free logs are preserved in the project directory which allows clear trace back of what caused the error, and is quite valuable in case of overnight/scalable scraping sessions. playwright.sync\_api:Playwright serves as the motor driving automation within the Web. It performs actions like page openings, scrolling, clicking buttons and many more just like a human does. Modern websites that use JavaScript for rendering content demand such capabilities. In your case, the version that is easier to follow (synchronous called sync\_playwright) is employed, so that the code remains straight forward by executing one operation at a time. time: As mentioned Before, the time module offers basic functionalities, but is an integral part of abiding by scraping rules. While human emotions are expected, using time.sleep() to pause is a natural action among humans. This step further decreases the chances of getting flagged as a bot and creates the impression of actual browsing behavior. BeautifulSoup: BeautifulSoup is a library specifically designed for parsing sections of HTML to get the data wanted. Post a playwright’s performance in loading a webpage, BeautifulSoup is able to access the content and systematically pull out clear and structured text like product names, amounts, or descriptions. It simplifies the task of navigating complex webpage layouts. json:This module deals with the saving of data in a standard, small structure that is easily shared, backed up, or understood. In this case, it serves for the creation of .jsonl (JSON Lines) files where every single product is stored in its own line. This is ideal for backup and debugging, or simply sharing data with others. urllib.parse: URLs are sometimes overly long, messy, or random. The module urllib.parse takes care of cleaning duplicate URLs using a simplified method so that the same link is not stored twice. It organizes and rebuilds the URLs in a specific format. Pymongo: Now your scraper can be linked into MongoDB, a flexible NoSQL database, with pymongo. Unlike SQLite, which organizes information in rows and tables, MongoDB employs a document-information structure resembling the JSON format. This is suitable for product data since not all items contain the same fields. With all these tools combined, they make a very powerful advanced web scraping infrastructure for professional-grade responsible data extraction, cleaning, storing, and backing up. Each library has a unique purpose and together they reinforce the entire system and its flexibility. ## STEP 1: Scraping Product URLs from Women's Perfumes on Boutiqaat ### Importing Libraries ``` import sqlite3 import logging from playwright.sync_api import sync_playwright import time ``` The scraping adventure would commence from the women's section of Boutiqaat through combining a set of libraries that possessed distinct features geared toward the process’s efficiency, reliability, and scalability. First, its core, sqlite3 has a lightweight file based database system. It has mechanisms for tracking and storing product URLs, which from further manual processes would enable resuming from where scraping had left off. This eliminates repetitive tasks which is beneficial for larger projects. Also the logging module is very important for purposefully organizing records or logs of wheels—record accomplishments warnings successes and or errors—which simplifies debugging and progress tracking in long scraping sessions.Finally, control of a browser in real-time is executed using playwright.sync\_api from Playwright, permitting multi-step interactions with JavaScript laden websites such as Boutiqaat which do not render fully with static HTML. This is also very important when scraping modern e-commerce platforms. Together with adding simulation of natural browsing behavior to reduce chances of being flagged, the time module built into Python contributes through the addition of pauses in scripts where needed. All these stated libraries provide a cohesive environment with reduced risk and increased efficiency when extracting data from dynamic websites. ### Keeping Track of Progress with Logging ``` # Setup Logging logging.basicConfig( filename="boutiqaat_url_scraper.log", level=logging. INFO, format="%(asctime)s - %(levelname)s - %(message)s" ) """ Configures the basic logging system for the application. Logging Setup: ------------- The logging configuration creates a detailed record of the script's operation: - All logs are written to the file 'boutiqaat_url_scraper.log' - Only messages with INFO level priority or higher are recorded (INFO, WARNING, ERROR, CRITICAL) - Each log entry includes: * Timestamp: When the event occurred * Level: How important the message is (INFO, WARNING, ERROR, etc.) * Message: Description of what happened """ ``` To make the scraper efficient and problem traceable, logging is a significant part of this project. Python’s built-in logging module is used by the script to monitor the scraping procedure, logging detailed operational activities of the scraper. The log data is outputted to a file named boutiqaat\_url\_scraper.log which is created, if not available, automatically. The logging system is configured to log all messages files including but not limited to warnings, errors, and critical failures at or above the INFO level. Each log message is preceded with a timestamp marking and explicitly tagged with the event’s severity so that developers know when and where something happened and they can debug easily. This is particularly helpful when scraping hundreds of product pages because it gives insight into progress and knows exactly which pages failed or were skipped. Rather than having developers wonder where the scraper could have left off, the logs serve as a timeline of the script's progress—a virtual paper trail that helps debugging and monitoring become so much easier. In much the same way that databases monitor data, logging monitors the health and activity of your scraper. ### Database and Website Configuration ``` # Database Configuration DB_NAME = "/home/anusha/Desktop/DATAHUT/Boutiqaat/Data/boutiqaat_full_urls.db" BASE_URL = "https://www.boutiqaat.com/en-kw/" """ Database and Website Configuration: DB_NAME:  The full path to the SQLite database file where scraped URLs will be stored.  This database acts as persistent storage for all collected product links. BASE_URL:  The root domain of the Boutiqaat website. This is used to convert relative URLs  to absolute URLs when necessary.  """ ``` The configuration section of the script sets the foundation for storing and organizing the data extracted from the Boutiqaat Women’s section. The DB\_NAME variable points to the full file path of the SQLite database — a lightweight and file-based database system. In this case, the database is named boutiqaat\_full\_urls.db and is located in a folder named Data on the user's desktop. This database acts as a secure container for holding all the URLs of the women's perfume products scraped from the website. Next, the BASE\_URL is defined as "[https://www.boutiqaat.com/en-kw/](https://www.boutiqaat.com/en-kw/?ref=blog.datahut.co)", which serves as the main address of the website. It plays an essential role in combining partial links (called relative URLs) from the site with this base address to form complete product URLs (absolute URLs). This setup ensures that even if the site only gives part of a link, the script can convert it into a full, usable web address. Altogether, these settings are crucial for organizing and guiding the scraping process smoothly from the very beginning. ### Setting Up User-Agent Headers for Safe and Seamless Scraping ``` # User-Agent Headers HEADERS = {    "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36" } """ Browser Headers: HEADERS:  Contains the User-Agent string that identifies our web browser to the website.  This helps the script appear as a regular web browser rather than an automated tool.  The User-Agent mimics Chrome on Linux to ensure compatibility with the website.  Proper headers help avoid being blocked by the website's security measures. """ ``` When we scrape information from websites such as Boutiqaat's Women's page, we should ensure that our script does not appear suspicious or bot-like. That is where browser headers are useful — namely, the HEADERS dictionary in our script. This section contains a User-Agent, and this is simply a tiny bit of data that informs the website that we are using what type of browser. In this instance, the User-Agent is mimicking Google Chrome on a Linux machine, something that lots of actual users utilize. Adding this is assisting our scraping script to fit in and act like a typical person surfing the website. Without this, the site may see that a bot is making a visit and block or limit access to the pages. ### Building the Foundation: Creating a Database for Perfume Links ``` def create_db(): """ Initialize the SQLite database with a 'links' table. This function creates a new SQLite database (if it doesn't exist) with a table structured to store product URLs and their associated categories. It includes: - id: An auto-incrementing primary key - url: A unique text field to store product URLs - category: A text field to store the category of each product If the database or table already exists, this function will not modify them. Returns: None Raises: Exception: Logs any error that occurs during database creation """ try: conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS links ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, category TEXT ) """) conn.commit() conn.close() logging.info("Database initialized.") except Exception as e: logging.error(f"Database creation failed: {e}") ``` We should have a way to save links to the product page of each perfume safely before we start data collection on women's perfumes from the Boutiqaat website. All of this part of the script is for that purpose only. It creates a small but efficient database using SQLite, which is a Python library that allows you to create and use databases without having to download anything else. The role played here is create\_db(), and it's meant to specify the database structure. Inside it, we create a table named links, where each row will have one perfume product URL and its category (e.g., "Niche Perfumes" or "Arabic Fragrances"). Each record is assigned an automatically incremented id number, which makes each link distinct. Even when this function gets executed more than once, it will not foul up any old data — it first checks to determine whether or not the table already exists, and only then creates it if it doesn't. This both makes the code reusable and makes it safe. If there happens to be a problem during this operation (such as if the database connection isn't working correctly), the script will record what's gone wrong through the use of the logging system, which makes it straightforward to determine where things went awry. In brief, this section of code goes quietly about laying the groundwork for your scraping efforts — a spot to house and keep tabs on each perfume link prior to further data extraction. ### Saving Product URLs Without Duplicates ``` # Save URL to database def insert_link_to_db(link, category):    """    Insert a product URL and its category into the database.       This function attempts to add a new URL to the database. The UNIQUE constraint    on the 'url' field prevents duplicate entries. If a URL already exists in the    database, it will be skipped and logged accordingly.       Args:        link (str): The complete product URL to be stored        category (str): The category name associated with the product    """        try:        conn = sqlite3.connect(DB_NAME)        cursor = conn.cursor()        cursor.execute("INSERT OR IGNORE INTO links (url, category) VALUES (?, ?)", (link, category))        if cursor.rowcount == 0:            logging.info(f"Duplicate URL skipped: {link}")        else:            logging.info(f"Inserted URL: {link}")        conn.commit()        conn.close()        logging.info(f"Inserted URL: {link}")    except Exception as e:        logging.error(f"Error inserting link: {link} → {e}") ``` When web scraping women's perfume products on the Boutiqaat website, you should store each product URL in some sort of organized way so you'll be able to come back and scrape in-depth information later. This part of the code does that—it saves each perfume's URL along with its category ("Niche Perfume" or "Arabic Fragrance") into a database. This function is declared here: insert\_link\_to\_db(link, category). It's called to add perfume product URLs into a SQLite database, which is such a clean place to save all your links. Whenever the script finds a new product link, it calls this function to store it. But you don't need to save the same link more than once. To avoid duplications, the database rule is "only unique URLs allowed." So if the script attempts to add an existing URL, it will silently skip it and log a message in the log file indicating that it was a duplicate. If the link is new, it's saved and a successful message is logged. It also employs something referred to as logging to monitor what's going on—whether the link was saved successfully or skipped because it already existed. If something fails, such as a database connection error, it logs that as well, so it's easier to debug later. In short, this function ensures that all of the perfume product links are saved cleanly, without repetition, and any issues encountered along the way are logged for examination. This is an important step because it sets the foundation for the second stage of your project—extracting the real perfume data from each of these saved links. ### Smart Scrolling to Gather Perfume URLs from Boutiqaat ``` def fetch_page_content(url, category, max_scrolls=300):      """ Visit a category page and extract all product URLs by scrolling through the page. This function uses Playwright to automate a browser session that: 1. Navigates to the specified category URL. 2. Repeatedly scrolls down to trigger lazy loading of products. 3. Extracts product URLs from anchor tags and ensures they are unique. 4. Stores these unique URLs in the database. 5. Detects when no new content is loaded after scrolling and stops scrolling. It also includes mechanisms to: - Close any popup dialogs that might interfere with scraping. - Track URLs that have already been processed to avoid duplicates. - Stop scrolling if no new content is added after several attempts. Args:      url (str): The category page URL to scrape. This is the page that contains the product links.      category (str): The category of products being scraped (e.g., 'Niche Perfumes', 'Arabic Fragrances').      max_scrolls (int, optional): The maximum number of scroll operations to perform. Defaults to 300.                                    This ensures that the scraping process doesn't run indefinitely if there are a large number of products. Returns:      None: This function doesn't return any values. It performs scraping and data insertion directly. """    try:        with sync_playwright() as p:          """Launch a visible browser for debugging purposes            Set to 'False' for debugging, change to 'True' for headless mode"""            browser = p.chromium.launch(headless=False)            context = browser. new_context(user_agent=HEADERS["User-Agent"])             page = context. new_page()            logging.info(f"Navigating to {url}")            page.goto(url, timeout=60000)            page.wait_for_load_state("networkidle")            time.sleep(3)                 """Navigate to the page with a 60-second timeout                Wait until the page is fully loaded                Allow some time for the page to stabilize"""            try:                if page.is_visible("button.close"):                    page.click("button.close")                elif page.is_visible(".popup-close"):                    page.click(".popup-close")            except Exception:                pass             """Close popup if visible                  Check if there's a close button for a popup                 Check for another possible popup close button                  If no popup is found or couldn't be closed, continue"""            seen_urls = set()            same_scroll_count = 0            last_height = 0         """Set to track URLs we've already processed          Counter to track if new content is loaded after a scroll          Track the page's height to detect when it's not scrolling anymore"""            for scroll in range(max_scrolls):                page.evaluate("window.scrollTo(0, document.body.scrollHeight)")                time.sleep(3)                current_height = page.evaluate("document.body.scrollHeight")                anchors = page.query_selector_all("a[href*='/p/']")                logging.info(f"Scroll #{scroll+1}: {len(anchors)} links found")                   """ Scroll up to `max_scrolls` times                        Scroll to the bottom                        Wait for the page to load new content                        Get the new page height                        Find all product links by matching 'href' attribute                        Log the number of links found"""                for a in anchors:                    try:                        href = a.get_attribute("href")                        if href and href.endswith("/p/") and "/product/" not in href:                            full_url = BASE_URL + href if href.startswith("/") else href                            if full_url not in seen_urls:                                seen_urls.add(full_url)                                insert_link_to_db(full_url, category)                    except Exception as link_error:                        logging.warning(f"Failed to process link: {link_error}")                    """Loop through all found anchor tags (product links)                  Get the href attribute (URL)                      Filter out invalid links                        Complete the URL if it's relative                        Only process new URLs                        Add to seen URLs                        Insert the link into the database                  Log any issues with processing links"""                if current_height == last_height:                    same_scroll_count += 1                    if same_scroll_count >= 5:                        logging.info("No more new content, scrolling ended.")                        break                else:                    same_scroll_count = 0                    last_height = current_height            browser.close()    except Exception as e:        logging.error(f"Exception while scraping {url} → {e}")                       """Log any errors that occur during scraping""" ``` This method, fetch\_page\_content, is designed to navigate to a perfume category page on the Boutiqaat site (e.g., "Niche Perfumes" or "Arabic Fragrances") and automatically extract the URLs for individual perfume items. Since all the items are not loaded simultaneously, the page must be scrolled several times—just as a human user would—to cause new items to become visible. The primary intention is to scroll down and retrieve all the links of the products, recognize every product on the page, and store that link into a database where it could be analyzed subsequently. Controlling the Browser with Playwright: The function employs a library named Playwright, which is a robot that opens and operates a browser (Chrome/Firefox) like a human. It can scroll down, press buttons, and wait for pages to load. The browser is initiated in visible mode (headless=False) which is perfect for debugging — you actually get to see the script in action in real time as it clicks, scrolls, and fetches links. Only when satisfied everything runs perfectly fine, you can turn it "invisible" (headless) mode so you can execute it quicker and behind the scenes. Dealing with Popups and Page Loading: A few websites display popups that may obstruct the screen (such as offers or cookie alerts). The script tries to detect and close such popups automatically. It waits for the page to load fully, giving a couple of seconds to settle before it starts scrolling and scraping data. This prevents the scraper from being interrupted by unrelated distractions that could prevent it from accessing the perfume products. Smart Scrolling to Load More Products: The product pages use a technique called lazy loading, so all items are not shown at once. Instead, when you scroll, more items are shown. The function mimics this behavior by: - Scrolling to the bottom page - Waiting a few seconds for more content to be loaded - Checking how far it scrolled by comparing the page height - Repeating this cycle anywhere from 300 times (or until nothing new loads) If it realizes that no new content loads for five scrolls consecutively, it stops scrolling—because that likely indicates that everything has loaded. Gathering and Saving Product Links: Whenever new products load, the function searches for certain links (those with "/p/" in their URL, which typically indicates it's a product page). It ensures: - The link is correct and not a duplicate - It hasn't been saved before - It gets converted to a full URL (if necessary) - It gets stored in a database to be analyzed later A set named seen\_urls is utilized to keep track of which links have already been gathered, so there's no redundancy. Logging Every Event Which Happens: With every action, the aim logs itself. It logs if: - A page is navigated - Every scroll is finished and the number of links encountered - A link has been saved successfully or skipped successfully. - There has been an error. If something goes wrong (like the page not loading or the structure not being the same), the script doesn't fail silently; instead, it logs the exact error so that it can be easily fixed later. In Short This tool is a complete automation process that: - Traverses a perfume category on Boutiqaat - Scrolls to load all items available - Detects and closes popups - Collects and stores all perfume product links - Keeps everything tidy and logged It's built to run smoothly, prevent duplicates, and handle errors gracefully—making it a trustworthy tool for bulk data collection from an e-commerce website like Boutiqaat. ### Starting the Scraper: Organizing Perfume Data Collection from Boutiqaat ``` # Main function def main():  """ Main entry point of the scraper. This function performs the following tasks: 1. Calls `create_db()` to initialize or ensure the SQLite database is set up correctly. 2. Defines a dictionary of fragrance category URLs and their corresponding labels:      - "Niche Fragrances"      - "International Fragrances"      - "Arabic Fragrances" 3. Iterates over each URL-category pair:      - Logs the start of the scraping process for each category.      - Calls `fetch_page_content(url, category)` to scrape and process product data from that category. This setup helps organize and automate scraping for multiple product categories on the Boutiqaat website. """    create_db()    urls = {        "https://www.boutiqaat.com/en-kw/women/fragrances/niche-perfumes-1/l/": "Niche Fragrances",        "https://www.boutiqaat.com/en-kw/women/fragrances-1/c/": "International Fragrances",        "https://www.boutiqaat.com/en-kw/women/arabic-fragrances-1/c/": "Arabic Fragrances"    }    for url, category in urls.items():        logging.info(f"Starting category: {category}")        fetch_page_content(url, category) ``` The main() function is such that it is the beginning or control center of your web scraping script. It's the point where the whole process of data collection starts and gets coordinated in a structured manner. Step-by-Step Breakdown 1.Creating a Database This function first calls create\_db(). This guarantees that the perfume database you intend to create is prepared. A database is created if one doesn't already exist. If it is present, it guarantees that everything is set up properly. This keeps your data organized and accessible in the future. 2\. Fragrance Categories & Their URLs Then, a dictionary called urls is created. This dictionary contains the top categories of women's perfumes you wish to scrape from the Boutiqaat website. It contains: - Niche Fragrances (more high-end or luxury fragrance) - International Fragrances (common international brands) - Arabic Fragrances (perfumes mimicking typical Arabian fragrances) 3\. Each category has a particular web address (URL) that takes one to the products under it. 4\. Looping Through Each Category Then, the method goes through every category individually. For each category of perfume: - A message is logged (stored) using the logging tool. Such a message is, for example, "Starting category: Arabic Fragrances," useful to follow progress or debug in case something goes wrong. - It then calls a function called fetch\_page\_content(url, category). This function is responsible for visiting the webpage of the category, loading all the products, and fetching the data. ### Entry Point to Execute the Script ``` if name == "__main__":    main()  """    Entry point when run as script:    - Executes main scraping function    - Handles any top-level errors    """ ``` Here, main() is the central function which initiates the whole process of scraping the perfumes from the Boutiqaat site. Hence, when you execute this script — e.g., by simply double-clicking the Python script or running it through the terminal — the following line ensures that the main() function initiates and everything goes into action: from fetching product links, opening perfume pages, scraping the information, and writing them out. If this file were imported elsewhere (e.g. in some other program), Python would not automatically execute main(). This is helpful when your script may be reused in a larger project. ## STEP 2: Extracting Complete Product Information from Each Link ### Importing Libraries ``` import sqlite3 import logging import json from pymongo import MongoClient, errors from bs4 import BeautifulSoup from playwright.sync_api import sync_playwright from urllib.parse import urlparse, urlunparse ``` This set of libraries collaborates to enable web scraping from Boutiqaat's women section effortless and trustworthy. The playwright.sync\_api library is utilized to automate the browser in loading product pages just as a human being would—clicking, scrolling, and waiting for content to render—most important for those heavily dependent on JavaScript. BeautifulSoup parses the downloaded page to retrieve clean data, like names, prices, and descriptions of scents. sqlite3 keeps an in-memory local database with product URLs and whether they have been processed to avoid scraping the same product more than once. pymongo allows storing information in MongoDB, a light-weight NoSQL database well-suited to dealing with large and complex data sets. logging keeps a record of everything that's going on in the background, allowing developers to see where problems are and how far along the scraper has gotten. Lastly, urllib.parse cleans up and normalizes URLs before they're used. Overall, these are an easy-to-use yet powerful way to extract and clean product data. ### Key Configuration Setup for Structured Data Collection ``` # Constants BASE_URL = "https://www.boutiqaat.com"  DB_PATH = "/home/anusha/Desktop/DATAHUT/Boutiqaat/boutiqaat_full_urls.db"   BACKUP_FILE = "boutiqaat_data_backup_.jsonl"  MONGO_URI = "mongodb://localhost:27017/" MONGO_DB = "boutiqaat"   MONGO_COLLECTION = "products"  """ Application configuration settings: -BASE_URL: The root URL of the website being scraped - DB_PATH: Location of SQLite database file              - MONGO_URI: MongoDB connection string - MONGO_COLLECTION: This sets the table name to "products" where all scraped items will be saved              - MONGO_DB: MongoDB database name  - BACKUP_FILE: Backup location for MongoDB data """ ``` This section of the code determines the underlying settings for where and how the data scraped from the Boutiqaat Women's Perfume category will be stored and processed. These constants are like a map—nominally referring to the tools and files that the scraper will be using in the process. BASE\_URL specifies the top-level website address where the scraper starts gathering information, in this case, women's perfumes on Boutiqaat. The DB\_PATH instructs the script where to store the gathered product links in the form of SQLite, a light database that is stored in a local file. BACKUP\_FILE defines where a copy of the extracted data will be saved in a .jsonl format—this acts as a safety net in case anything goes wrong. For long-term storage and scalability, data is also saved in a MongoDB collection, which is set up using MONGO\_URI, MONGO\_DB, and MONGO\_COLLECTION. These settings help the scraper stay organized, efficient, and fail-safe by ensuring data is always backed up and easily accessible for further analysis or review. ### Tracking Progress with Logging for Reliable Scraping ``` # Setup Logging logging.basicConfig(    filename="boutiqaat_data_scraper.log",     level=logging. INFO,      format="%(asctime)s - %(levelname)s - %(message)s"  ) """ Configures the logging system for the Boutiqaat data scraper. This setup configures logging to output messages to a file named `boutiqaat_data_scraper.log`. The log messages will include: - The timestamp when the log entry was created. - The log level (e.g., INFO, WARNING, ERROR). - The actual log message. This is essential for tracking the script's progress, debugging issues, and auditing scraping activities. Parameters: ----------- - filename (str): The name of the log file where log entries are saved. In this case, `boutiqaat_data_scraper.log`. - level (int): The logging level that determines which log messages are captured. `logging. INFO` captures informational messages and above (INFO, WARNING, ERROR, CRITICAL). - format (str): The format for log messages. It includes the timestamp (`asctime`), log level (`levelname`), and the log message itself (`message`). Logging Levels: --------------- - `INFO`: Captures general information about the scraping process (e.g., progress updates, data extraction successes). - `WARNING`: Used for warnings (e.g., when a product is skipped due to missing data). - `ERROR`: Used for error messages (e.g., when an exception occurs during scraping or data extraction). """ ``` This section of the script sets up a logging system to track everything that happens when scraping. Instead of flooding your console with messages, logs are neatly written to a file called boutiqaat\_data\_scraper.log. This log is akin to a diary for your scraper—it keeps track of important actions, successes, warnings, and even errors that occur when scraping data from the Boutiqaat site. The logging mode is designed to show the time when something happens exactly, the level of the message (e.g., INFO for regular updates or ERROR if there is a break), and a simple description of what is happening. Having the logging level as INFO is saying that all the significant updates, warnings, and errors are being recorded without drowning in too much technical chatter. This simplifies debugging and also gives a clean history of the scraping session, which is particularly helpful when dealing with large amounts of data or running automated scrapers for extended periods. It's a tiny setup with a huge effect on having control and visibility over your scraping process. ### Seamless Data Storage with MongoDB Integration ``` # MongoDB connection # Creates a connection to the MongoDB database for storing scraped product data mongo_client = MongoClient(MONGO_URI)  # Connect to MongoDB server mongo_collection = mongo_client[MONGO_DB][MONGO_COLLECTION]  # Access specific collection ``` This section of the script establishes a connection to MongoDB, a dynamic NoSQL database for storing all the product information scraped from the Boutiqaat website. While normal databases store data in rigid tables, MongoDB enables us to store information in a JSON-like format, which is ideal for dealing with web data that is often inconsistent in structure. The script then starts by getting the connection to the MongoDB server via MongoClient and the previously defined MONGO\_URI. Upon connection, it is accessing the respective database and collection where the product information is going to be stored. In this way, the entire perfume information gathered—brand names, prices, and descriptions—is maintained in an organized and accessible format. Using MongoDB is easier to handle a lot of data in an effective way, especially when handling modern web content that is not necessarily in a standard format. ### Cleaning Up URLs for Consistent Data Storage ``` def clean_url(url):       """    Cleans the input URL by removing duplicate segments in the URL path.    This is helpful in ensuring uniformity and avoiding multiple entries for    the same product due to slight differences in the URL structure.    Args:        url (str): The original product page URL.    Returns:        str: A sanitized URL with duplicate segments removed.    """    parsed = urlparse(url)    path_parts = []    seen = set()    for part in parsed.path.split("/"):        if part and part not in seen:            path_parts.append(part)            seen.add(part)    cleaned_path = "/" + "/".join(path_parts)    return urlunparse((parsed.scheme, parsed.netloc, cleaned_path, '', '', '')) ``` This section of the script is intended to normalize product URLs using a function called clean\_url. When scraping product pages in Boutiqaat, the same product might be listed with slightly varying URLs due to redundant or unnecessary elements in the web address. These small differences can cause duplicate records in the database if not cleaned, i.e., dirty or inconsistent data. To avoid this, the clean\_url function employs Python's urlparse and urlunparse functions to parse the URL into its various components. It then reads through the path component of the URL, eliminating any duplicate segments but maintaining their order. The end result is a clean, streamlined version of the original URL, such that every product ends up with only one unique, consistent link stored in the database. It is a small step, but it is all about a big leap in the quality and integrity of the scraped data, particularly when you are scraping product links in hundreds or thousands. ### Setting Up the SQLite Database to Manage Product URLs ``` def setup_database():    """    Initializes and sets up the SQLite database to store product URLs and    their processing status.    - Creates a `links` table if it doesn't already exist.    - The table contains:        - id: Primary key        - url: The product URL (must be unique)        - category: The category to which the product belongs        - processed: An integer flag (0 or 1) to track if data has been scraped    - Adds the `processed` column if it's missing.    This database helps in managing progress across multiple runs.    """    conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    cursor.execute("""        CREATE TABLE IF NOT EXISTS links (            id INTEGER PRIMARY KEY AUTOINCREMENT,            url TEXT UNIQUE,            category TEXT,            processed INTEGER DEFAULT 0        )    """)    try:        cursor.execute("ALTER TABLE links ADD COLUMN processed INTEGER DEFAULT 0")    except sqlite3.OperationalError:        pass    conn.commit()    conn.close() ``` The function setup\_database is very crucial in data organization and maintenance of the scrape process. When we scrape product pages, we would need something that will be able to remember the URLs we have already processed and save other information like the category of the product. This function precisely does this by initializing a database based on SQLite, a minimalist and lightweight database system. This is the way the function operates: it opens an SQLite database, storing the data through a path specified by DB\_PATH. After being opened, the script creates a table named links if the table does not already exist. This is the table that we utilize to store the URLs for the products that we scrape and some other data such as the category for the product and a flag known as processed. The processed flag is initialized to 0, indicating the product has not yet been scraped for data. After creating the table, the function looks for the processed column if it does not exist and tries to add it if needed, ensuring the table remains updated at all times. The purpose of using a database like SQLite is to keep track of what URLs are yet to be scraped and what URLs have been scraped. It makes the scraping more efficient, particularly when you need to execute the script several times or when you are scraping a huge number of product pages. The function ensures that there will be no data loss and you can track the progress of your scraping by keeping product URLs stored and organized. This setup is done for the ease of suspending and resuming the scraping operation at will without losing track of which URLs need to be scraped or re-scraped again. It makes the process both efficient and correct throughout the entire web scraping procedure. ### Fetching Unprocessed Product Links for Scraping ``` def fetch_unprocessed_links():    """    Fetches all product URLs from the database that have not yet been processed.    Returns:        list of tuples: Each tuple contains:            - id (int): ID of the database row            - url (str): Product page URL            - category (str): Product category    """    conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    cursor.execute("SELECT id, url, category FROM links WHERE processed = 0")    results = cursor.fetchall()    conn.close()    return results ``` The fetch\_unprocessed\_links function is used to load all the unprocessed product URLs. This is an important step in a prolonged web scraping process over the course of multiple runs or sessions because it will only process unprocessed URLs and will not repeat processing, instead prioritizing new, unvisited URLs. This is done the following way: the method initially establishes connection with the SQLite database in which all the URLs of products are stored. Having established connection, it performs a SQL query that retrieves all URLs stored in the database whose processed flag is set to 0, i.e., such URLs are yet to be scraped. The SQL query also retrieves the id of the record (to keep track of the record) and the category under which a product belongs. This is in order to get control over processing of scraping by type of the product that will be scraped. Once the data is fetched, the function stores it as a list of tuples containing id, url, and the category of the item. This list of unprocessed links is then passed on to the script, which can now use the data to scrape the corresponding product pages for further information. The importance of this function is that it allows the scraper to pick up where it stopped. In the event that the scraping process had been cut short or in the event that it is to be resumed, the function allows the previously scraped URLs not to be scraped again, hence conserving time and resources. Essentially, it assists in the effective handling of the process of scraping by only operating on the yet-to-be-scraped URLs. This method makes the whole web scraping process more structured and effective. ### Marking Product Links as Processed to Avoid Re-Scraping ``` def mark_as_processed(link_id):    """    Marks a product link as processed in the SQLite database, so it won’t be scraped again.    Args:        link_id (int): The ID of the link to be marked as processed.    """    conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    cursor.execute("UPDATE links SET processed = 1 WHERE id = ?", (link_id,))    conn.commit()    conn.close() ``` The mark\_as\_processed is an important component of the process as it makes sure that all product links are scraped only once. Once a link has been scraped and data is received, the function makes sure that progress is made by marking the database as the link has been processed. This is how it works: when you call the function, it takes a single link\_id as an argument. This link\_id is the individual product link just processed. The function establishes a connection to the SQLite database where all the product URLs are located. It then employs an SQL query to mark the processed field of the specific link\_id as 1, indicating that the link has been processed. This mark informs the scraper that this link has been scraped and will not be scraped again. Once the update is complete, the modifications are committed to the database, and the connection is closed. It is a tidy but effective mechanism that maintains the scraper clean and only unprocessed links are scraped when the scraper is executed the next session. With this mechanism, the script can maintain progress, avoid re-scraping already processed data, and enhance performance overall, especially with lengthy scraping sessions or when the scraper is executed repeatedly. ### Scraping Product Pages Using Playwright for JavaScript Rendering ``` def scrape_html(url):    """    Uses Playwright to render a JavaScript-powered product page and retrieve its HTML content.    Args:        url (str): The product page URL to load.    Returns:        str or None: The full HTML content of the page if successful, otherwise None.    Logs:        Errors in loading or rendering the page are logged for later review.    """    try:        with sync_playwright() as p:            browser = p.chromium.launch(headless=True)            page = browser. new_page()            page.goto(url, timeout=60000)            page.wait_for_selector("h1.product-name-h1", timeout=10000)            html = page.content()            browser.close()            return html    except Exception as e:        logging.error(f"Failed to load page {url}: {e}")        return None ``` The scrape\_html function would be applied on JavaScript-heavy product pages which load their content via JavaScript. Web scraping libraries would be unable to scrape such pages since the information we require may not be easily accessible in the HTML source code. Playwright, an automation library for browsers with high performance, bridges this gap and helps us simulate a real user web browsing experience and load JavaScript-heavy pages. The process starts by initializing the Playwright browser in headless mode, i.e., without displaying the browser window. The product page URL is provided to the process, and Playwright loads the page. In order to finish loading the page and wait until the page is ready to scrape, it waits until some specific element, say the product name (h1.product-name-h1), shows up on the page. This ensures that the page is finished rendering and all of the content of interest has loaded. When the page is ready, Playwright goes there. Once the page is loaded completely, the function retrieves the page's HTML by using page.content(). This is the raw page structure with all the product information that would be parsed subsequently. The browser instance is then closed, freeing up resources. If there is any form of error, e.g., timeout or page load error, the function raises the exception and prints an error. This error message records what URLs could have led to the error and makes it easier to debug later. This web scraping method is effective when sites load information dynamically using JavaScript, such as product descriptions, images, and prices, and therefore is an unrivaled approach to web scraping information from sophisticated online shopping sites. ### Extracting Structured Product Data from HTML ``` def extract_product_data(html, url, category):    """    Extracts structured product details from the HTML of a Boutiqaat product page.    It parses the HTML using BeautifulSoup and safely retrieves:        - Product name        - Brand name        - Current price        - Old price (if available)        - Discount percentage        - Review count        - Full product description        - Specifications        - Availability status (in stock / notify me)    Args:        html (str): The rendered HTML content of the product page.        url (str): The original URL of the product page.        category (str): The product's category (used for tagging).    Returns:        dict: A dictionary with all extracted and cleaned product information.    """    soup = BeautifulSoup(html, 'html.parser')    def get_price_with_kwd(selector):        tag = soup.select_one(selector)        return tag.get_text(strip=True) if tag else "N/A"    def safe_select_text(selector):        tag = soup.select_one(selector)        return tag.get_text(strip=True) if tag else "N/A"    def extract_description():        description_tag = soup.select_one("div.content-color")        if description_tag:            paragraphs = description_tag.find_all("p")            return "\n".join(p.get_text(separator=" ", strip=True) for p in paragraphs if p.get_text(strip=True))        return "N/A"    availability_raw = safe_select_text("div.pro-details-add-to-cart a")    if "Buy Now" in availability_raw:        availability = "yes"    elif "Notify Me" in availability_raw:        availability = "no"    else:        availability = "unknown"    return {        "url": url,        "category": category,        "product_name": safe_select_text("h1.product-name-h1"),        "brand_name": safe_select_text("a.brand-title strong"),        "price": get_price_with_kwd("div.pro-details-price.discount span. new-price"),        "old_price": get_price_with_kwd("div.pro-details-price.discount span.old-price"),        "discount_percentage": safe_select_text("div.pro-details-price.discount span.discount-price"),        "review_count": safe_select_text("div.product-review-order span"),        "description": extract_description(),        "specifications": safe_select_text("li.heading-tag-sku-h1 span.attr-level-val"),        "availability": availability    } ``` The extract\_product\_data function is built to extract key details about a product from the HTML content of a product page on Boutiqaat. Once the page is loaded and its HTML content is retrieved, this function takes over to parse and organize the information in a structured format. The function begins by using BeautifulSoup, a Python library designed to parse HTML. It safely navigates the HTML content, looking for specific elements that contain important product details such as the product's name, brand, price, description, availability, and more. To extract the prices, the function looks for elements containing the current price, old price (if available), and discount percentage. It uses a helper function called get\_price\_with\_kwd, which retrieves the price text from the HTML and handles cases where the price might not be available, returning "N/A" in such cases. Next, the safe\_select\_text helper function is used to safely extract text from various HTML elements, such as the product name, brand name, and review count. If a certain element is not found, the function ensures that it doesn't break the program by returning a default value ("N/A"). The product description is handled separately. The function looks for a specific section of the page that holds the detailed description. If found, it extracts all paragraphs and joins them together into a single text block, ensuring the description is properly formatted. The function also checks the product's availability by looking for indicators like "Buy Now" or "Notify Me" in the HTML. This helps determine if the product is in stock or out of stock, providing a simple "yes," "no," or "unknown" status. Finally, the function returns all this collected data as a dictionary, which includes: - Product name - Brand name - Current price - Old price (if available) - Discount percentage - Review count - Full product description - Specifications - Availability status By organizing the extracted data into a dictionary, this function ensures that all product details are cleanly structured, making it easier to analyze, store, and later use in any application or database. This approach is crucial for handling large amounts of product data efficiently, especially in e-commerce scraping. ### Storing Data Efficiently with JSON Lines (JSONL) ``` def append_to_jsonl(data, file_path):    """    Appends a single dictionary entry to a JSON Lines (JSONL) file.    JSONL format stores one JSON object per line and is efficient for streaming or incremental backups.    Args:        data (dict): The data dictionary to append.        file_path (str): The full path to the backup file.    Logs:        - Success messages when data is appended.        - Errors if the file cannot be written to.    """    try:        with open(file_path, "a", encoding="utf-8") as f:            f.write(json.dumps(data, ensure_ascii=False) + "\n")        logging.info(" Backup entry written to JSONL.")    except Exception as e:        logging.error(f" Failed to write to JSONL backup: {e}") ``` The append\_to\_jsonl function is responsible for saving the scraped product data to a local file in a reliable and organized manner. Instead of saving all data at once or using a bulky format, it takes a more efficient route—appending one entry at a time to a JSON Lines (JSONL) file. JSONL is a simple and powerful file format where each line is a standalone JSON object. This makes it perfect for logging and storing large amounts of structured data incrementally. For instance, if your scraper collects thousands of product entries, you don’t have to keep all of them in memory or risk losing everything if the script crashes. Instead, each product's data is written line-by-line, one at a time. The function receives two inputs: - data: A Python dictionary containing the product's details (like name, price, availability, etc.). - file\_path: The full path where the JSONL file will be saved or updated. Inside the function, it opens the file in append mode using open(...), which ensures the file is closed properly after writing. The json.dumps() method converts the dictionary into a JSON-formatted string, and ensure\_ascii=False ensures that non-English characters (like Arabic or accented letters) are preserved. Each JSON object is then written to the file with a newline character, making it easy to parse later. To maintain reliability and traceability, the function logs messages throughout the process. If the entry is successfully written, a confirmation is logged. If something goes wrong (like the file being locked or missing permissions), it catches the exception and logs an error message with a clear description. This design not only keeps the data backed up incrementally but also allows the scraper to resume seamlessly in case of an interruption—an essential feature when working with large e-commerce websites like Boutiqaat. JSONL files are also compatible with many data processing tools, making them a great choice for long-term data storage and analysis. ### The Heart of the Scraper: Coordinating the Workflow with main() ``` def main():    """    Main entry point for running the Boutiqaat product scraper.    This function Coordinates the entire scraping process, from setting up the database to processing product links    and saving the scraped data to MongoDB and a local JSONL file.    Workflow:    ---------    1. Database Setup: Initializes the SQLite database by setting up the required table and columns if they do not already exist.    2. Fetch Unprocessed Links: Retrieves all unprocessed product links from the SQLite database. This ensures that only new products are scraped.    3. Scraping Loop: For each unprocessed product link:        - Cleans the URL to standardize it and remove duplicates.        - Loads the product page using Playwright, waits for the page to fully render, and retrieves the HTML content.        - Extracts product details (such as name, price, brand, description, availability) using BeautifulSoup.        - Saves the extracted data to a local JSONL backup file.        - Attempts to insert the data into a MongoDB collection for persistence.        - Marks the URL as processed in the SQLite database to avoid scraping the same link again in the future.    4. Logging: Throughout the process, logs are generated to keep track of progress, successes, and errors:        - Successes are logged when a product is processed and saved successfully.        - Warnings are logged for duplicates in MongoDB or missing HTML content.        - Errors are logged for any failures during scraping or data extraction.           The function ensures that the scraper runs smoothly, continues processing links, and avoids duplicate scraping.    This function is executed when the script is run directly, and serves as the main driver for the entire scraping workflow.    Error Handling:    ---------------    - If the page cannot be loaded (e.g., due to timeouts or missing elements), a warning is logged and the link is skipped.    - If data extraction fails, the error is logged and the scraper moves on to the next link.    - Duplicate entries in MongoDB are detected, and a warning is logged without attempting to insert them again.    """       setup_database()    unprocessed_links = fetch_unprocessed_links()    logging.info(f"Found {len(unprocessed_links)} unprocessed product links.")    for link_id, url, category in unprocessed_links:        logging.info(f"Processing: {url}")        url = clean_url(url)        html = scrape_html(url)        if not html:            logging.warning(f"Skipping (no HTML): {url}")            continue        try:            extracted = extract_product_data(html, url, category)            # Save to JSONL backup (no id)            append_to_jsonl(extracted, BACKUP_FILE)            # Insert into MongoDB (no id)            try:                mongo_collection.insert_one(extracted)            except errors.DuplicateKeyError:                logging.warning(f" Duplicate MongoDB entry for URL: {url}, skipping insert.")            mark_as_processed(link_id)            logging.info(f" Success: {url}")        except Exception as e:            logging.error(f" Error processing {url}: {e}")            continue ``` At the center of this entire scraping project lies the main() function. This function acts as the command center—responsible for coordinating every moving part of the scraper, from setting up the database to saving product data into both local files and a cloud database. Let’s break down what it does in an intuitive, step-by-step way. 1\. Setting Up the Environment The first step consists of calling the setup\_database() function. From there, the script takes care of setting up the SQLite database in case it isn’t ready yet. During the first run of the script, the required table (links) is created. If it is already created, it just continues. This allows dealing with hundreds or thousands of product pages since the links which are processed already, are tracked. 2\. Obtaining Product Links That Have Not Been Processed Proceeding with the same logic, the script fetches unprocessed links using fetch\_unprocessed\_links(). As the name suggests, this filtering ensures only new or missed links are scraped, avoiding unnecessary redundancy. Each record also contains the product ID, corresponding link, and category (for example, “Colognes” or “Niche Perfumes”). 3\. Looping Through Each Product Page The main logic is embedded in a for loop that goes through each unprocessed link. This is how each part works within the loop: - URL Cleaning: Before doing anything, the URL is cleaned using clean\_url() to ensure consistency and eliminate potential duplicates due to URL formatting issues. - HTML Rendering with Playwright: The script loads the complete webpage content by using the scrape\_html() function. Boutiqaat's pages are JavaScript-dependent, and hence having a tool like Playwright becomes imperative to get the entire HTML after the rendering of the page. - Skip if HTML is Missing: If for some reason the page does not load or display correctly (e.g., due to slow internet or server issues), the script logs a warning and skips over such a link. This maintains efficiency and prevents unnecessary crashes. 4\. Extract and Store Product Data When the HTML document is retrieved, the function calls extract\_product\_data(), which retrieves the product name, price, brand, availability status, and description. This function uses BeautifulSoup to process the HTML and extract pertinent information. - Local Backup with JSONL: Whether MongoDB is reachable or not, the append\_to\_jsonl() method is always called which saves the information into a .jsonl file. This local backup ensures that no data is lost even when there is a connectivity problem. - Storing in MongoDB: Subsequent to Backup, the information is also added to a collection in MongoDB. MongoDB serves as the main database for maintaining a catalog of structured product information. When using the script for the first time, it attempts to insert a document with a unique primary key for each product. If the product already exists, which is checked using a duplicate key, then the script logs a warning and does not include it again—no crashes, no clutter. 5\. Mark As Processed URLS After performing all the necessary tasks for a product, mark\_as\_processed() is called to set the new status in the SQLite database in case the product has been processed. It is worth noting that this action does not seem very significant at first sight – however, it does guarantee that the same product is scraped only once in the life of the program unless its flag is forcefully changed. 6\. Logging for Transparency Throughout the operation, precise log messages are created. They consist of successful scrapes, skipped links, duplicate detection, and any unanticipated errors. Such live feedback is invaluable when you're debugging or watching over lengthy scraping sessions. ### Entry Point to Execute the Script ``` # Entry point of the script if name == "__main__":    """    Entry point when run as script:    - Executes main scraping function    - Handles any top-level errors    """    main() ``` The final piece in our Boutiqaat scraper is the entry point, invoked when the script is run in direct mode. This is treated in a dedicated Python clause: if name == "\_\_main\_\_": While this may seem a bit enigmatic to beginners, it takes a long way in keeping your code modular and clean. In essence, this block guarantees that the scraper only executes when the file is run independently — not when it's imported as a module into another program. Within this block, we invoke the main() function, which serves as the core of the whole scraping process. By doing so, we initiate everything in order: from database setup and product page scraping, to data saving and progress logging. This is a Python best practice because it lets all other functions — such as HTML scraping or data extraction — be reused or tested independently without running the complete scraping process automatically. It's a neat yet effective way to keep your script tidy, secure, and production-ready. ## Conclusion In a world where thousands of products are listed online, doing everything by hand just isn’t practical. That’s why we built a smart and simple tool that automatically scrolls through Boutiqaat’s perfume pages, finds every product, and saves their links for us. It works just like a human browsing the site—scrolling down, closing popups, and picking out the right links—but does it faster, without getting tired. This automation helps us collect a full list of perfumes in one go, making it easier to study trends, compare prices, or build cool features like search tools or dashboards. This is the equivalent of having a digital assistant that can manage boring tasks which allow us to direct our energy and time towards analyzing new data and obtaining actionable insights. Whether you are a beginner in coding and just want to know how online data collection works, this project serves as a perfect example of effective and impactful automation. ## Libraries and Versions Name: pymongo Version: 4.10.1 Name: playwright Version: 1.48.0 Name: beautifulsoup4 Version: 4.13.3 ## FAQ SECTION 1\. Is it legal to scrape data from Boutiqaat? Web scraping publicly accessible data from Boutiqaat may be legally permissible for personal or research purposes. However, it’s crucial to review their [Terms of Service](https://www.boutiqaat.com/?ref=blog.datahut.co) and ensure compliance with ethical scraping practices, such as respecting robots.txt. 2\. What kind of perfume-related data can I scrape from Boutiqaat? You can extract product names, prices, discounts, brand information, bottle sizes, customer ratings, reviews, fragrance notes, gender segmentation, and availability status to analyze perfume market trends. 3\. Which tools or technologies are best for scraping Boutiqaat? Python libraries like BeautifulSoup, Scrapy, or Selenium work well. For scalable or automated scraping, you may consider using Playwright, Puppeteer, or a no-code tool like n8n. 4\. How can scraped data from Boutiqaat help in market analysis? The data can uncover pricing trends, top-selling fragrances, consumer preferences by brand or scent type, gender-based product distribution, and identify emerging perfume brands in the GCC market. 5\. What are the risks of scraping Boutiqaat manually? Manual scraping can trigger anti-bot protections, lead to IP bans, or result in outdated insights due to slow updates. Automated tools or scraping services can reduce these risks and provide cleaner, timely data. ### Tata CLiQ Personal Care Data Scraping for Market Insights URL: https://www.blog.datahut.co/post/how-to-scrape-tata-cliq-for-reliable-personal-care-product-insights/ Last updated: 2026-07-23T07:48:32.000Z Let's say you had a really smart assistant who could go into a website and collect everything important for you; for example, products, prices, sales, and you'd never have to copy-and-paste anything. This is basically the equivalent of web scraping! It's like you have a robot that can go through a website and pick out the parts you care about in a timely manner. It's especially functional for retail as businesses can monitor changing prices, promotions, best-selling items, etc., and then put themselves in a better position to make good decisions. In our case as shoppers, it gives us the ability to find the best deals quickly for our preferred brands. When done right, web scraping does not harm the website or contravene any rules; it just works in the background without any obfuscation. ## About Tata CLiQ Fashion Tata CLiQ Fashion isn't just a shopping website -- it's as good as a friend when it comes to your online shopping needs for clothing, beauty, and lifestyle products. Tata CLiQ Fashion was started in 2016 by Tata group, which is pretty much the world's most trusted company (Tata group owned Tanishq and Tata motors, which will be appearing later on). What makes Tata CLiQ Fashion special is that it doesn't sell fake and cheap rubbish online -- Tata CLiQ Fashion only sells products that they acquire from more than 6,000 branded shops, so you know everything you purchase is not fake. Whether you are in the market for clothes, accessories, skincare, or stuff for your home, there is plenty of choice. This project looked at the "personal care" part of the total items that they sell -- the lotions, face washes, etc. The aim is to save shoppers time by showing what shoppers are happy to spend their money on, and also provide business with insight into how shoppers spend their money and what's happening with pricing considerations. In the end, if both parties on either side of the commerce will understand each other, it will help make the online shopping process smooth and enjoyable for both sides. ## Effortless Data Extraction Our web scraping process contains two primary steps. URL Collection Phase and Data Extraction from Product Links. ### URL Collection Phase We began by scraping all the product URLs from the Personal Care category from the Tata CLiQ website. Tata CLiQ's website doesn't load all the products on the screen at once; it loads them as the user scrolls. So we used an open-source tool called Playwright Stealth that makes our script behave like a real user browsing the web to bypass the website's attempts to block us. As our tool was browsing through the category pages, it captured all of the product URLs as it was scrolling the screen. Then we encountered a problem: a screen popup asking users to subscribe was constantly blocking everything on the screen. So we added code to have the scraper recognize and close the popup before completing any action. Also, instead of a "next page" button, Tata CLiQ has a "Show More Products" button that loads more products. We took precautions to ensure our tool clicked that button every time we were collecting links so that we would not miss any products. ### Data Extraction from Product Links After saving all the product links to a simple database (named SQLite), the script visits each link, one by one, to gather more information — kind of like opening each product page in a browser to check its name, price, description, size, color, and if its in stock. It uses a tool called Playwright which allows the script to act like a real human while scrolling and clicking, so websites do not block it. Instead of saving the information to a file and uploading it, the information goes directly into another database, called MongoDB — this way, it keeps things faster and seamless. It also saves a backup as CSV files (similar to saving an Excel sheet) on the computer just to be safe, so nothing gets lost. This process is simple, straightforward, and keeps it under the radar while still obtaining essential information. ### Data cleaning Once you are ready to use your collected data, the next step (and very important step) is cleaning the data so that it makes sense and can be put to use. You can compare this to a pile of messy notes - they need to be sorted and cleaned up before you can use them for studying. Raw data typically contains a great deal of "mess" in the form of duplicate data, formatting issues, and other issues that don't seem to match up. If you haven't figured out your "messy data" situation, one great tool for this is OpenRefine. It is sort of a smart editing app for your data - you can identify and eliminate duplicates, bring standardization to attributes, and adjust any inconsistencies, all with a few clicks. Working with OpenRefine can be like organizing a messy spreadsheet while looking under a microscope! If you have more complex problems - such as unwanted HTML tags left in your text, or turning a column of numbers into a proper date - Python, with its library called pandas, will save the day. Pandas gives you the functionality to make these intricate edits, similar to having an intelligent assistant managing the background to sort your data. While it might not seem like the most exciting step in the process, cleaning your data will ultimately be a compulsory step if you want your final results to be of value! ## Powerful Tools and Libraries for Smarter Data Extraction The relevant code undertakes usage of several fundamental Python libraries that are distinct from each other in the sense that they make web scraping effective, effortless, and undetectable. I will highlight each of them in an easy to understand professional manner so that even those who are not well versed with the subject will find it easy to comprehend. Asyncio: The first in the list is asyncio, which is a built-in track of a function in Python – simply put, it is multitasking. It acts as a ‘smart traffic controller’ that makes sure that various parts of your code run smoothly without bringing the whole program to a standstill waiting for one task to be completed. Ordinarily, after a webpage is scraped, there is time taken to load a webpage and you do not want other parts of the program to sit idle waiting for that event to execute. By utilizing asyncio, your script can ask for many webpages to be processed at the same time, thus saving a lot of time. logging : Keeping track of the flow of the program is done through logging. Picture this: you’ve left a scraper running for hours, and later discover that an error midway caused the scraper to stop. Without sufficient logs, you would have no idea what went wrong. By noting key activities, errors, and important messages, logging assists in error resolution.Instead of cluttering the console with print statements, logs are saved systematically so they can be reviewed whenever necessary. Playwright: Most of the code works on Playwright, a high-level web automation library. Unlike bare-bones scraping libraries that do nothing more than fetch webpage text, Playwright allows true interaction with web pages—just as a human does. It will scroll, press buttons, submit forms, and even handle pop-ups. It is useful in scraping websites whose data are dynamically loaded using JavaScript. Both async\_playwright and sync\_playwright are present in the code, with the former allowing multiple tasks to run at the same time to save time and the latter keeping it simple by running one task at a time. Either of these can be used based on the type of scraping task. Playwright\_stealth : But the majority of sites actively try to detect and prevent scrapers. This is where playwright\_stealth comes in handy. In regular situations, sites check for bot-like behavior by looking at how the browser and the page interact. If the interaction is unnatural or is too fast, the site will deny access. playwright\_stealth makes the interactions appear more human-like, avoiding sites from detecting that Playwright is controlling the browser, reducing the chances of being blocked. SQLite : Then, sqlite3 offers a method of storing and handling data in a light database. SQLite is a file-based, self-contained database system that does not need an independent server, so it is a great option for small to medium projects. When you scrape data, you have to have somewhere to put it so you can analyze it later, and SQLite is there to assist you in doing that effectively. Rather than storing all the information in memory, which may make your program slow or crash, SQLite puts it in a structured format so that it can be easily retrieved when required. Time: Another critical library in action is time, and it is primarily utilized to introduce delays when necessary. Sites have rate limits or anti-bot functionality, so sending requests too frequently may trigger alarm. A plain time.sleep() statement can retard things just sufficiently to make scraping seem more organic and prevent security features from being triggered. Pymongo : To save databases, pymongo is used to interact with MongoDB, a NoSQL database. SQLite saves information in tables (like an Excel spreadsheet), while MongoDB saves information in a more loosely structured manner, similar to JSON. It is perfect for handling large amounts of unstructured or semi-structured data, such as web-scraped data. When thousands of entries need to be saved without the need to care about strict table structures, MongoDB is a great choice. ## STEP 1 :Product URL Scraping From Tata Cliq Personal care appliances ### Importing Libraries ``` import asyncio    import sqlite3    import logging from playwright.async_api    import async_playwright from playwright_stealth    import stealth_async ``` This configuration brings together strong tools to ensure web scraping is efficient and effective. asyncio enables multiple tasks to be performed at the same time, and hence data extraction is fast. sqlite3 offers a lightweight database to save scraped urls in an orderly fashion. logging makes it possible to trace back the errors and execution status to ensure smoothness. async\_playwright makes web interaction easier, with handling sophisticated sites that integrate JavaScript elements. stealth\_async makes it easy to conceal from bot discovery to enable easy scraping without being blocked. These libraries together form a solid foundation for effective data extraction, proper handling, and addressing common challenges faced in web scraping. ### Setting Up Logging for Tracking and Debugging ``` # Logging setup logging.basicConfig(    filename="Log/tatacliq_original.log",                level=logging. INFO,                  format="%(asctime)s - %(levelname)s - %(message)s",   )                           """ Configures the basic logging system for the application. What this does: - Saves all log messages to 'tatacliq_original.log' in the Log folder - Records messages of level INFO and above (ignores DEBUG messages) - Formats each log entry with: [Timestamp] - [Log Level] - [Message] Note: - The log file will be created automatically if it doesn't exist - Existing log file will be appended to (not overwritten) - Useful levels: DEBUG (detailed), INFO (normal), WARNING (problems), ERROR (failures) """ ``` This script is equipped with a logging system for monitoring the scraping process and it is simple to monitor, identify issues, and debug. It stores all log messages in a file "tatacliq\_original.log" in the Log directory. The level of logging is INFO, and hence only required messages (i.e., status of script, warnings, errors) will be logged—no information at the debug level is logged to keep the log brief. Each log message follows a standard format: it has a timestamp (when the event happened), the log level (INFO, WARNING, ERROR), and a message (what happened). This makes all important events during scraping well-documented, a clean record of execution left behind. The log file is updated in real-time, with new entries being added instead of overwriting the old ones. This makes it possible to track script performance over time, diagnose failure, and seamless data extraction. ### Setting Up the Database: Creating a Reliable Storage for Product URLs ``` # Database setup def setup_database():    """    Sets up and prepares the SQLite database for storing product URLs.    What this function does:    1. Creates a connection to 'Tata_cliq_original.db' in the Data folder    2. Creates a table named 'product_urls' if it doesn't already exist    3. The table has two columns:       - id: Auto-numbered unique identifier (automatically increases)       - url: Web address of the product (must be unique, no duplicates allowed)    4. Saves (commits) these changes    5. Returns the active database connection    """    conn = sqlite3.connect('Data/Tata_cliq_original.db')    cursor = conn.cursor()    cursor.execute('''        CREATE TABLE IF NOT EXISTS product_urls (            id INTEGER PRIMARY KEY AUTOINCREMENT,            url TEXT NOT NULL UNIQUE        )    ''')    """    - Creates the database file if it doesn't exist    - Won't overwrite existing data if table already exists    - 'UNIQUE' ensures no duplicate URLs can be stored    - Remember to close connection when done    """    conn.commit()    return conn ``` The function, setup\_database(), establishes the database for storing product URLs that are scraped from the Tata Cliq Fashion site. It begins by creating a connection to an SQLite database file Tata\_cliq\_original.db that is located in the Data directory. SQLite will automatically create the file if it does not already exist. Under this database, a product\_urls table is created (if it has not already been created). This table contains two significant columns: id, a unique identifier auto-increasing for every new entry, and url, which holds product web addresses with no duplicate links being stored, courtesy of the UNIQUE constraint. After setting the structure, the function commits (saves) such changes in the database and returns the active database connection. This setup provides a structured way of holding product links, free of redundancy and ease of tracking and retrieval of scraped data in the future. It is advisable to close the database connection after use to release system resources and avoid complications. ### Saving Product URLs Without Duplicates ``` # Save URL to database def save_url_to_db(conn, url):    """    Saves a product URL to the database, preventing duplicates.    What this does:    - Takes a database connection and a URL    - Tries to add the URL to the 'product_urls' table    - If URL already exists, logs a warning and skips it    - If URL is new, saves it permanently to the database    Parameters:    - conn: An open database connection (from setup_database())    - url: The product URL to save (as a string)    What happens if URL exists:    - Logs: "Duplicate URL skipped: [url]"    - Doesn't crash or stop the program    - Just moves on without adding duplicate    """    cursor = conn.cursor()    try:        cursor.execute('INSERT INTO product_urls (url) VALUES (?)', (url,))        conn.commit()    except sqlite3.IntegrityError:        logging.warning(f"Duplicate URL skipped: {url}") ``` With the method save\_url\_to\_db, product URLs can be added to the database without violating uniqueness constraints and resulting in corruption of data due to repetitive entries. Two arguments are passed to the method which are the connection to the database (conn) and the product URL (url) and its attempts to insert the URL into the "product\_urls" within the database. If the URL does exist within the database, the program does not crash, but rather captures a warning “Duplicate URL skipped: \[url\]”, enabling it to continue executing smoothly. The function begins asynchronously creating a cursor for the database in order to run SQL commands. Next, he attempts to insert the new URL into the database with an SQL command INSERT. Upon successful capture of the new URL, it would be stored permanently in the database upon committing the transaction. If the captured URL does exist, then an sqlite3\. IntegrityError is raised because the database already has the URL stored and unique entries are restricted. The absence of the system crashing is made possible via exception handling of this error, with a log statement created for future examination. As a result the rest of the database is unaffected which makes it clean and effective making data storing reliable. ### Extracting Product URLs from Tata Cliq Efficiently ``` # Scrape URLs from the website async def scrape_tatacliq():    """    Main function to scrape product URLs from Tata CLiQ website.    What this function does:    1. Sets up database connection to store URLs    2. Opens a hidden browser (WebKit) to visit the website    3. Uses stealth techniques to avoid being blocked    4. Handles popups that might interfere with scraping    5. Collects all product page URLs while scrolling    6. Saves unique URLs to database    7. Shows progress in logs    Step-by-Step Process:    - Starts browser in visible mode (headless=False for debugging)    - Goes to personal care products page    - Tries to close any pop ups within 5 seconds    - Scrolls down to load more products    - Finds all product links on page    - Saves new URLs and counts them    Special Features:    - Remembers already seen URLs to avoid duplicates    - Logs every 10 URLs saved    - Uses 'network idle' wait to ensure page fully loads    - Gracefully handles popup errors without crashing    Note:    - Runs asynchronously (needs 'await' when called)    - Requires Playwright and stealth plugins    - Creates 'tatacliq_original.log' for progress tracking    - Needs internet connection to work    """        conn = setup_database()    base_url = "https://www.tatacliq.com/personal-care/c-msh1236?&icid2=nav:regu:audnav:m1236:mulb:best:08:R3:clp:bx:010"    async with async_playwright() as p:        # Launch browser in non-headless mode for debugging        browser = await p.webkit.launch(headless=False)        page = await browser. new_page()             # Apply stealth to avoid detection(# Apply stealth techniques to mimic human behavior)        await stealth_async(page)              # Navigate to the base URL        logging.info(f"Navigating to {base_url}")        await page.goto(base_url)        await page.wait_for_load_state('networkidle')        # Handle popup overlay        try:            await page.wait_for_selector('#wzrk-cancel', timeout=5000)            close_button = await page.query_selector('#wzrk-cancel')            if close_button:                is_disabled = await close_button.is_disabled()                if is_disabled:                    await page.evaluate("document.querySelector('#wzrk-cancel').removeAttribute('disabled')")                await close_button.click()                logging.info("Closed subscription popup.")        except Exception as e:            logging.warning(f"Popup handling failed: {e}")        unique_urls = set()        url_counter = 0        # Loop to load all products        while True:            # Scroll to bottom to trigger lazy loading            await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")            await asyncio.sleep(3)            # Extract all product links on current page            product_links = await page.query_selector_all('a.ProductModule__aTag')            new_urls_found = False            for link in product_links:                product_url = await link.get_attribute('href')                if product_url:                    # Ensure proper URL formatting                    full_url = product_url if product_url.startswith("http") else f"https://www.tatacliq.com{product_url}"                                       if full_url not in unique_urls:                        save_url_to_db(conn, full_url)                        unique_urls.add(full_url)                        url_counter += 1                        new_urls_found = True                        if url_counter % 10 == 0:                            logging.info(f"Saved {url_counter} URLs.")                       # Exit condition: no new URLs found in current iteration            if not new_urls_found:                logging.info("No new URLs found. Scraping complete.")                break            # Try clicking "Show More Products"            show_more_button = await page.query_selector('text=Show More Products')            if not show_more_button:                logging.info("No more products to load.")                break                       try:    """    Attempts to load more products by:    1. Finding and scrolling to the 'Show More product' button    2. Clicking it to load additional items    3. Waiting 3 seconds for content to load       Error Handling:    - If button missing: Stops with log message    - If click fails: Logs error and stops       Behavior:    - Scrolls button into view first    - Adds brief delay after click    - Breaks loop on any failure    """                await show_more_button.scroll_into_view_if_needed()                await show_more_button.click()                logging.info("Clicked 'Show More Products'.")                await asyncio.sleep(3)            except Exception as e:                logging.error(f"Error clicking 'Show More Products': {e}")                break        logging.info(f"Scraping complete. Total unique URLs saved: {url_counter}")        await browser.close()    conn.close() ``` The scrape\_tatacliq function is the heart of the web scraping process for Tata CLiQ’s personal care products section. The function handles the web scraping process of Tata CLiQ’s personal care products section by automating the process of visiting the website, gathering product page URLs, and saving them in a database. Overall, this function ensures efficiency and reliability. What this function does: 1. Establish Database Connection – The first step in scraping requires the function to connect with an SQLite database where the extracted URLs will be stored. 2. Open Browser – The next step is to navigate to the Tata CLiQ website which requires opening a WebKit based browser that is not in headless mode (this means it is visible for debugging purposes). 3. Stealth Mode – The next step is to perform some human simulation movements so that the chances of getting blocked while scraping the site are reduced. 4. Popup Handle – Subscription popups that appear when an abrupt site entry takes place will be attempted to be closed in 5 seconds to keep scraping uninterrupted. 5. Dynamic Product Scroll Down – To enhance user experience, Tata CLiQ loads products as the user scrolls down the page, so the function will automatically scroll down to load more products. 6. Capture Hyperlink To Product – The function attempts to locate product links in the webpage. This has to be done in a manner that the links are unique and formatted uniquely before capturing them. 7. Ignore Captured URL – The capture step will be avoided if the URL has already been captured to maintain uniqueness. 8. Logs Progress – The system keeps track of the scraping process in a log file while logging every 10 URLs saved. 9. Clicks “Show More Products” Button - The function searches and clicks the button “Show More Products”, if it exists and awaits the information to load before proceeding to scrape. 10. Gracefully Handles Errors– If an error occurs, or the button is missing the function will resolve the issue by logging it in the system and stopping without crashing. Step-by-Step Execution. - The function goes to the Tata CLiQ personal care products page. Instructed to navigate to the appropriate URL. - Waits for the page to load fully and attempts to close any popup that may show up. - Enters a loop of scrolling down the list extracting product URLS and then storing them in the database - Has the option of either resuming the process if URLS are detected or ceasing the process when none are detected. - Click the “Show More Products” button if it exists to be able to extract more products. - Logs the total amount of unique URLS obtained from the scrape, shuts the browser down, and reestablishes connection with the database after finishing. Because of the structure, the software can complete tasks without having to be monitored the whole time. Each step provides detailed instructions that ensure account for issues like popups, lazy loading, duplicates or even restrictions from the website. It really makes it easy and fast to scrap webs. ### Kicking Off the Scraping Process: Running the Main Function ``` # Main function async def main():           """ This is the starting point of the web scraping script. What this function does: - Logs the message "Starting scraping with WebKit..." to indicate the process has begun. - Calls the scrape_tatacliq() function to start scraping product URLs. - Uses 'await' to ensure the scraping process completes before the script moves forward. Why this is important: - Helps track when the scraping starts in the log file. - Ensures that the main script runs in an organized and structured manner. - Uses asynchronous execution for efficient handling of web scraping tasks. """    logging.info("Starting scraping with WebKit...")    await scrape_tatacliq() ``` The main() function serves as an official starting function in the Tata CLiQ web scraping operation. In other words, it is akin to hitting the “Start” button of the scraper. When the main() function is called, it logs a message on the log file which reads “Starting scraping with WebKit…” This is helpful for anyone who is tracking the process in knowing when the exact scraping activities started. After providing this log message, the function proceeds to call the scrape\_tatacliq() function with the await operator. This means the script will suspend execution until this function completes its work, so no further logic is executed till this scraping task is done. This ensures that all the product URLs are collected and saved properly without rushing or skipping any steps. This function is done in asynchronous fashion, which is quite helpful with time intensive action such as web scraping. Asynchronous execution enhances the efficiency of a program by multitasking and not getting stuck on a web page or on waiting or on other simulation tasks. While the main() function is short, it serves a very critical purpose which is to control all the scraping processes – coordinate the start, do the checks needed to confirm that all the scraping uses are done and conduct them in an orderly and systematic manner. From an execution point of view this design pattern adds to performance, but it also improves overall aesthetics and maintainability of the scraping script. ### Executing the Web Scraper Seamlessly ``` # Run the scraper asyncio.run(main()) ``` The line asyncio.run(main()) handles initiating and managing the entire web scraping workflow in an effective and systematic process. Because the main() function is asynchronous, it cannot be run like a traditional function. asyncio.run() starts the asynchronous event loop and manages it to ensure all asynchronous tasks such as loading web pages, scrolling, and collecting product URLs all execute to completion without any interruptions. This command serves as the final action that puts the entire web scraping process into action; effectively making sure the script runs smoothly from start to finish, managing multiple actions in a concurrent manner. ## STEP 2 :Scraping Comprehensive Product Details from Product Links ### Importing Libraries ``` import sqlite3 import time import logging import pymongo import subprocess from playwright.sync_api import sync_playwright from playwright_stealth import stealth_sync ``` This set of libraries gives decent functionality for scraping data from the web, and for managing that information as well. Pymongo offers a nice way to use MongoDB, a NoSQL database that is well suited for storing large volumes of semi-structured data. Subprocess allows system commands to be run from Python, which is useful when wanting to automate the use of opening a web browser or manipulating files. Sync\_playwright from playwright.sync\_api has added the capability of navigating and executing commands within complex websites using JavaScript, making automating web activities easier. Lastly, stealth\_sync from playwright\_stealth helps in avoiding bot detection by behaving like a human to avoid being blocked. These enable the user to scrape, store, and manipulate data with utmost ease whilst following the necessary precautions and being flexible to different website layouts. ### Essential Settings for Scraping and Data Storage ``` # Configuration """ Application configuration settings: - DB_PATH: Location of SQLite database file - TABLE_NAME: Name of the URLs table - LOG_FILE: Path for the log file - MONGO_COLLECTION_NAME: This sets the table name to "products" where all scraped items will be saved - MONGO_EXPORT_PATH: Backup location for MongoDB data """ DB_PATH = "/home/anusha/Desktop/DATAHUT/Tata_Cliq_Fashion/Data/Tata_cliq_original.db" TABLE_NAME = "product_urls" LOG_FILE = "/home/anusha/Desktop/DATAHUT/Tata_Cliq_Fashion/Log/Data_scraper_original.log" MONGO_DB_NAME = "tata_cliq_db_original" MONGO_COLLECTION_NAME = "products" MONGO_EXPORT_PATH = "/home/anusha/Desktop/DATAHUT/Tata_Cliq_Fashion/Data/mongodb_data_original.bson" ``` This part of the code specifies key configuration settings that describe where data will be stored, logged, and backed up when scraping the web. These settings make it easy and organized to collect and manage all data. Below is the description of each one: - DB\_PATH: This is the file path to the SQLite database. The SQLite database is a storage space for product URLs that were extracted during the scraping process to reference later. - TABLE\_NAME: This is the name of the table (in the SQLite database) the product URLs will be stored in; in this case, it is "product\_urls." - LOG\_FILE: The log file will store information about actions, errors, or events that occur during the scraping process that can be used for debugging and monitoring the scraping process. - MONGO\_DB\_NAME: Name of the MongoDB database where all extracted data will be saved. MongoDB is always a flexible NoSQL database that can efficiently store large and complex datasets. - MONGO\_EXPORT\_PATH: This is the file path where the MongoDB data will be backed up to. Backing up data allows your scraped data to become even safer, and allows restoring in case of failure. All of these settings guarantee a structured data storage and logging flow, which ultimately produces a better experience for the scraper by making it more efficient, reliable, and easier to manage. ### Setting Up Logging for Tracking and Debugging ``` # Logging Setup """ Configure logging system to: - Write logs to specified file - Record INFO level and above - Include timestamp, log level and message """ logging.basicConfig(filename=LOG_FILE, level=logging.INFO,                    format="%(asctime)s - %(levelname)s - %(message)s") ``` This code provides a basic logging system that we can use to log and track important events that occurred during the web scraping process in order to provide smooth operation and facilitate easier debugging. All log messages are written to a given file (LOG\_FILE) and we can therefore review them later instead of printing them (as we have done with the print statements). It uses a standard commit log level of INFO which means only meaningful updates, warnings and errors will be logged and unnecessary debug information can be filtered out. Each log item will shows the time stamp (to track exactly when the event occurred), the log level (to indicate whether it is an INFO or ERROR), and the message (to detail what event occurred). This makes it easier to maintain transparency, of course, but is also important to help with identifying and diagnosing problems quickly as information is core to achieving this, and lastly to also give assurance that the scraper runs smoothly over time. ### Connecting to MongoDB for Storing Scraped Data ``` # MongoDB Setup """ Initializes MongoDB connection: - Connects to local MongoDB instance - Selects database and collection - Will be used for storing product data """ client = pymongo.MongoClient("mongodb://localhost:27017/") db = client[MONGO_DB_NAME] collection = db[MONGO_COLLECTION_NAME] ``` The purpose of this code is to link the script with a MongoDB database that will store the scraped product data. It begins by creating a client connection via pymongo.MongoClient("mongodb://localhost:27017/"), which tells the script to connect to the MongoDB server that should be running locally on the machine, at port 27017 (the default port for MongoDB). If the server is running, this connection will allow you to interact with the database. Next, db = client\[MONGO\_DB\_NAME\] allows you to select a specific database where the scraped product information will be stored with a specific structure, making it much easier to write and retrieve information. The last line, collection = db\[MONGO\_COLLECTION\_NAME\], selects a collection from the database. Although some may refer to it as a table, it has a similar concept in that multiple entries of product information will be stored. This is important in order to effectively manage and organize a large amount of scraped information. Whereas SQL databases employ a rigid schema, MongoDB uses a more flexible document-based structure which is suitable for a variety of product details with easier retrieval and analysis. Because the script will connect to the database and collection at the beginning of the scraping process, you will thus have a way to save all product information that is pulled in during the execution of the script. ### Database Query Executor for SQLite Operations ``` def execute_query(query, params=None, fetch=False):    """Helper function to execute SQLite queries.    Args:        query: SQL query string        params: Optional parameters for query        fetch: Whether to return results    Returns:        Query results if fetch=True, else None    Handles:        - Opening/closing connections        - Parameterized queries        - Commit operations    """    conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    cursor.execute(query, params or ())    result = cursor.fetchall() if fetch else None    conn.commit()    conn.close()    return result ``` The provided function, execute\_query, is an auxiliary function that interacts with an SQLite database and is especially relevant for ensuring efficiency. It prevents the script from repeating similar connection and execution logic whenever an SQL statement needs to be executed. The function begins by getting a connection to the SQLite database located at the DB\_PATH to make sure that any database interaction occurs with the appropriate database file. The function also creates a cursor, which is used to execute presumably safe SQL commands. The function will take in three parameters: query, which are the SQL commands to be executed; params, which are optional (allowed to pass the parameters to prevent SQL injection and be safer); and fetch, to request if it returns results (useful for SELECT statements). If fetch principle equals True, then it will return the queried results if run, or otherwise it would run SQL without return. Finally, the function is created to repeat the process by committing all changes once the query is run, it will then close the connection to release resources. Centralizing interactivity within this function for database management enables a script to continue being clean and readable while executing all required pre-tasks of opening and closing the database, executing safe parameterized queries, committing, or discarding changes—all safely and efficiently. ### Database Setup for Tracking Scraping Progress ``` def setup_database():    """    Initializes and prepares the database structure for URL tracking.    Purpose:    - Ensures the database table has a 'processed' column to track which URLs have been scraped    - Runs automatically at script startup to maintain data integrity    What it actually does:    1. Checks existing table structure for columns    2. If 'processed' column doesn't exist:       - Adds new column with default value 0 (unprocessed)       - Logs this modification    3. If column exists:       - Simply confirms its presence in logs    Database Schema Modification:    - Adds column: processed (INTEGER)    - Default value: 0 (False/Unprocessed)    - Values:       0 = URL not yet processed       1 = URL successfully scraped            Why this matters:    - Prevents duplicate scraping of the same URL    - Enables restart capability if script stops midway    - Maintains scraping progress tracking    """      logging.info("Setting up the database...")    columns = execute_query("PRAGMA table_info(product_urls)", fetch=True)    column_names = [col[1] for col in columns]    if "processed" not in column_names:        execute_query("ALTER TABLE product_urls ADD COLUMN processed INTEGER DEFAULT 0;")        logging.info("Added 'processed' column to track processed URLs.")    else:        logging.info("'processed' column already exists.") ``` The setup\_database function is a critical component of the scraping pipeline, as it will ensure that we set up a structure that can easily accommodate tracking scraping progress. This function is run automatically when the script runs, so every URL that is stored in the database will provide a status indicator to demonstrate whether it has been scraped before or not. The function begins by logging that the database is currently in setup; this demonstrates what is currently occurring within the execution flow. The function will then call the existing structure of the product\_urls table using an SQL command to check the PRAGMA table\_info(product\_urls); this will return the metadata for all columns in the table. The function first retrieves the rows from the column names and checks if the processed column is already included. The processed column is integral to knowing what URLs have been scraped successfully, so that we can understand when we are scraping duplicate URLs, which can reduce efficiency as well. In the event that the column does not exist, the function will modify the schema of the database by issuing an ALTER TABLE SQL command to insert a new column called processed, with an integer type, and a default integer value of 0-- a value of zero means that the URL has not been processed, while a value of 1 will be assigned to the column when the URL has been processed without issues. This allows the function to restart from where it left off, rather than starting over, after an unplanned interruption. After it finishes adding the column the function will also log that it added the column to the logging database. In the event that the column already exists, the function will log the message that the column existed and no changes were made. This way of tracking the ETL process will improve the overall performance and reliability of the web scraping process with a systematic way to confirm progress, without necessitating extra scraping of data (through NULL values), and make resuming an objective upon interruptions. ### Removing Duplicate URLs to Keep Data Clean and Unique ``` def remove_duplicates(): """ Removes duplicate URLs from the database to ensure only unique records are retained.      Purpose: - This function eliminates duplicate entries in the database while preserving the earliest recorded instance of each unique URL. - Helps maintain data integrity and efficiency by reducing redundant entries. How It Works: 1. Logs the start of the duplicate removal process. 2. Executes an SQL query that identifies and deletes duplicate URLs.     - The query groups records by the `url` column.     - It retains only the entry with the smallest `rowid` (earliest inserted record).     - All other occurrences of the same URL are deleted. 3. Logs a confirmation message once the process is complete. SQL Query Explanation: - The subquery `SELECT MIN(rowid) FROM {TABLE_NAME} GROUP BY url` finds the smallest `rowid` for each unique URL. - The main `DELETE` query removes all rows where the `rowid` is NOT in the list of earliest `rowid` values. - This effectively deletes all duplicate URLs while keeping only the first recorded instance.  Logging: - The function logs both the start and completion of the duplicate removal process. - This helps track database maintenance activities and ensures visibility into cleanup operations.             """    logging.info("Removing duplicate URLs from the database...")    execute_query(f"""        DELETE FROM {TABLE_NAME}        WHERE rowid NOT IN (SELECT MIN(rowid) FROM {TABLE_NAME} GROUP BY url)    """)    logging.info("Duplicates removed successfully.") ``` The remove\_duplicates function helps to clean the database from duplicate product URLs and keeps only unique entries available. Duplicate URLs can happen as a result of multiple scraping sessions, a bad network connection, or simply because a URL was added again, resulting in processing the same data multiple times, and/or needlessly consuming space in memory that handling the data requires. The remove\_duplicates function will start by logging an information log message that says it is going to start the process of duplicate removal. Following the information log, it runs an SQL query that will systematically identify and remove duplicate records, while leaving the first instance of each unique URL in the database.This process is accomplished using the rowid column which is unique to each row within a SQLite table. The SQL group by clause groups all identical URLs together and the select MIN(rowid) effectively removes each character added to the database but keeps the first character added to the database. In this manner, the duplicate removal function is able to keep the product data valid and remove the extra duplicates once the first instance has been maintained. After the query finishes running the duplicates are removed successfully, an information log is logged to notify duplicates are removed successfully. This process will help keep the database clean and efficient, and it reduces storage overhead while also preventing the scraper from repeatedly scraping the same URLs again and again. By running the function regularly, you can maintain the dataset as streamlined, improving the accuracy and reliability of the web scraping process overall. ### Fetching unprocessed URLs for Scraping ``` def get_unprocessed_urls():    """    Retrieve unprocessed product URLs from the SQLite database, ordered by ID    Returns:        List of (id, url) tuples ordered by ID    Logs:        - Count of found URLs    Filters:        - Only URLs with processed=0        - Ordered by ascending ID    """    logging.info("Fetching unprocessed URLs ordered by ID...")    result = execute_query(f"SELECT id, url FROM {TABLE_NAME} WHERE processed = 0 ORDER BY id ASC", fetch=True)    logging.info(f"Found {len(result)} unprocessed URLs.")    return result ``` The get\_unprocessed\_urls function retrieves product URLs from the database that have yet to be processed by the web scraper, allowing this programmatic script to avoid reprocessing a URL from a previous scraping session, and to resume scraping from this last session. The function logs that it is fetching unprocessed URLs so that you are auditing the process of scraping. The function afterwards runs an SQL select all query that selects all rows from the database products table that have a value of 0 in the processed column, meaning that the URL has not yet been scraped. The rows will be ordered in ascending order by id, to ensure the flow of processing information is organized and predictable. Finally, the function logs the count of unprocessed URLs found, to help you keep track of progress in the scraping process. The function then returns the taken URLs in a list of tuples, where each tuple contains an id, and url. This setup allows other sections of the script to loop through and process the URL in an efficient manner. By excluding URLs that have already been processed, this function minimizes the effort of scraping the eligible URLs, as there will not be any duplicate requests increasing overall performance. ### Marking URLs as Processed to Prevent Duplicate Scraping ``` def mark_url_processed(url):    """    Mark a URL as processed in the database    Updates a URL's status to 'processed' in the database to prevent re-scraping.    Functionality:    - Sets the 'processed' flag (1) for a specific URL in the database    - Provides feedback via logs about success/failure    - Ensures proper database connection handling    Parameters:        url (str): The exact product URL to mark as processed    Database Operation:    - Executes: UPDATE product_urls SET processed = 1 WHERE url = [provided_url]    - Uses parameterized queries to prevent SQL injection    - Commits changes immediately       Logging Behavior:    - Success: "Successfully marked as processed: [url]"    - Failure: "Failed to mark as processed: [url]" (if URL not found)    Error Handling:    - Implicit: Fails gracefully if URL doesn't exist (rowcount = 0)    - Explicit: Closes database connection even if errors occur    """       conn = sqlite3.connect(DB_PATH)    cursor = conn.cursor()    cursor.execute(f"UPDATE {TABLE_NAME} SET processed = 1 WHERE url = ?", (url,))    conn.commit()       if cursor.rowcount > 0:        logging.info(f"Successfully marked as processed: {url}")    else:        logging.warning(f"Failed to mark as processed: {url}")       """    Verification and logging of database update status.    What this checks:    - cursor.rowcount: Number of rows affected by the UPDATE query    - > 0 means the URL was found and marked successfully    - == 0 means no matching URL was found    Logging Behavior:    - Success: Logs INFO level message with the processed URL    - Failure: Logs WARNING level message with the problematic URL     """       conn.close() ``` The mark\_url\_processed function plays a crucial role in making sure that a product URL gets scraped only once. It updates the database to denote that a URL has been "processed," which avoids it getting scraped again, and more generally, to optimize the scraping process. When this function is called, it first connects to the SQLite database and creates a cursor object to use for connecting to the database. This function then executes an UPDATE SQL statement that sets the processed column to 1 for a specific URL. 1 denotes that the URL has been scraped and added to the database, and it will not scrape this URL again in future executions of the script. Also, to improve security, and to avoid SQL injection, the function uses parameterized queries, which ensure that the specified url is treated as data and is not part of the query structure. After the update is executed, the function commits the transaction, saving the changes made in the database before the database connection is closed. The function then checks whether the update was successful by testing the cursor.rowcount which returns the number of rows affected by the operation. If rowcount > 0, it means that the URL was found in the database and was marked as processed. In this case, a log message is recorded at the INFO level to indicate that the URL has moved on to processed. If rowcount == 0 URL was not found in the database. This does not indicate an error, but that the URL could not be marked. This could be for various reasons but is most often because a misspelled URL was input or the entry is new. In this case, the log message indicates this situation at WARNING level that the URL could not be marked as processed. Ultimately, the function guarantees that the database connection is always closed, no matter if the operation completed successfully or not. This is important for resource management to avoid memory leaks and database locks. Using this update method allows the scraping script to effectively track program progress, so it can always partially resume where it left off if something was interrupted. ### Extracting Text from Web Pages using a CSS selector ``` def extract_text(page, selector):    """    Extract text from the page using a CSS selector, handling missing elements gracefully    Args:        page: Playwright page object        selector: CSS selector string    Returns:        Extracted text or "N/A" if not found    """    try:        return page.locator(selector).text_content().strip()    except:        return "N/A" ``` The extract\_text function is used to get text content from a webpage via a CSS selector without causing the script to fail if the element does not exist. It has two parameters: page, which is a Playwright page object for the current webpage being processed, and selector, a string that defines the CSS selector for the target element. The function attempts to locate the element with Playwright's locate function and retrieve its text content. If successful, it strips leading/trailing spaces and returns the cleaned text. However, if the element cannot be found or if there is a failure in extraction, the function bypasses the error by returning "N/A" instead of killing the script. This gives stability to the web scraping process, preventing errors from stopping data harvesting when an assumed element is not found on the page. ### Extracting general\_features from Tata CLiQ Pages ``` def extract_general_features(page):    """    Extracts product features from a Tata CLiQ product page using Playwright.    This function:    - Finds all feature containers on the page    - For each container, extracts feature names (headers) and their corresponding values    - Returns a dictionary of {feature_name: feature_value} pairs    Parameters:        page (playwright. page): The Playwright page object currently on a product page    Returns:        dict: A dictionary where:              - Keys are feature names (e.g., "Material", "Warranty")              - Values are the corresponding feature details              - Returns empty dict if no features found or error occurs    Error Handling:    - Catches and logs any exceptions during extraction    - Returns empty dict on error to allow graceful continuation             """    try:        features = {}        elements = page.locator(".ProductFeatures__content").all()               for element in elements:            headers = element.locator(".ProductFeatures__header.ProductFeatures__description").all()            values = element.locator(".ProductFeatures__description").all()                       if len(headers) > 0 and len(values) > 1:                key = headers[0].text_content().strip()                value = values[1].text_content().strip()                features[key] = value               return features    except Exception as e:        logging.error(f"Error extracting general features: {e}")        return {} ``` The extract\_general\_features function scrapes important product information from the product page on Tata CLiQ using Playwright. When the function runs, it will examine the webpage for containers specifying product features and extract relevant information from those containers (i.e., material, warranty, specifications, etc.). Specifically, the function starts by declaring an empty dictionary, features, which will hold (or store) all details collected. From there, the function searches through the page for all instances of class .ProductFeatures\_\_content, which holds product features. These instances will be stored as a list, and the function will loop through each one of those to extract feature names and values. For each feature container, the function will look for its corresponding headers (examples "Material" and "Warranty") and their related descriptions. Headers will be found using the CSS selector .ProductFeatures\_\_header.ProductFeatures\_\_description and values will be accessed through .ProductFeatures\_\_description. The function checks, before it continues to match values to headers, that there is at least one header present and more than one value. The text that has been extracted will be cleaned of unnecessary white space and stored in the dictionary as key-value pairs. Should an error occur during the extraction process (e.g., the element is missing or the expected structure of the website changes), the function simply raises an exception and logs an appropriate error message. The scraping process does not terminate; the function returns an empty dictionary, allowing the script to continue execution without interruption. This is a robust error-handling process that guarantees the scraping script remains stable and reliable, which means it won't fail merely due to a minor inconsistency on the webpage. ### Extracting Complete Product Details from Tata CLiQ Pages ``` def fetch_product_details(url, page):    """    Extracts comprehensive product details from a Tata CLiQ product page.    This function:    - Navigates to the product page URL    - Waits for key elements to load    - Extracts multiple product attributes using CSS selectors    - Returns structured product data    - Handles errors gracefully with detailed logging    Parameters:        url (str): The complete product page URL to scrape        page (playwright.page): Playwright page object for browser automation    Returns:        dict: Structured product data with these fields:            - url (str): Product page URL            - product_name (str): Name of the product            - brand_name (str): Manufacturer/brand name            - brand_info (str): Additional brand information            - price (str): Current selling price            - mrp (str): Original maximum retail price            - discount (str): Discount percentage/amount            - rating_value (str): Numeric rating (e.g., "4.2")            - rating_count (str): Number of ratings            - review_count (str): Number of written reviews            - product_description (str): Full product description            - general_features (dict): Key-value pairs of product features        Returns None if scraping fails    Error Handling:    - Logs detailed error messages including the failed URL    - Returns None on failure to allow graceful error handling    - Uses generous timeouts (60-80 seconds) for slow-loading pages    Implementation Details:    - Uses a lambda helper 'extract()' for consistent element handling    - Returns "N/A" for missing fields rather than failing    - Combines both direct selector extracts and feature extraction    - Dependent on extract_general_features() for feature details    - Includes debug logging for tracking progress    Selector Notes:    - All selectors target specific Tata CLiQ DOM structures    - Uses :not(:empty) to avoid blank price elements    - Relies on itemprop attributes for rating metadata    """    try:        logging.info(f"Scraping URL: {url}")        page.goto(url.strip(), timeout=80000)        page.wait_for_selector(".ProductDetailsMainCard__linkName > div:nth-child(1)", timeout=60000)        extract = lambda selector: page.locator(selector).text_content().strip() if page.locator(selector).count() > 0 else "N/A"        product = {            "url": url,            "product_name": extract(".ProductDetailsMainCard__linkName > div:nth-child(1)"),            "brand_name": extract("#pd-brand-name > span:nth-child(1)"),            "brand_info": extract("div.ProductDescriptionPage__detailsHolder:nth-child(1) > div:nth-child(1) > div:nth-child(4) > div:nth-child(2) > div:nth-child(1)"),            "price": extract(".ProductDetailsMainCard__price *:not(:empty)"),            "mrp": extract(".ProductDetailsMainCard__cancelPrice"),            "discount": extract(".ProductDetailsMainCard__discount"),            "rating_value": extract(".ProductDetailsMainCard__reviewElectronics[itemprop='ratingValue']"),            "rating_count": extract(".ProductDetailsMainCard__ratingLabel[itemprop='ratingCount']"),            "review_count": extract(".ProductDetailsMainCard__ratingLabel[itemprop='reviewCount']"),            "product_description": extract("div.ProductDescriptionPage__detailsHolder:nth-child(1) > div:nth-child(1) > div:nth-child(1) > div:nth-child(2) > div:nth-child(1)"),            "general_features": extract_general_features(page)        }        logging.info(f"Successfully scraped: {url}")        return product    except Exception as e:        logging.error(f"Error scraping {url}: {e}")        return None ``` The fetch\_product\_details function is intended to extract detailed product details from a Tata CLiQ product page using Playwright. The function is designed to handle the fetching of data, using the URL of the product page which gets processed into visiting that page. It makes sure that the key elements have been loaded before extracting the product title, brand, price details, discounts, ratings & reviews, and a product description in-depth. In addition to extracting those data points, the function extracts additional product specifications in the separate extract\_general\_features() function. One of the key features of this function is the extract helper lambda function. The purpose of the extract function is for the consistency and cleanliness of the data retrieval itself. The way it works is that it checks for whether an element is present for you to extract text from, and if it is not, the function will assign a default value of "N/A" rather than create a function error. The procedure begins with a log of the URL of the product being processed, before then navigating to the product page in a longer time limit to accommodate slower page-load speeds. Once the main items are established, the product data is extracted according to the predetermined CSS selectors that apply to the Tata CLiQ Web page style. Once documented, the product dictionary is then filled with that extracted data to offer the structured data in a more human-readable format. Some notable attributes would include product\_name, the capturing of its title, brand\_name, for the name of the manufacturer, price and mrp prices reflecting the present and original prices respectively, and discount, which is the capture of any available discounts. Customer metrics are also available, such as rating\_value, rating\_count, and review\_count, to provide a picture of customer interaction. To enhance reliability, the function is equipped with a handful of error-handling techniques. If there is ever an issue detected while the program executes, something like a missing element or an unexpected change to the website, it will log the error and the URL of the failed process and return None in a controlled manner, rather than halt the execution all together. This gives the scraper the opportunity to continue on processing other URLs. The function also permits generous timeouts to load a page or an element so that it does not lose data for slow loading networks. In addition, all of the members of the logs provide a chance for members to glean scraping as it happens, and to try to make the 'dotting the I's and crossing the T's' on errors as easy as possible. Debug logs connect all previous discussions with the event that indicates the scraping is still progressing. By linking structured data extraction, robustness to failure, and log progress, the fetch\_product\_details function offers a way to reliably scrape and collect product data from Tata CLiQ to ingest for analysis and/or storage. ### Saving Scraped Product Data to MongoDB Without Duplicates ``` def save_to_mongodb(data):    """    Safely saves scraped product data to MongoDB while preventing duplicates.    This function:    - Checks if the product URL already exists in MongoDB    - Transforms and stores new records with proper ID mapping    - Provides detailed logging of all operations    - Gracefully handles errors during database operations    Parameters:        data (dict): A dictionary containing product details with these required keys:            - id (int): The SQLite primary key (will become _id in MongoDB)            - url (str): The product URL (used for duplicate checking)            - Other product attributes (name, price, etc.)    Behavior:    1. Duplicate Check:       - Uses the URL field to check for existing records       - Skips insertion if URL already exists (idempotent operation)    2. Data Transformation:       - Moves SQLite 'id' → MongoDB '_id' field       - Removes the original 'id' field to avoid data duplication    3. Database Operations:       - Performs atomic insert if record is new       - Commits changes immediately    4. Logging:       - Success: "Saved to MongoDB: [url]"       - Duplicate: "Skipped duplicate: [url]"       - Errors: "Error saving to MongoDB: [error_details]    Error Handling:    - Catches and logs all database exceptions    - Prevents crashes from duplicate key errors    - Maintains data consistency    Notes:    - Depends on global 'collection' MongoDB collection object    - Designed for use with Tata Cliq scraping pipeline    - Preserves original SQLite record relationships via _id    """    try:        if collection.find_one({"url": data["url"]}) is None:            # Use the SQLite `id` as the `_id` in MongoDB            data["_id"] = data["id"]            # Remove the `id` field to avoid duplication            del data["id"]            # Insert the document into MongoDB            collection.insert_one(data)            logging.info(f"Saved to MongoDB: {data['url']}")        else:            logging.info(f"Skipped duplicate: {data['url']}")    except Exception as e:        logging.error(f"Error saving to MongoDB: {e}") ``` The save\_to\_mongodb function stores web-scraped product information into a MongoDB database without inserting duplicate records. As the function calls, it will check first whether the product URL has already been inserted into the MongoDB collection or not. This function's duplicate-checking functionality prevents a product from being inserted twice, resulting in no redundancy of data being stored and making the database performance-optimized. The operation anticipates data in dictionary format, with the product information of SQLite id, url, etc., and additional information like name and price. It also transforms the SQLite id field into the MongoDB \_id field to create a uniform and unique identifier. To avoid undesirable duplication, it drops the original id field from the dictionary before inserting it into MongoDB. If the product URL is not found in the database, the function adds the new record and outputs a success message. If the URL already exists, it outputs that the record was skipped and maintains processed entries. In addition, the function is error resilient—if there is any error during execution of the database operation, for example, errors related to the connection or unknown errors, it catches the exception and outputs an error message rather than crashing the script. This is to ensure smooth operation and data integrity. With the integration of deep logging, the function provides real-time feedback on database interactions, and debugging and monitoring of the web scraping pipeline becomes easier. The method is particularly tailored for the Tata Cliq web scraping process and is instrumental in storing structured product data in a way that does not influence the integrity of the dataset. ### Backing Up MongoDB Database to a Compressed File ``` def export_mongodb():    """    Exports the entire MongoDB database to a compressed archive file using mongodump.    This function:    - Creates a gzipped backup of the specified MongoDB database    - Saves the backup to the predefined export path    - Provides detailed logging of the export process    - Handles errors gracefully with specific error logging    Workflow:    1. Initiates mongodump command with these parameters:       - --db: Specifies the database name (from MONGO_DB_NAME)       - --archive: Outputs to a single compressed file (MONGO_EXPORT_PATH)       - --gzip: Enables compression to reduce file size    2. Logs start/stop messages for tracking:       - "Exporting MongoDB database..." (when starting)       - "MongoDB export successful!" (on completion)       - Specific error messages if failed    Configuration Requirements:    - MONGO_DB_NAME: Must be set to a valid database name    - MONGO_EXPORT_PATH: Must be a writable file path with .bson extension    - mongodump must be installed and in system PATH    Error Handling:    - Catches subprocess.CalledProcessError specifically    - Logs detailed error message including the actual command failure    - Does not crash the application on failure    Notes:    - Requires MongoDB tools installed (mongodump specifically)    - Runs as a blocking operation (will pause script during export)    - Output file uses BSON format (MongoDB's binary JSON)    - Compression reduces file size significantly (--gzip flag)    - Preserves all collections in the database    """    try:        logging.info("Exporting MongoDB database...")        subprocess. run(            ["mongodump", "--db", MONGO_DB_NAME, "--archive=" + MONGO_EXPORT_PATH, "--gzip"],            check=True        )        logging.info("MongoDB export successful!")    except subprocess.CalledProcessError as e:        logging.error(f"MongoDB export failed: {e}") ``` The export\_mongodb function has been implemented to export the whole MongoDB Database and store it in compressed format. Therefore, all scraped product information is secured and can be retrieved later when needed. The function calls the mongodump command, which is a native tool from MongoDB that will dump the data in the database to a defined bson, which is binary json. The output file is then compressed using the gzip feature enabled by the --gzip flag, which heavily compresses the output without sacrificing data integrity. When executed, the function writes a message first indicating that export has started. It then executes the mongodump command with arguments that pass the database name (MONGO\_DB\_NAME) and output path (MONGO\_EXPORT\_PATH). The script completes the exporting process successfully with the use of the check=True parameter that causes it to terminate with an error if any command fails. It returns a success success message that indicates export was successful on successful export.In case any process fails within it—i.e., installing mongodump, setting the database wrongly, or permissions—it traps the exception and publishes an extensive message for the failure. This does not cause the whole scraping process to crash but gives useful debugging information. mongodump must be installed and available on the system PATH for this function to work. Also, MongoDB database name and export file location must be set correctly. The backup file is saved in BSON format, so thereafter it can be restored by using MongoDB's mongorestore command. This position is critical for data safeguarding and preventing precious scraped information from loss due to system failures or crashes. ### Scraping and Storing Product Data from URLs in MongoDB ``` def scrape_all():    """    Scrape all product URLs stored in the database in order of ID and save to MongoDB progressively.    This main controller function:    1. Initializes and prepares the database    2. Cleans existing URL data    3. Processes all un-visited product pages systematically    4. Stores results in MongoDB    5. Exports the final dataset    Workflow Steps:    ----------------------------    1. Database Setup:       - Ensures proper schema exists       - Removes duplicate URLs       - Retrieves unprocessed URLs ordered by ID    2. Browser Initialization:       - Launches headless WebKit browser       - Configures stealth mode to avoid detection       - Creates fresh browsing context    3. URL Processing:       - Processes URLs sequentially       - Validates URL format       - Extracts product details       - Handles failures gracefully    4. Data Management:       - Enriches data with SQLite ID       - Saves to MongoDB (with duplicate prevention)       - Updates URL status in SQLite       - Includes 2-second delay between requests    5. Completion:       - Exports MongoDB collection       - Cleans up resources       - Provides completion logging    Error Handling:    - Skips invalid URLs (logged as warnings)    - Continues on individual URL failures    - Preserves state between runs via 'processed' flags    - Comprehensive error logging at each stage    Configuration:    - Uses headless browser for efficiency    - 2-second delay between requests (adjustable)    - Timeout values inherited from helper functions    Notes:    - Maintains state via SQLite 'processed' flags    - Processes URLs in ID order for consistency    - Requires MongoDB and SQLite to be properly configured    - All actions are logged for progress tracking    """       setup_database()    remove_duplicates()    url_entries = get_unprocessed_urls()    if not url_entries:        logging.info("No unprocessed URLs found. Exiting...")        return    with sync_playwright() as p:        browser = p.webkit.launch(headless=True)        context = browser. new_context()        page = context. new_page()        stealth_sync(page)        for product_id, url in url_entries:            if not url.startswith("http"):                logging.warning(f"Skipping invalid URL: {url}")                continue            try:                product = fetch_product_details(url, page)                if product:                    # Add the SQLite `id` to the product data                    product["id"] = product_id                    # Save to MongoDB                    save_to_mongodb(product)                    # Mark the URL as processed in SQLite                    mark_url_processed(url)                    logging.info(f"Successfully processed and saved: {url}")                else:                    logging.warning(f"Failed to scrape product details for: {url}")            except Exception as e:                logging.error(f"Error processing URL {url}: {e}")            time.sleep(2)        browser.close()    export_mongodb()    logging.info("Scraping process completed successfully.") ``` The scrape\_all function is the primary function that controls the complete data scraping process of Tata Cliq fashion products pages. It retrieves all the stored product URLs in the database one by one, fetches the necessary information, and saves it to mongodb while making sure that the program runs efficiently without causing any errors. This function has its own workflow which the user divides in several steps. Firstly, it goes to the database and initializes it by creating the appropriate schema and getting rid of repeat urls so that unique urls are the only tags that get processed. Then, it fetches all the unsorted product urls in the database which have not yet been worked on, ordered by the database id of the table, which helps in maintaining consistency. If there isn’t an unprocessed url, the function captures this and exits so as to not exhaust system resources. Once the URLs are obtained, the next thing the function does is start up a headless WebKit browser through Playwright because this enables quick and discreet scraping. The way the browser works will decrease the likelihood of being spotted by the anti scraping techniques of Tata Cliq. After that it creates a new browsing context for every session which provides a blank and separate space from which data can be fetched. Every url will be handled one at a time so that no product page will be missed.Before scraping, the function verifies if the URL is valid, and if it is not properly formatted (e.g., missing "http"), it logs a warning and moves to the next URL without crashing the entire process. For each valid product URL, the function calls fetch\_product\_details, which extracts all relevant product information. If successful, it enriches the extracted data by adding the SQLite database ID before storing the structured information into MongoDB. The save\_to\_mongodb function is responsible for handling the database insertion while preventing duplicate entries. Once data is successfully stored, the function marks the URL as processed in the SQLite database using mark\_url\_processed, ensuring that the same URL is not re-scraped in future runs. If any issues occur during the scraping process, such as a page failing to load or missing data, the function catches the exception, logs an error message, and moves on to the next URL without interrupting the entire workflow. To avoid overloading the website with requests and to mimic human-like behavior, the function includes a two-second delay between each request. This delay can be adjusted as needed. After processing all URLs, the function closes the browser to free up system resources and then calls export\_mongodb to create a compressed backup of the MongoDB database. This ensures that all scraped data is safely stored and can be restored later if needed. Finally, it logs a completion message indicating that the entire scraping process was executed successfully. Throughout the process, extensive logging is used at every step to track progress, identify issues, and ensure transparency. This function is designed to be highly robust, preserving its state across multiple runs by using the processed flag in SQLite. This means that if the script is interrupted for any reason, it can resume from where it left off without having to start over. The scrape\_all function is a crucial component of the Tata Cliq data scraping pipeline, integrating various helper functions to efficiently extract, store, and manage product data in an automated and structured manner. ### Entry Point to Execute the Script ``` if name == "__main__":    """    Entry point when run as script:    - Executes main scraping function    - Handles any top-level errors    """    scrape_all() ``` The if name == "\_\_main\_\_": clause is the principal entry point of the script and guarantees that the scraping operation occurs only when the script is run directly. It is a popular Python convention preventing accidental execution in case the script is imported as a module for another program. In this block, the scrape\_all() function is invoked, which is the main controller for handling the overall web scraping process for Tata Cliq fashion items. This function coordinates all major operations such as database initialization, fetching of product information, and storing within MongoDB, making sure that the pipeline of scraping works effectively from beginning to end. Further, this entry point is designed to gracefully handle root-level errors in order to keep the script from crashing when facing unexpected failure. In the event of any failure in the process, it would be logged and hence will make debugging easy, and also in subsequent runs it will continuously run. The design structure makes the script steadier and more modular to make it composite or reusable into more complex data processing streams. In effect, this block ensures that whenever the script is run as a stand-alone application, the scraping operation automatically occurs in a contained manner. ## Conclusion This data pipeline of the Tata Cliq Fashion website design is a manually automated property harvesting process that is well organized. The system employs Playwright for automated web browsing, SQLite for tracking progress, and MongoDB for data storage, ensuring that these goals are met on reliable, scalable, and complete systems.Having error handling, duplicate prevention and logging at every point makes this scraper efficient, robust and accurate all at the same time, which reduces data loss. This approach enables the system to be easily maintained and improved, thus providing a facility for changing data scraping requirements. From market analysis, product analysis to competition analysis, this scraping solution aids in accurate and dependable extraction of quality e-commerce data. ## Libraries and Versions Name: pymongo Version: 4.10.1 Name: playwright Version: 1.48.0 Name: playwright-stealth Version: 1.0.6 ## FAQ's 1\. Is it legal to scrape Tata CLiQ for product data? Scraping publicly available data is often legal for personal or research use, but it's essential to review Tata CLiQ’s Terms of Service. Always respect their robots.txt file and avoid violating copyright or usage policies. 2\. What kind of personal care data can I extract from Tata CLiQ? You can scrape data such as product names, prices, discounts, ratings, reviews, ingredients, brand details, availability, and product descriptions. 3\. Which tools are best for scraping Tata CLiQ? Popular tools include Python libraries like BeautifulSoup, Scrapy, and Selenium. For large-scale scraping, tools like Playwright or Puppeteer can help with dynamic content rendering. 4\. How can I ensure the scraped data is accurate and up to date? Implement regular scraping intervals, data validation checks, and deduplication logic to ensure fresh and reliable insights from Tata CLiQ. 5\. What are some use cases for scraping Tata CLiQ’s personal care section? Use cases include competitor analysis, pricing strategy development, market trend tracking, brand benchmarking, and building recommendation engines for e-commerce. ### Why DIY Web Scraping Projects Often Fail for Businesses URL: https://www.blog.datahut.co/post/5-reasons-you-shouldn-t-diy-your-web-scraping-projects/ Last updated: 2026-09-07T09:44:17.000Z Ever wondered if building your own web scraper is really worth the effort? It might look like a quick way to save money and get custom data, but what you don't see upfront are the legal landmines, technical debt, and hidden costs that come with it. In the age of digital intelligence, businesses run on data. From tracking competitor prices to forecasting market demand, data is the edge every decision-maker needs. Web scraping—the process of extracting data from websites—has emerged as a critical tool in this transformation. But while it's tempting to build your own scraping scripts in-house, the reality is far more complex than it appears. What starts as a weekend hack can quickly become a legal, technical, and strategic nightmare. Here's why you shouldn't DIY your web scraping projects and why partnering with experts is the smarter route. ## What Is Web Scraping and Why Businesses Use It Web scraping, or web data extraction, is the automated collection of structured data from websites. It has use cases across e-commerce, real estate, logistics, finance, and media, helping teams gather insights that were once inaccessible or prohibitively expensive. ### Key Business Use Cases: - Competitive Price Monitoring: Track rival pricing in real-time to stay competitive. - Stock Availability Tracking: Know when your competitors run out of popular products. - Content Aggregation: Consolidate news, reviews, or product listings from multiple sources. - Lead Generation: Pull data from directories, job boards, or industry listings. - Trend Analysis: Monitor market sentiment through social media or blog content. Despite its usefulness, scraping the web isn’t as simple as running a script. Let's dive into the top five reasons DIY web scraping can backfire. ![Why DIY web scraping is a risky move ](https://www.blog.datahut.co/content/images/2026/07/img-378.png.webp) ## 1\. The Technology Is Complex and Always Evolving We've worked with Engineering teams at FAANG companies and they spend between 40–60% of total project time just on scraper maintenance and updates—not building them. 70% of our customers who previously attempted DIY scraping encounter frequent issues like anti-bot blocks and broken scripts. Today’s websites aren't static HTML documents. They use JavaScript, AJAX, lazy loading, infinite scrolls, and interactive user elements. Scraping them requires a deep understanding of how browsers render pages and how servers detect bots. ### DIY Challenges: Learning Curve: You’ll need to master tools like Scrapy, Selenium, Playwright, BeautifulSoup, and handle cookies, headers, and sessions manually. Bot Detection: Many sites employ anti-scraping tools like Cloudflare, Akamai, and Distil Networks. DIY setups often get blocked or blacklisted. Constant Breakage: Website structure changes often. One minor change in the HTML can break your script and corrupt your data. Infrastructure Management: Scraping at scale means managing headless browsers, rotating proxies, and ensuring your IPs don’t get banned. "You wouldn’t build your own CRM in 2025\. So why build and maintain your own scraper?" Related article: [Web Scraping vs API: What's the best way to extract data](https://www.blog.datahut.co/post/web-scraping-vs-api/) ## 2\. Legal and Compliance Risks Are Not Optional According to a 2023 report from Apify, more than 62% of businesses cited legal uncertainty as a top reason for outsourcing their scraping efforts. With regulations like GDPR and CCPA constantly evolving, the risk landscape is simply too complex for most internal teams to manage effectively. Web scraping laws are murky and jurisdiction-dependent. What might be technically possible isn't always legally safe. Countries are tightening rules around data privacy, platform access, and ethical data use. ### DIY Legal Pitfalls: Terms of Service Violations: Most websites explicitly prohibit scraping in their user agreements. Privacy Laws: GDPR (EU), CCPA (California), and others place strict controls on how data, especially personal data, can be accessed and stored. CFAA & Other Regulations: In the U.S., the Computer Fraud and Abuse Act has been invoked in web scraping lawsuits. Legal Threats: DIYers often receive cease-and-desist letters or worse—find themselves named in legal proceedings. Unless your team has in-house counsel or strong legal SOPs, you're putting your company at risk. Explore more on[ ](https://en.wikipedia.org/wiki/Computer%5FFraud%5Fand%5FAbuse%5FAct?ref=blog.datahut.co)[data privacy laws for web scraping](https://en.wikipedia.org/wiki/Computer%5FFraud%5Fand%5FAbuse%5FAct?ref=blog.datahut.co) ## 3\. Data Privacy, Security & Quality Often Get Overlooked A survey by ScrapingAPI found that over 50% of DIY web scraping projects suffered from low data reliability due to inconsistent formats, stale content, or incomplete extraction. This is often due to lack of validation, audit trails, and compliance tooling—features typically included in managed scraping platforms. One of the most common blind spots in DIY scraping is what happens after the data is collected. Raw scraped data is often messy, incomplete, and vulnerable. ### Key Risks: Unclean Data: Duplicates, HTML artifacts, and malformed entries require extensive cleaning. No Anonymization: DIY tools often miss the step of removing personally identifiable information (PII). No Audit Trails: Regulations require logs of how and where the data was sourced. DIY setups usually don’t track this. Security Gaps: Insecure storage or transfers of scraped data can lead to breaches. A scraping partner builds in these safeguards from the start—ensuring what you extract is usable, secure, and compliant. ## 4\. DIY Costs More Than You Think In a 2023 ProWebScraper case study, companies reported spending 30–50% more in engineering hours trying to fix and maintain DIY scrapers compared to outsourcing to specialized providers. What appears to be a cost-saving effort often ends up draining valuable engineering and business resources. The myth that building your own scraper is "free" dissolves quickly once hidden costs add up. ### True Costs of DIY: Developer Time: Expect weeks of coding, debugging, and patching. Every site breakage needs immediate attention. Proxy & Infrastructure Costs: Need to rotate IPs? Use headless browsers? Expect to pay for proxies, servers, and uptime monitors. Delays in Decision-Making: If you're spending all your time gathering data, who’s analyzing it? By contrast, outsourcing web scraping delivers: - Rapid Setup: Go from spec to dashboard in days, not months. - Focus on Analysis: Your teams focus on driving value, not fixing broken scrapers. - Scale On Demand: Scrape one site or a thousand without investing in new hardware. ## 5\. DIY Scrapers Often Deliver Inaccurate or Incomplete Data When your business decisions depend on scraped data, accuracy matters. Unfortunately, DIY tools often silently fail, collect partial data, or capture the wrong information altogether. ### Symptoms of Bad Scraping: Stale Data: No scheduler = outdated content Broken HTML: A small UI change can crash your parser Inconsistencies: Multiple formats for the same field (e.g., price with and without tax) Missed Data: JavaScript-rendered content often gets skipped without proper rendering engines Bad data leads to bad decisions. In regulated industries, it can also lead to legal consequences. "The cost of making a bad decision from bad data is often 10x the cost of good data." ## Why Outsourcing Web Scraping Is the Better Option Working with a professional data extraction provider gives you peace of mind and a competitive advantage. ![DIY Web scraping vs professional services ](https://www.blog.datahut.co/content/images/2026/07/img-379.png.webp) ### What You Get: Legally Compliant Processes: Vendors stay up-to-date on regional regulations and adapt accordingly. Technical Resilience: Teams maintain scraper health, detect breakages, and fix them before you notice. Scalable Infrastructure: Whether you want 10 records a week or 10 million a day, capacity isn’t a problem. Flexible Delivery: Choose JSON, CSV, dashboards, or API delivery tailored to your workflow. ### Real-World Example: A leading consumer electronics brand outsourced scraping of 100+ e-commerce websites. The result? - Pricing decisions sped up by 3x - Ad campaigns timed around competitor stockouts - A 21% increase in campaign ROI over a quarter ## Final Thoughts: Focus on Insights, Not Infrastructure Your job isn’t to build scrapers. Your job is to make smarter business decisions with better data. The DIY route may seem appealing for a quick win, but it comes with too much technical debt, legal ambiguity, and opportunity cost. A professional scraping partner is like having an elite data engineering team on call—without the hiring headaches. ### Key Takeaways: - Building and maintaining scrapers is a full-time engineering function - DIY puts you at risk of violating laws you may not even be aware of - Data quality is as important as data quantity - Outsourcing saves time, reduces cost, and improves outcomes ## Next Steps: Scrape Smarter with [Datahut](https://datahut.co/?ref=blog.datahut.co) At Datahut, we help businesses extract high-quality, ready-to-use data from websites across the globe. From product listings and pricing to real estate and financial data—our team handles the complexity so you can focus on growth. Want to see how it works? Schedule a call or explore our blog for real-world use cases. Extract better. Grow faster. With [Datahut](https://blog.datahut.co/?ref=blog.datahut.co). ### FAQs 1\. Why is DIY web scraping not recommended for businesses? DIY web scraping often leads to technical debt, frequent script breakages, legal risks, and unreliable data. Without a dedicated team, it’s hard to manage evolving websites and anti-bot mechanisms. 2\. What are the legal risks of DIY web scraping? DIY scraping may violate website terms of service, data privacy laws like GDPR and CCPA, and can even trigger legal actions under the Computer Fraud and Abuse Act (CFAA). 3\. What are the hidden costs of building your own scrapers? DIY scraping incurs high developer costs, infrastructure expenses (like proxies and servers), and leads to delays in decision-making due to poor data reliability and constant maintenance needs. 4\. Can DIY scrapers collect complete and accurate data? Often, no. DIY tools frequently miss dynamic content, fail silently, or return inconsistent results—leading to poor business decisions based on flawed or incomplete datasets. 5\. What are the benefits of outsourcing web scraping? Outsourcing provides legally compliant processes, robust technical infrastructure, accurate data delivery, faster setup, and allows your team to focus on strategic insights instead of technical issues. ### How to Extract Hotel Data from Booking.com URL: https://www.blog.datahut.co/post/how-to-scrape-data-from-booking-com/ Last updated: 2026-09-07T09:44:19.000Z Have you ever wondered how websites like [Booking.com](https://www.booking.com/?ref=blog.datahut.co) show so many hotel details so quickly? What if you wanted to collect that information yourself—automatically? In this post, I’ll walk you through a real project where we scraped hotel details from [Booking.com](https://www.booking.com/?ref=blog.datahut.co), focusing on stays in San Francisco. [Web scraping ](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/)is a method used to collect data from websites. Think of it like copying information by hand from a web page only much faster and smarter, because a computer does it for you. This can be incredibly useful when you want to gather data that isn’t easily available in a downloadable format. For our project, we chose Booking because it's one of the largest travel sites, with listings from all around the world. It offers tons of useful information: hotel names, prices, locations, user reviews, and more. Our goal is to collect this data in a clean, organized format that we can later analyze or use for research. Here’s how we approached it. First, we wrote a script to go through Booking’s search results for San Francisco and collect the links to individual hotel pages. Since the website doesn’t show all results on one page, we also handled the “Next” buttons—this process is called pagination. Once we had the list of hotel URLs, we moved on to the second part: visiting each hotel’s page and gathering details like name, address, price, room types, amenities, and user ratings. To make all this work smoothly, we used a few handy tools. Playwright helped us load the website like a real user would, which is especially useful when pages rely on JavaScript to display content. Then we used Beautiful Soup to read the underlying HTML and pull out just the parts we needed. Finally, we saved all our collected data in a SQLite database, which is a simple and lightweight way to store information. By the end of this project, we’ll have turned a bunch of regular web pages into a neat, structured dataset—ready for analysis. In the next sections, I’ll guide you step-by-step through the code and logic behind each part, so you can build your own scraper even if you're just starting out. ## Url Collection Phase Now that we’ve set the stage, let’s talk about what this scraper is actually built to do. This project isn’t just about scraping hotel data for one specific day. Instead, it’s designed to work across multiple date ranges. That means it can automatically go through Booking's search results for different check-in and check-out dates and collect hotel page links for each of those time periods. So, imagine you want to see which hotels are available in San Francisco from December 23rd to December 24th, then again from the 24th to the 25th, and so on for a whole week. Doing that manually would take a lot of clicking, copying, and pasting. But with this solution, the whole process is automated. You just feed in the dates and destination, and the script takes care of everything—opening the search pages, scrolling through the results, and collecting the links to individual hotel pages. This step is important because each hotel’s availability and pricing might change depending on the dates. By capturing links across different days, we set ourselves up to gather more accurate and complete data in the next phase. In the next section, we’ll look at how we actually write the code to do this—step by step. ### Import and Logging Setup ``` from playwright.sync_api import sync_playwright import sqlite3 import logging from typing import List, Tuple # Setup logging for better visibility logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') ``` To get started with building our scraper, we first bring in some important tools—what we call libraries in Python. These are like ready-made toolkits that help us do specific tasks without having to write everything from scratch. The first one we use is playwright.sync\_api. This is the key to our automation—it lets our script behave like a real user visiting the website in a browser. It can click buttons, scroll through pages, and wait for content to load, just like you would. Next, we have sqlite3, which helps us save the data we collect. Think of it as a tiny, portable database that lives inside a file on your computer. It's perfect for projects like this where we need to store lots of hotel links in an organized way. We also import logging. This might sound a bit boring, but it's actually very helpful. Logging lets us keep a record of what’s happening while the script runs—like when it starts, what it’s doing, and if anything goes wrong. Each message includes a timestamp so we know exactly when things occurred, which makes troubleshooting much easier later. Finally, we use typing. This isn’t something the script needs to run, but it’s useful for making our code cleaner and easier to understand. It lets us label the kind of data we expect in different parts of the script, which helps avoid mistakes. Once we’ve imported everything, we set up the logging system. This makes sure that all messages—like successes or errors—are printed with the date and time, so we always have a clear record of what happened while our scraper was running. ### \`scrape\_booking\_links\` Function: The Core Scraping Mechanism ``` def scrape_booking_links(url: str, checkin: str, checkout: str) -> List[Tuple[str, str, str]]: """ Scrape hotel links from Booking.com search results for a specific date range. This function uses Playwright to: 1. Navigate to the Booking.com search results page 2. Scroll and load more results 3. Extract unique hotel page links Args: url (str): Full Booking.com search results URL checkin (str): Check-in date in format 'DD/MM/YY' checkout (str): Check-out date in format 'DD/MM/YY' Returns: List[Tuple[str, str, str]]: A list of tuples containing: - Hotel page URL - Check-in date - Check-out date Raises: Exception: For navigation, scraping, or browser-related errors Notes: - Uses Chromium in non-headless mode for better debugging - Implements scroll and "Load more" strategies to capture more results - Adds user agent to mimic real browser behavior """ links = [] try: with sync_playwright() as p: # Launch browser with more stable settings browser = p.chromium.launch( headless=False, # Add more browser launch options for stability args=[ '--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage' ] ) context = browser.new_context( # Add user agent to mimic a real browser user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" ) page = context.new_page() try: # Navigate to the page with error handling try: page.goto(url, timeout=30000, wait_until='networkidle') except Exception as nav_error: logging.error(f"Navigation error: {nav_error}") return links # Max attempts to load more results max_scroll_attempts = 5 scroll_attempts = 0 while scroll_attempts < max_scroll_attempts: # Scroll to bottom and top to trigger lazy loading page.evaluate("window.scrollTo(0, document.body.scrollHeight)") page.wait_for_timeout(2000) page.evaluate("window.scrollTo(0, 0)") page.wait_for_timeout(1000) page.evaluate("window.scrollTo(0, document.body.scrollHeight)") page.wait_for_timeout(2000) # Try to find and click "Load more results" button try: load_more_button = page.locator("text='Load more results'").first if load_more_button.is_visible(): load_more_button.click() page.wait_for_timeout(3000) scroll_attempts = 0 # Reset attempts if button was clicked else: scroll_attempts += 1 except Exception as load_more_error: logging.info(f"No more 'Load more' button or error: {load_more_error}") scroll_attempts += 1 # Extract hotel links hotel_elements = page.query_selector_all( "h3.aab71f8e4e > a.a78ca197d0" ) # Capture unique links unique_links = set() for element in hotel_elements: link = element.get_attribute('href') if link and link not in unique_links: # Ensure full URL if not link.startswith('http'): link = f"https://www.booking.com{link}" unique_links.add(link) links.append((link, checkin, checkout)) logging.info(f"Scraped {len(links)} unique hotel links") except Exception as e: logging.error(f"Scraping error: {e}", exc_info=True) finally: # Ensure browser resources are closed try: page.close() context.close() browser.close() except Exception as close_error: logging.error(f"Error closing browser resources: {close_error}") except Exception as setup_error: logging.error(f"Playwright setup error: {setup_error}", exc_info=True) return links ``` Now we come to the heart of the script—a function called scrape\_booking\_links. This part does the heavy lifting. It’s in charge of visiting a Booking.com search results page and collecting all the individual hotel links you see listed there. This function needs three things to get started: 1. The URL of the search results page 2. the check-in date, 3. and the check-out date. With these inputs, it opens the Booking.com page using a tool called Playwright , which lets our script control a browser just like a human would. We specifically use the Chromium browser (a cousin of Google Chrome) for this task. To make the browser behave more like a real user, we add a custom user agent —this is like a little ID card that tells websites what kind of device or browser is visiting. By doing this, we reduce the chance of our scraper being blocked or flagged as a bot. Once the page opens, we use a smart trick called scrolling . Booking.com, like many modern websites, doesn’t load everything at once. It loads more results only as you scroll down—this is known as lazy loading. So, the function scrolls down to the bottom of the page, waits a bit, then scrolls back up and repeats. This pattern encourages the site to load more and more hotel listings each time. Sometimes, there’s also a “Load more results” button on the page. The function tries to click that button too—several times if needed—but not forever. We set a limit on how many times it can try, just in case something goes wrong or the site behaves differently. After all the visible hotels are loaded, we extract the links using something called CSS selectors . These are like address labels for elements on the page—in this case, for hotel cards that contain the URLs we want. As the links are collected, we make sure there are no duplicates by using a set (a type of list that only keeps unique items). Finally, each hotel link is saved along with the date range it belongs to. This way, we know not just which hotel was listed, but when it was available. ### \`save\_links\_to\_db\` Function: Persistent Storage ``` def save_links_to_db(links: List[Tuple[str, str, str]], db_name: str = "booking_links.db"): """ Save scraped hotel links to a SQLite database with duplicate prevention. This function: 1. Creates a SQLite database if not exists 2. Creates a 'links' table to store unique hotel URLs 3. Inserts new links with their associated dates 4. Prevents duplicate entries Args: links (List[Tuple[str, str, str]]): List of tuples containing: - Hotel page URL - Check-in date - Check-out date db_name (str, optional): Name of the SQLite database file. Defaults to "booking_links.db". Returns: None Raises: sqlite3.Error: For database connection or insertion errors Notes: - Uses INSERT OR IGNORE to prevent duplicate links - Logs the number of newly inserted unique links """ try: conn = sqlite3.connect(db_name) cursor = conn.cursor() # Create table with unique constraint cursor.execute(''' CREATE TABLE IF NOT EXISTS links ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, checkin_date TEXT, checkout_date TEXT ) ''') # Use INSERT OR IGNORE to prevent duplicate entries for link, checkin, checkout in links: try: cursor.execute( "INSERT INTO links (url, checkin_date, checkout_date) VALUES (?, ?, ?)", (link, checkin, checkout) ) except sqlite3.Error as insert_error: logging.error(f"Error inserting link: {insert_error}") # Commit changes and get number of inserted rows conn.commit() inserted_count = conn.total_changes logging.info(f"Inserted {inserted_count} new unique links") except sqlite3.Error as db_error: logging.error(f"SQLite Error: {db_error}", exc_info=True) finally: if conn: conn.close() ``` Once we've collected the hotel links from the website, the next step is to save them somewhere safe. That’s where the save\_links\_to\_db function comes in. This function takes all the links we just scraped and stores them in a small, local database using SQLite . If you’re not familiar with SQLite, think of it like a digital notebook—it lets you store and organize data in tables, just like a spreadsheet, but it works right inside your Python project. Here’s how it works: if the database file doesn’t already exist, the function will create one. Inside it, it sets up a table called hotel\_links. This table includes a few key pieces of information for each hotel link: - a unique ID number, - the actual URL of the hotel page, - the check-in date, - and the check-out date. When it comes time to insert the links, the function uses a smart method called “INSERT OR IGNORE.” This means that if a link has already been saved before, the function won’t insert it again. This helps us avoid having the same hotel link appear multiple times—even if we run the scraper again for the same dates. After trying to insert all the links, the function counts how many new, unique ones were added to the database. It then logs that number so you can easily keep track of what’s being stored. Just in case something goes wrong—like a database error—it also includes a safety check. If there’s a problem during the saving process, it will record an error message in the logs instead of crashing the entire script. This makes the function reliable and keeps the data clean and organized. ### \`main()\` Function: Orchestrating the Scraping Process ``` def main(): """ Main execution function for Booking.com link scraping process. This function: 1. Defines a list of Booking.com search result URLs for different dates 2. Iterates through URLs to scrape hotel links 3. Collects all unique links across different date ranges 4. Saves collected links to a SQLite database Args: None Returns: None Notes: - Handles exceptions during scraping for individual URLs - Logs total number of links scraped - Warns if no links were scraped """ urls_with_dates = [ ("https://www.booking.com/searchresults.en-gb.html?ss=San+Francisco&ssne=San+Francisco&ssne_untouched=San+Francisco&label=gen173nr-1BCAEoggI46AdIM1gEaGyIAQGYAQm4ARnIAQzYAQHoAQGIAgGoAgO4AqO5yroGwAIB0gIkMzE4NDYxMTYtNjRlMC00NzQ0LWFhNGYtYzU2YmI4Y2FkMjUw2AIF4AIB&sid=74da2c31c035c6df8deb313a85e24f8e&aid=304142&lang=en-gb&sb=1&src_elem=sb&src=index&dest_id=20015732&dest_type=city&checkin=2024-12-23&checkout=2024-12-24&group_adults=2&no_rooms=1&group_children=0", "23/12/24", "24/12/24"), # add more links here.. ] # Total links across all scraping attempts all_links = [] for url, checkin, checkout in urls_with_dates: logging.info(f"Scraping: {url}") try: links = scrape_booking_links(url, checkin, checkout) all_links.extend(links) except Exception as e: logging.error(f"Error scraping {url}: {e}") if all_links: logging.info(f"Total links scraped: {len(all_links)}") save_links_to_db(all_links) else: logging.warning("No links were scraped.") ``` At the center of everything is the main() function. You can think of it as the conductor of the script—making sure each part plays its role at the right time to complete the full scraping process smoothly. This function starts by creating a list of search result URLs from Booking, each with different check-in and check-out dates. By covering several date ranges, we make sure our scraper doesn’t miss any hotels that might only appear on specific days. This is especially useful if hotel availability changes often. Once the list is ready, the script goes through each URL one by one. For every URL, it calls the scrape\_booking\_links function—the one we discussed earlier—to collect hotel page links. All the links found during each run are added to a single, growing list that holds everything we've gathered so far. Of course, sometimes things don’t go perfectly. A page might not load correctly, or Booking.com might behave unexpectedly. That’s okay—this function is built to handle such moments gracefully. If there’s an error while working on one of the URLs, it doesn’t stop everything. Instead, it logs the error (so you’ll know what happened), and then it moves on to the next link in the list. This way, one problem doesn’t ruin the entire run. After the script finishes checking all the URLs, it looks at what it collected. If we ended up with any hotel links, it sends them over to the save\_links\_to\_db function to be safely stored in the database. But if no links were found—maybe because of a technical issue or no availability—it logs a message to let you know. ### Execution Flow ``` # Run the main function if __name__ == "__main__": main() ``` If you’re new to Python, this might look a bit strange—but here’s what it does in simple terms: it tells Python, “Only run the main() function if this file is being run directly.” That means, when you open a terminal and run the script, this condition becomes true, and main() is called. This single line is what kicks off the entire scraping process. It launches the browser, navigates through all the Booking com pages you’ve specified, scrolls and clicks to load hotel listings, collects the hotel links, and finally saves them neatly into your database. Without this part, the script would just sit there, defining functions but never actually doing anything. So, think of it as the green light—the command that sets everything into motion. ## Data Collection Phase Now that we’ve gathered hotel links and saved them in a database, the next phase is all about digging deeper—visiting each individual hotel page and collecting detailed information. This is where our second scraper comes in. It’s a bit more advanced, designed to go through each hotel link stored in the SQLite database, one by one. Using Playwright, the scraper opens each hotel page, waits for the content to fully load (including anything built with JavaScript), and then begins carefully pulling out the key details. Each of these elements gives us a clearer picture of what the hotel offers and what kind of experience guests might expect. But websites like Booking don’t always present data in the same way on every page. Some hotel pages might be missing certain sections—like pricing or a detailed description. That’s why the scraper is built to handle these situations gracefully. If any piece of information isn’t found on a particular page, it simply marks it as “Not Available” and moves on. This makes the tool reliable and robust, even when the website content isn’t perfectly consistent. Another smart feature is the tracking mechanism. Once the scraper finishes collecting data from a link, it marks that link as “scraped” in the database. This means if the script stops or crashes partway through, it can resume later from where it left off—without starting over or re-scraping the same pages. It’s an efficient way to manage large-scale data extraction. So, in short, this scraper doesn't just grab links—it explores each one thoroughly, handles missing content with care, and keeps track of its progress. It’s a powerful way to turn Booking.com’s hotel pages into clean, structured data you can actually work with. ### Import Section ``` import sqlite3 import asyncio from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` As usual, the script starts by importing some built-in Python modules: sqlite3, asyncio, playwright and [beautifulsoup](https://www.crummy.com/software/BeautifulSoup/bs4/doc/?ref=blog.datahut.co). ### connect\_to\_database(db\_name) Function ``` def connect_to_database(db_name): """ Establish a connection to the SQLite database. This function creates a connection to the specified SQLite database. If the database doesn't exist, it will be created automatically. Args: db_name (str): The name or path of the SQLite database file. Returns: sqlite3.Connection: A connection object to the specified database. """ return sqlite3.connect(db_name) ``` Every good scraping project needs a place to store the data, and for that, we use a database. But before we can start saving anything, we need to set up a connection between our script and the database file. That’s exactly what the connect\_to\_database function does. You can think of this function as the first handshake—it opens the line of communication between your scraper and the database where all the hotel information will be stored. The best part? It’s completely automatic. You don’t need to create the database manually or set anything up in advance. The function checks whether the database file already exists: - If it does, it simply connects to it. - If it doesn’t, SQLite quietly creates a brand-new file on the spot and gets it ready to use. This smart design keeps things simple and avoids unnecessary setup steps. As soon as the connection is made, the rest of the scraper can start inserting hotel links and detailed data without worrying about whether the storage space is ready. So in short, connect\_to\_database is the gateway to saving and managing your scraped data. It’s a small piece of code that plays a big role in keeping your workflow organized and reliable. ### check\_and\_add\_scraped\_column(conn) Function ``` def check_and_add_scraped_column(conn): """ Verify and add a 'scraped' column to the 'links' table if it doesn't exist. This function checks the structure of the 'links' table and adds a 'scraped' column with a default value of 0 if it's not already present. This column is used to track which links have been processed during scraping. Args: conn (sqlite3.Connection): An active database connection. """ cursor = conn.cursor() cursor.execute("PRAGMA table_info(links);") columns = [col[1] for col in cursor.fetchall()] if 'scraped' not in columns: cursor.execute("ALTER TABLE links ADD COLUMN scraped INTEGER DEFAULT 0;") conn.commit() ``` Another important part of this scraper is keeping track of which hotel links have already been processed—and which ones are still waiting to be scraped. That’s where the check\_and\_add\_scraped\_column function comes in. You can think of this function as a helper that keeps our database organized and ready. It looks into the database—specifically the links table—and checks whether there’s a column called scraped. This column acts like a status tag for each link: if a link hasn’t been scraped yet, its value is 0; once it’s been processed, the value is updated to 1. The smart part is how this function works. Instead of assuming the column already exists, it inspects the table dynamically. In other words, it checks what columns are actually present in the database at that moment. If the scraped column is missing, the function adds it automatically—no need for you to do anything manually. By doing this, the script becomes much more flexible and reliable, even if the database structure wasn’t perfectly set up in advance. It adapts on its own and ensures that every hotel link can be marked with a clear status. This simple check gives us a clean and effective way to track progress, avoid re-scraping the same pages, and pick up right where we left off if the script is ever interrupted. ### get\_unscraped\_links(conn) Function ``` def get_unscraped_links(conn): """ Retrieve all links that have not yet been scraped from the database. This function queries the 'links' table to fetch links where the 'scraped' column is set to 0, along with their associated check-in and check-out dates. Args: conn (sqlite3.Connection): An active database connection. Returns: list: A list of tuples containing (link_id, url, checkin_date, checkout_date) for unscraped links. """ cursor = conn.cursor() cursor.execute("SELECT id, url, checkin_date, checkout_date FROM links WHERE scraped = 0;") return cursor.fetchall() ``` Once we’ve marked which links have been scraped and which haven’t, the next step is to pull out only the ones that still need to be processed. That’s exactly what the get\_unscraped\_links function does. Think of this function as building a to-do list for the scraper. It looks inside the database, finds all the hotel links that still have a scraped value of 0, and prepares them for the next round of data collection. But it doesn’t just fetch the link alone—it also brings along the check-in and check-out dates tied to each hotel search. This extra bit of information helps the scraper stay smart about context. For example, it knows not just which hotel page to visit, but also which specific dates were used during the search. That kind of detail can be important when displaying prices or room availability. The result of this function is a list of tuples—each one holding a hotel link along with its associated dates. This list becomes the input for the asynchronous scraping engine, which means the scraper can begin working through the queue, one hotel at a time, without repeating any links or missing key information. ### mark\_as\_scraped(conn, link\_id) Function ``` def mark_as_scraped(conn, link_id): """ Update the status of a specific link to indicate it has been scraped. This function sets the 'scraped' column to 1 for the given link ID, marking it as processed in the database. Args: conn (sqlite3.Connection): An active database connection. link_id (int): The unique identifier of the link to be marked as scraped. """ cursor = conn.cursor() cursor.execute("UPDATE links SET scraped = 1 WHERE id = ?;", (link_id,)) conn.commit() ``` Once the scraper finishes gathering data from a hotel page, it needs a way to mark that job as done—and that’s exactly what the mark\_as\_scraped function does. This function updates the database to say, “Hey, we’ve already scraped this link.” It does this by setting the scraped column value to 1 for that specific hotel link. That little update may seem simple, but it plays a huge role in keeping the entire scraping process organized and efficient. Why is this important? Because scraping large websites like Booking.com can take time. Sometimes, the script might stop halfway through due to a network issue, a system reboot, or any number of unexpected reasons. Without this tracking system, you’d either have to start over or risk scraping the same hotel pages again. But thanks to mark\_as\_scraped, the scraper knows exactly where it left off. When you run the script again, it will skip over links that have already been processed and continue with the rest—making the scraping workflow resumable and much more reliable. ### save\_scraped\_data(conn, data) Function ``` def save_scraped_data(conn, data): """ Save the scraped hotel information to the database. This function creates a 'scraped_data' table if it doesn't exist and inserts the scraped information for a specific hotel, including details like title, location, description, pricing, and reviews. Args: conn (sqlite3.Connection): An active database connection. data (tuple): A tuple containing scraped hotel information in the order: (url, checkin_date, checkout_date, title, location, location_score, description, facilities, room_type, price, sale_price, tax_amount, rating, rating_score, review_count) """ cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS scraped_data ( id INTEGER PRIMARY KEY, url TEXT, checkin_date TEXT, checkout_date TEXT, title TEXT, location TEXT, location_score TEXT, description TEXT, facilities TEXT, room_type TEXT, price TEXT, sale_price TEXT, tax_amount TEXT, rating TEXT, rating_score TEXT, review_count TEXT ); """) cursor.execute(""" INSERT INTO scraped_data ( url, checkin_date, checkout_date, title, location, location_score, description, facilities, room_type, price, sale_price, tax_amount, rating, rating_score, review_count ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?); """, data) conn.commit() ``` Once all the data has been scraped from a hotel page—names, locations, prices, reviews, and more—it needs to be saved in a safe and structured way. That’s where the save\_scraped\_data function steps in. It’s the final step in our scraping journey, and it plays a key role in making sure nothing gets lost. This function creates a new table in the database called scraped\_data, if it doesn’t already exist. The table is designed to hold all the rich details we collected from each hotel page—things like hotel names, addresses, descriptions, amenities, room types, prices, and user reviews. But what makes this function especially useful is its flexibility. It builds the table programmatically, which means the scraper can handle changes in the hotel page layout without needing you to manually update the database structure. If tomorrow they adds a new field or moves things around, the scraper can still adapt and store the data properly. With every successful scrape, a new entry is added to the database, turning raw webpage content into organized, usable information. Over time, this builds a valuable dataset that can be used for data analysis, trend tracking, competitor research, or business decisions. ### \`parse\_title(soup)\` Function ``` def parse_title(soup): """ Extract the hotel title from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: The hotel title, or "Not Available" if not found. """ title = soup.select_one("#hp_hotel_name > div > h2") return title.get_text(strip=True) if title else "Not Available" ``` This function is in charge of getting the hotel’s name from the webpage. It works with the parsed HTML content using BeautifulSoup, a tool that helps us navigate and extract data from web pages. The function looks for the part of the page where the hotel’s name is usually placed—typically inside a specific HTML tag. It does this using a CSS selector, like pointing a finger at the exact spot where the name is expected to be. But here’s the smart part: websites sometimes change. A hotel name might be missing or placed somewhere unexpected. Instead of crashing or giving an error, this function gracefully handles the situation. If it can’t find the name, it simply returns “Not Available.” That small fallback keeps the scraper running smoothly, even if the page doesn’t look exactly the same every time. It’s an example of resilient scraping—where the goal is to collect as much data as possible, without breaking when something is missing or slightly different. Similarly,below given the parsing functions for location, location score, description, room type, price details and ratings. ### \`parse\_location(soup)\` Function ``` def parse_location(soup): """ Extract the hotel location from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: The hotel location, or "Not Available" if not found. """ location = soup.select_one("#wrap-hotelpage-top > div:nth-child(4) > div > div > span:nth-child(2) > div") return location.get_text(strip=True) if location else "Not Available" ``` ### \`parse\_loc\_score(soup)\` Function ``` def parse_loc_score(soup): """ Extract the location score from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: The location score, or "not available" if not found. """ loc_score=soup.select_one("#reviewFloater > div.best-review-score.best-review-score-with_best_ugc_highlight.hp_lightbox_score_block > span > span") return loc_score.get_text(strip=True) if loc_score else "not available" ``` ### \`parse\_description(soup)\` Function ``` def parse_description(soup): """ Extract the hotel description from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: The hotel description, or "Not Available" if not found. """ description = soup.select_one("#basiclayout > div.hotelchars > div.page-section.hp--desc_highlights.js-k2-hp--block > div > div.bui-grid__column.bui-grid__column-8.k2-hp--description > div.hp-description > div.hp_desc_main_content > div > div > p.a53cbfa6de.b3efd73f69") return description.get_text(strip=True) if description else "Not Available" ``` ### \`parse\_roomtype(soup)\` Function ``` def parse_roomtype(soup): """ Extract the room type from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: The room type, or "not available" if not found. """ room=soup.select_one("a.hprt-roomtype-link > span.hprt-roomtype-icon-link ") return room.get_text(strip=True) if room else "not available" ``` ### \`parse\_prices(soup)\` Function ``` def parse_prices(soup): """ Extract price-related information from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: tuple: A tuple containing (original_price, sale_price, tax_amount), with "not available" used if any information is missing. ""” price=soup.select_one("div.bui-f-color-destructive.js-strikethrough-price.prco-inline-block-maker-helper.bui-price-display__original") sale_price=soup.select_one("span.prco-valign-middle-helper") tax_amount=soup.select_one("div.prd-taxes-and-fees-under-price") return price.get_text(strip=True) if price else "not available",sale_price.get_text(strip=True) if sale_price else "not available",tax_amount.get_text(strip=True) if tax_amount else "not available" ``` ### \`parse\_ratings(soup)\` Function ``` def parse_ratings(soup): """ Extract rating-related information from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: tuple: A tuple containing (rating, rating_score, review_count), with "not available" used if any information is missing. """ rating=soup.select_one("span.a3b8729ab1.e6208ee469.cb2cbb3ccb") rating_score=soup.select_one("div.a3b8729ab1.d86cee9b25") review_count=soup.select_one("span.a3b8729ab1.f45d8e4c32.d935416c47") return rating.get_text(strip=True) if rating else "not available",rating_score.get_text(strip=True) if rating_score else "not available",review_count.get_text(strip=True) if review_count else "not available" ``` ### \`parse\_facilities(soup)\` Function ``` def parse_facilities(soup):   """ Extract hotel facilities from the BeautifulSoup parsed page. Args: soup (BeautifulSoup): The parsed HTML content of the hotel page. Returns: str: A comma-separated string of hotel facilities, or an empty string if none found. """ facilities = [] items = soup.select("#basiclayout > div.hotelchars > div.page-section.hp--desc_highlights.js-k2-hp--block > div > div.bui-grid__column.bui-grid__column-8.k2-hp--description > div.hp--popular_facilities.js-k2-hp--block > div:nth-child(2) > div > div > ul > li") for item in items: facilities.append(item.get_text(strip=True)) return ", ".join(facilities) ``` This function is responsible for gathering the list of facilities or amenities that a hotel offers—things like Wi-Fi, gym access, parking, or air conditioning. It works by going through the HTML content of the page and looking for the list items that usually contain each facility. Using BeautifulSoup, it loops through each of these items and pulls out the text—one by one. Once all the facilities are collected, the function joins them into a single string, separating each item with a comma. This makes the final result easy to read and store—like a quick summary of everything the hotel provides. What’s especially useful about this approach is that it can handle different facility setups. Whether a hotel lists five items or fifteen, the function adjusts accordingly. It turns a messy block of HTML into a clean, readable string that tells you what to expect from your stay. ### fetch\_page\_content(url, page) Function ``` async def fetch_page_content(url, page): """ Fetch the HTML content of a webpage using Playwright. This async function navigates to the specified URL, scrolls to the bottom of the page to trigger any lazy-loaded content, and waits for a short time to ensure page rendering is complete. Args: url (str): The URL of the webpage to scrape. page (playwright.async_api.Page): An active Playwright page instance. Returns: str: The fully rendered HTML content of the page. """ await page.goto(url) await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await page.wait_for_timeout(12000) return await page.content() ``` The fetch\_page\_content function is like the scraper’s eyes and patience—it doesn’t just open a webpage, it makes sure everything is fully loaded and ready before collecting any data. Modern websites like Booking don’t load all their content in one go. Instead, they load things as you scroll or as needed, using JavaScript. That’s why this function uses Playwright, a powerful tool that controls a real browser, to visit the hotel’s page just like a human would. Once on the page, the function scrolls to the bottom and then waits. This is a smart trick. The act of scrolling triggers the website to load more information—like photos, room types, or reviews—that wouldn’t appear right away. The wait gives it enough time to finish loading all that data. By the time the function captures the page content, it’s not just a half-loaded snapshot—it’s the fully expanded version of the hotel page, with everything visible and ready to be parsed. This way, the scraper doesn’t miss out on anything important that might have been hidden during the first few seconds of loading. ### parse\_content(content) Function ``` def parse_content(content): """ Parse the HTML content and extract all relevant hotel information. This function uses BeautifulSoup to parse the HTML and calls multiple parsing functions to extract different pieces of hotel information. Args: content (str): The HTML content of the hotel page. Returns: tuple: A tuple containing extracted hotel details in the order: (title, location, location_score, description, facilities, room_type, price, sale_price, tax_amount, rating, rating_score, review_count) """ soup = BeautifulSoup(content, 'html.parser') title = parse_title(soup) location = parse_location(soup) location_score=parse_loc_score(soup) description = parse_description(soup) facilities = parse_facilities(soup) room_type=parse_roomtype(soup) price=parse_prices(soup)[0] sale_price=parse_prices(soup)[1] tax_amount=parse_prices(soup)[2] rating=parse_ratings(soup)[0] rating_score=parse_ratings(soup)[1] review_count=parse_ratings(soup)[2] return title, location,location_score, description, facilities,room_type,price,sale_price,tax_amount,rating,rating_score,review_count ``` At the heart of the Booking.com Hotel Scraper is the parse\_content() function. Think of this as the brain of the operation—the place where raw, messy webpage content gets transformed into clean, structured data you can actually use. Here’s how it works: once the scraper fetches the HTML of a hotel’s page, parse\_content() steps in to make sense of it. It uses BeautifulSoup, a library that helps us dig into the webpage and pick out just the parts we need—like the hotel’s name, location, room details, prices, and more. But instead of trying to extract everything at once, the function calls on a team of helper functions—each designed to grab one specific piece of data. For example, parse\_title() gets the hotel’s name, parse\_location() finds the address, and other functions handle things like facilities, reviews, and pricing. Each of these helpers knows exactly where to look and what to collect. Once all the pieces are gathered, parse\_content() combines them into a single, organized bundle—what we call a tuple. This tuple represents a complete snapshot of the hotel’s information and is ready to be saved into the database. What makes this function especially powerful is its flexibility and resilience. If a hotel page is missing something—say the location score or the description—the function won’t break. Instead, it fills in a placeholder like "Not Available", so the scraper keeps running smoothly. This way, even if some hotel pages are a bit messy or incomplete, the scraper still pulls out as much information as possible. In the end, parse\_content() ensures that every hotel we scrape ends up as a well-structured, information-rich record. And since each part of the function is modular, it’s easy to update or improve one piece without affecting the whole workflow. ### scrape\_links(db\_name) Function ``` async def scrape_links(db_name): """ Main asynchronous function to scrape hotel links from the database. This function performs the following steps: 1. Connect to the database 2. Ensure the 'scraped' column exists 3. Retrieve unscraped links 4. Launch a Playwright browser 5. Iterate through links, scraping and saving data 6. Mark each link as scraped upon successful scraping Args: db_name (str): The name of the SQLite database containing links to scrape. """ conn = connect_to_database(db_name) check_and_add_scraped_column(conn) unscraped_links = get_unscraped_links(conn) async with async_playwright() as playwright: browser = await playwright.chromium.launch(headless=False) context = await browser.new_context() page = await context.new_page() for link_id, url, checkin_date, checkout_date in unscraped_links: try: print(f"Scraping: {url}") content = await fetch_page_content(url, page) title, location, location_score, description, facilities, room_type, price, sale_price, tax_amount, rating, rating_score, review_count = parse_content(content) # Include checkin and checkout dates in the data to be saved save_scraped_data(conn, ( url, checkin_date, checkout_date, title, location, location_score, description, facilities, room_type, price, sale_price, tax_amount, rating, rating_score, review_count )) mark_as_scraped(conn, link_id) print(f"Successfully scraped: {url}") except Exception as e: print(f"Error scraping {url}: {e}") await browser.close() conn.close() ``` The scrape\_links() function is like the control center of the entire scraping operation. It pulls everything together—connecting to the database, opening the browser, getting the links, and guiding the scraper through each hotel page, step by step. This function is asynchronous, which means it can handle tasks in a non-blocking way. That’s useful for scraping because hotel pages can take time to load, and the scraper needs to wait for them without getting stuck. Here’s what it does: First, it connects to the SQLite database, which holds all the hotel links we plan to scrape. Then, it filters out the ones we've already processed and fetches only the unscraped ones. These are the fresh pages we still need to explore. After gathering the list, the function launches a Playwright browser—a tool that lets our script browse the web just like a person would. Then, for each hotel link, it loads the page, extracts the data, saves the result, and finally updates the database to mark that hotel as "scraped." What’s impressive about this function is how it handles errors gracefully. Websites don’t always behave the way we expect—some links may not load, others may be missing data, or the internet connection could drop. If something goes wrong with one hotel page, scrape\_links() doesn’t give up or crash. Instead, it logs the issue and moves on to the next one. This thoughtful design makes the scraper strong and reliable, ready to work through hundreds of hotel pages without breaking. It strikes the right balance between being thorough (by checking every hotel) and being resilient (by continuing even when a few pages have problems). In short, scrape\_links() keeps the entire process moving smoothly and efficiently. ### Execution Flow ``` # Run the script if __name__ == "__main__": asyncio.run(scrape_links("booking_links.db")) ``` The script is designed to be run directly, with the if name == "\_\_main\_\_": block triggering the asynchronous scraping process on a specified database of hotel links. ## Conclusion Throughout this project, we saw how [web scraping](https://www.blog.datahut.co/post/how-to-choose-the-best-web-scraping-service/) when done with a clear plan and the right tools, can turn messy website content into clean, usable data. By collecting hotel details, cleaning the results using OpenRefine, and storing everything neatly in an SQLite database, we transformed scattered web pages into structured, meaningful information ready for analysis. This journey shows that web scraping doesn’t have to be complicated. With some practice and the right mindset, anyone can learn to gather valuable data from websites. Projects like this are a great way to build confidence, sharpen your skills, and see real results. Over time, you’ll find yourself able to take on more advanced scraping tasks, build your own tools, or even feed this data into larger projects whether it’s for research, analytics, or app development. Contact [datahut](https://www.datahut.co/?ref=blog.datahut.co) for all your web scraping needs! AUTHOR I’m Shahana, a Data Engineer at Datahut, where I help turn complex, unstructured web data into organized, valuable insights—especially for travel, e-commerce, and hospitality sectors. At Datahut, we’ve worked with businesses around the globe to automate data collection from websites that are often difficult to scrape, including those that use JavaScript and dynamic loading. In this blog, I’ll walk you through a real-world project where we scraped hotel listings from Booking using Playwright, BeautifulSoup, and SQLite. The goal was to collect reliable hotel information for San Francisco over a range of dates handling dynamic content, pagination, and data storage, all in a way that’s scalable and beginner-friendly. If your team is looking to collect structured travel or accommodation data at scale—or simply wants to learn how scraping can support smarter market insights—feel free to reach out through the chat widget on the right. We’re always happy to explore solutions that fit your data goals. ### Frequently Asked Questions (FAQ) 1. Is it legal to scrape data from Booking.com?Scraping publicly available data may violate their terms of service, even if it is not always illegal. It’s important to review their robots.txt file and legal terms before scraping. 2. What type of data can I extract from Booking.com?You can extract hotel names, prices, locations, reviews, ratings, availability, room types, and amenities — depending on your scraping setup and compliance approach. 3. Which tools are best for scraping Booking.com?Tools like Selenium, BeautifulSoup, and Scrapy are often used. Selenium is ideal when dynamic content (JavaScript-rendered) needs to be handled. 4. Does Booking.com block web scrapers?Yes, Booking.com actively uses anti-bot mechanisms such as CAPTCHA, rate limiting, and IP blocking. You need to rotate proxies, use user-agents, and throttle requests. 5. Can I scrape Booking.com without coding knowledge?While DIY scraping requires coding, you can use scraping service providers or no-code tools — but always ensure compliance with their terms. ### How Can Web Scraping Help Extract Refrigerator Data from Best Buy ? URL: https://www.blog.datahut.co/post/how-can-web-scraping-help-extract-refrigerator-data-from-best-buy/ Last updated: 2026-07-23T07:48:32.000Z When it comes to electronics in the U.S., Best Buy is one of the biggest and most familiar names. Their online store has a huge range of home appliances, and refrigerators make up a big chunk of that—everything from compact, space-saving models to smart fridges that connect to Wi-Fi and even have screens on the front. At first glance, collecting all this product information might seem simple. You might think, “Just write a script, grab the data, and you're done.” But in reality, it’s not that straightforward. Best Buy’s website uses JavaScript to load a lot of its content, which means the information doesn’t appear right away when you look at the raw page. If you try to scrape it using basic tools, you’ll likely get nothing useful. So, to collect the data properly, we need smarter tools that behave more like a real person browsing the site. That means waiting for things to load, scrolling through pages, and clicking when needed. And here’s an interesting twist—Best Buy has some strong security measures to spot bots. We found that if we let our scraper run with the browser visible—so the browser actually opens up and shows what it’s doing—it had a much better chance of getting through without being blocked. To make things easier, we broke the project into two clear steps. First, we focused on collecting all the links to the individual refrigerator products by scrolling through the listing pages automatically. Then in the second step, we visited each of those product pages one by one to gather more detailed information—like prices, features, and customer reviews. This project is a great example of how web scraping today often means doing more than just grabbing data. It’s about figuring out how a website really works, handling content that changes as you browse, and making your scraper act as naturally as possible. And just as important—it’s about doing all of this in a respectful, thoughtful way that works with the website, not against it. ## URL COLLECTION ### Imports and Initial Setup ``` import asyncio import random from playwright.async_api import async_playwright from bs4 import BeautifulSoup import sqlite3 from datetime import datetime import logging ``` Before we jump into scraping anything, we need to gather all the tools our script will use—just like packing your toolkit before starting a project. This happens in the import section of our code, where we bring in different Python modules. Each one has a job, and together they help us build a smart, smooth, and reliable scraper. Let’s start with Playwright. Think of it like a remote control for your web browser. It can open pages, click buttons, scroll down, and wait for things to load—just like a human would. This is super useful for websites like Best Buy, where much of the content only appears when you interact with the page. Without Playwright, we’d miss a lot of the data we need. Once a page is fully loaded, we pass it to BeautifulSoup. If the web page were a messy room full of random papers, BeautifulSoup would be the person who calmly walks in and picks out the exact note you’re looking for. It helps us dig through all the HTML code and pull out just the useful stuff—like product names, prices, and links—without getting lost in everything else. Then there’s SQLite, our simple way to store data. You can think of it as a built-in spreadsheet that quietly lives on your computer. It doesn’t need a fancy setup or internet connection, but it gives us an easy way to save everything we collect and organize it nicely. It also helps us track which product links we’ve already visited and which ones are still waiting. To speed things up, we use asyncio, a tool that lets our program multitask. It’s like having several tabs open in your browser, all doing different things at once. This way, our scraper can handle more pages in less time. We also include a few helpful sidekicks. The random module lets us add small delays between actions, like waiting a few seconds before visiting the next page. This makes our scraper behave more like a real person, which helps us avoid getting blocked by the website. The datetime module is there to add time stamps—useful for keeping track of when we scraped something or organizing our data by date. And finally, there’s logging. This tool keeps a quiet log of everything that happens while the scraper runs. Whether something was scraped successfully, failed, or skipped, logging writes it all down. That way, if something goes wrong, we can look back and figure out what happened. With all these tools packed and ready, we’re set to start building the heart of our scraper—ready to explore dynamic websites, stay organized, and move smartly like a real user. ### Logging Configuration ``` # Set up logging with timestamp and log level for better debugging logging.basicConfig( level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s' ) ``` The logging.basicConfig() function sets up the basic configuration for logging in your script. level=logging.INFO means the logger will capture all messages that are INFO level or higher (like WARNING, ERROR, etc.). format='%(asctime)s - %(levelname)s - %(message)s' customizes the log message format to include the timestamp, the log level (INFO, WARNING, etc.), and the actual log message. This setup helps you track the flow of your scraper by recording when events happen and how serious they are — making it much easier to troubleshoot issues if something breaks during scraping. ### User Agent Management ``` def load_user_agents(file_path): """ Load user agent strings from a file to rotate during scraping. Args: file_path (str): Path to the text file containing user agent strings Returns: list: List of user agent strings with empty lines removed Note: Each user agent should be on a separate line in the file """ with open(file_path, "r") as f: return [line.strip() for line in f if line.strip()] ``` Every time you open a website, your browser quietly introduces itself by saying, “Hi, I’m Chrome,” or “I’m Firefox on Windows,” or something similar. This little introduction is known as a user agent, and it helps the website understand what kind of browser you're using. But it can also be used to tell whether a visitor is a human or a bot. In our case, since we’re building a scraper, we don’t want the website to immediately recognize us as a bot. So, we use a clever trick—our scraper puts on a disguise. That disguise is a user agent, and we’ve created a function that acts like a “disguise manager.” Imagine having a closet full of outfits. Instead of always wearing the same one, you change your outfit every time you go out, making it harder for someone to recognize you. That’s exactly what our function does—it reads from a file that contains many different browser identities and picks one at random for each visit. It even tidies things up first by removing any blank lines in the file, so we don’t end up using broken or empty disguises. In the end, this simple step makes a big difference in helping our scraper blend in and avoid being blocked. ### Page Scrolling Handler ``` async def scroll_to_bottom(page): """ Scroll to the bottom of the page to trigger lazy loading of products. Args: page: Playwright page object This function implements an infinite scroll detection mechanism by: 1. Getting the current page height 2. Scrolling to bottom 3. Waiting for new content 4. Comparing new height with previous height 5. Breaking if heights are equal (no more content loaded) """ previous_height = await page.evaluate("document.body.scrollHeight") while True: await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await page.wait_for_timeout(2000) # Wait for lazy-loaded content new_height = await page.evaluate("document.body.scrollHeight") if new_height == previous_height: break previous_height = new_height ``` Modern websites have become quite clever—they don’t show you everything right away. Instead, they load more content only when you scroll down, much like how Instagram or Facebook works. This is great for saving bandwidth and improving speed for real users, but it adds a bit of a challenge for web scrapers. To deal with this, we use what’s called a scroll handler. Here’s how it works: first, it checks how tall the page currently is. Then it scrolls down a bit, waits a second or two to let new content load, and checks again to see if the page got taller. If the height doesn’t change after scrolling, it means there’s no more content to load—we’ve reached the bottom. What makes this scroll handler especially useful is its ability to handle delays and timeouts. Sometimes websites are slow to load new items, and a less careful scraper might move on too soon and miss out on data. But this function waits patiently, checking carefully, and only stops once it's sure there’s nothing left to fetch. It’s like having a responsible assistant who doesn’t rush and always double-checks before finishing the job. ### Database Management ``` def initialize_db(db_name="bestbuy_refrigerators.db"): """ Initialize SQLite database and create products table if it doesn't exist. Args: db_name (str): Name of the SQLite database file Returns: sqlite3.Connection: Database connection object The products table schema: - id: Primary key - url: Unique product URL - date_scraped: Date when URL was scraped - scraped: Flag indicating if product details have been scraped (0/1) """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS products ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, date_scraped DATE, scraped INTEGER DEFAULT 0 ) ''') conn.commit() return conn ``` When we gather data, it’s important not to just grab it and forget it—we need a safe place to store everything we find. That’s where our database manager comes in. In our case, we use a SQLite database, which you can imagine as a supercharged version of an Excel sheet. It’s built to handle not just hundreds, but even millions of rows of data without slowing down. We set up a table inside this database to store key details: the product’s URL, the date we found it, and whether we’ve already collected its full details or not. But one of the best things about this setup is how reliable it is. Let’s say your computer crashes in the middle of scraping. Normally, that would mean starting over—but not here. Our system keeps track of progress as it goes, almost like an autosave feature in a video game. So even if something goes wrong, you can just pick up right where you left off. This makes the whole process smoother, safer, and a lot less stressful. ### URL Storage ``` def insert_product_url(conn, url): """ Insert a product URL into the database if it doesn't exist. Args: conn (sqlite3.Connection): Database connection object url (str): Product URL to insert Note: Uses INSERT OR IGNORE to handle duplicate URLs gracefully """ cursor = conn.cursor() try: cursor.execute(''' INSERT OR IGNORE INTO products (url, date_scraped, scraped) VALUES (?, ?, ?) ''', (url, datetime.now().date(), 0)) conn.commit() except sqlite3.Error as e: logging.error(f"Error inserting URL {url}: {e}") ``` This function works a lot like a meticulous librarian—one who’s determined to keep the catalog tidy and free of duplicates. Each time a new URL comes in, the function first checks if it’s already in our database. After all, there's no need to store the same link more than once—that would only create unnecessary clutter. To do this smartly, it uses a feature in SQLite called "INSERT OR IGNORE." That might sound technical, but the idea is simple: “Only add this new entry if it’s not already there.” It’s like having a filing system that quietly prevents you from putting the same paper in the drawer twice. Clean, efficient, and organized. Now, what if something goes wrong? Maybe the database is busy at the moment, or there’s an oddly formatted URL that causes a hiccup. Instead of stopping the entire process, the function stays calm. It writes down the issue using our logging system and moves on. This is incredibly important when you’re dealing with thousands of URLs—you don’t want one bad link to bring your whole project to a halt. It’s built to be steady, reliable, and quietly persistent, just like that skilled librarian who keeps the shelves in perfect order, no matter what. ### HTML Parsing ``` def parse_product_urls(content): """ Extract product URLs from the page HTML content using BeautifulSoup. Args: content (str): HTML content of the page Returns: list: List of complete product URLs Note: URLs are constructed by appending the extracted href to the base Best Buy URL """ soup = BeautifulSoup(content, "html.parser") product_urls = [] for link in soup.select("li.sku-item > div > div > div > div.shop-sku-list-item > div.list-item.lv > div.column-left > a.image-link"): href = link.get("href") if href: full_url = f"https://www.bestbuy.com{href}" product_urls.append(full_url) return product_urls ``` Here’s where we start turning the raw web content into something meaningful. The tool that helps us do this is BeautifulSoup, which allows us to move through the HTML structure of a webpage. It lets us see exactly where the valuable pieces—like product links—are hiding. This function is trained to look for very specific patterns in the HTML. It’s kind of like reading a treasure map where an ‘X’ marks the spot. When it finds what looks like a product link, it checks whether the link is complete. If it’s not, it cleverly builds the full URL by adding Best Buy’s domain to the front—ensuring we always get a valid, working link. We’ve also taken care to make this process more reliable. Websites often change how they look, but not everything changes at once. So instead of relying on elements that might move around, we focus on patterns in the HTML that tend to stay the same. That way, our scraper remains useful and accurate, even as the site evolves over time. ### Main Scraping Logic ``` async def scrape_product_urls(base_url, user_agents_file, db_name="bestbuy_refrigerators.db"): """ Main scraping function that coordinates the entire scraping process. Args: base_url (str): Template URL for Best Buy's refrigerator category pages user_agents_file (str): Path to file containing user agent strings db_name (str): Name of the SQLite database file This function: 1. Loads user agents for rotation 2. Initializes the database 3. Iterates through pages 4. Handles browser automation using Playwright 5. Manages country selection popups 6. Coordinates scrolling and content extraction 7. Stores results in the database Anti-detection measures: - Random user agent rotation - Random delays between pages - Standard viewport size - Proper handling of lazy loading """ user_agents = load_user_agents(user_agents_file) conn = initialize_db(db_name) async with async_playwright() as p: for page_num in range(1, 10): # Loop through pages 1 to 9 current_page = base_url.format(page_num=page_num) logging.info(f"Scraping page {current_page}...") user_agent = random.choice(user_agents) browser = await p.chromium.launch(headless=False) # Set headless to False to see the browser context = await browser.new_context( user_agent=user_agent, viewport={'width': 1920, 'height': 1080} # Set a standard viewport size ) page = await context.new_page() try: await page.goto(current_page, timeout=120000) # Handle country selection if it appears if "Choose a country" in await page.content(): logging.info("Detected country selection page, navigating to the US site...") await page.click("body > div.page-container > div > div > div > div:nth-child(1) > div.country-selection > a.us-link") await page.wait_for_load_state("domcontentloaded", timeout=120000) # Scroll to load all products await scroll_to_bottom(page) await page.wait_for_timeout(200000) content = await page.content() product_urls = parse_product_urls(content) logging.info(f"Total URLs scraped from page {page_num}: {len(product_urls)}") for url in product_urls: insert_product_url(conn, url) # Add a small delay between pages await page.wait_for_timeout(random.uniform(2000, 4000)) except Exception as e: logging.error(f"Error on page {page_num}: {e}") continue finally: await page.close() await context.close() await browser.close() conn.close() ``` This part of our project is the true brain of the entire scraping operation. Think of it as a well-organized project manager, quietly coordinating a team where each member has a specific role—browsing pages, handling pop-ups, collecting data, and carefully filing it away. When the function starts, it doesn’t just rush in. First, it gets everything ready. It loads up the different user agents we’ll use as disguises, sets up the database for storing our findings, and applies any configuration settings we need. This preparation is like a researcher setting up their tools before diving into a stack of library books—calm, focused, and methodical. Then, as it goes through each page of product listings, it behaves more like a real person than a robot. It pauses for a random amount of time between actions, mimicking natural human behavior. It even knows how to deal with pop-ups—like a prompt asking you to select your country—by interacting with them just the way a real visitor would. And throughout this process, it rotates through different user agents to help avoid detection. What really makes this function smart is how it handles problems. If one page doesn’t load or causes an error, it doesn’t stop everything. It simply makes a note of what went wrong in the log and moves on, just like a determined researcher who doesn’t let one missing book throw off the entire study. It’s thoughtful, adaptable, and persistent—exactly what a good scraping engine needs to be. ### Script Entry Point ``` if __name__ == "__main__": # Base URL template for Best Buy's refrigerator category # Parameters: # - page_num: Page number for pagination # - Additional query parameters filter for in-stock items and category base_url = "https://www.bestbuy.com/site/searchpage.jsp?_dyncharset=UTF-8&browsedCategory=pcmcat1637590307724&cp={page_num}&id=pcat17071&iht=n&ks=960&list=y&qp=soldout_facet%3Dname~Exclude%20Out%20of%20Stock%20Items&sc=Global&st=pcmcat1637590307724_categoryid%24pcmcat367400050001&type=page&usc=All%20Categories" try: asyncio.run(scrape_product_urls(base_url, "user_agents.txt")) except KeyboardInterrupt: logging.info("Scraping interrupted by user") except Exception as e: logging.error(f"Fatal error: {e}") ``` This is where our scraper program actually begins to run, like turning the key in a car's ignition. It sets up the initial URL we want to scrape (Best Buy's refrigerator category in this case) and kicks off the scraping process. The entry point includes all safety features for when things go wrong. It is capable of handling two types of stops: when you deliberately want to stop the script (like you pressed Ctrl+C), and when unexpected errors occur. Imagine that you have both regular and emergency brakes in your car. ## DATA COLLECTION ### Import Section ``` import sqlite3 from playwright.async_api import async_playwright from bs4 import BeautifulSoup import asyncio import json ``` The import section for this part of the project looks very similar to what we used earlier when scraping product URLs. It brings in all the essential tools—like Playwright for browser automation, BeautifulSoup for navigating through HTML, asyncio for asynchronous programming and SQLite for handling our database. But there’s one new addition here: the json library. This small but powerful tool is especially handy when we're dealing with structured data, like product specifications. Often, product details are stored on a webpage in neat, organized blocks—almost like little data files hidden inside the HTML. These blocks are usually written in JSON format. ### Database Connection Functions #### connect\_db ``` # Database Connection def connect_db(db_path='bestbuy_refrigerators.db'): """ Create and return a connection to the SQLite database. Args: db_path (str): Path to the SQLite database file Returns: sqlite3.Connection: Database connection object """ return sqlite3.connect(db_path) ``` The connect\_db function is where our data storage journey truly begins. Think of it as opening the front door to our data vault. When this function is called, it connects to a file-based SQLite database. If the file isn’t there yet, it quietly creates one for us—no extra steps needed. This connection is more than just a one-time handshake. It becomes a steady, reliable pipeline through which all our data flows—whether we’re saving new product information or checking what we’ve already collected. Without this connection, nothing else involving the database can happen. We’ve also made the function flexible. By default, it connects to a database file named 'bestbuy\_refrigerators.db', but if needed, we can tell it to use a different file. This makes the function reusable across different projects. Plus, it’s built to work well in asynchronous environments, meaning even if multiple parts of the program are trying to use the database at once, everything stays safe and efficient. So, you can think of this connection as setting up a dedicated phone line between our scraper and the database. It’s secure, always on, and forms the foundation of all the data-related tasks that follow. Without it, we’d have no way to store or retrieve anything—we’d be scraping into thin air. #### create\_tables ``` # Ensure necessary tables exist def create_tables(conn): """ Initialize database schema by creating necessary tables if they don't exist. Args: conn (sqlite3.Connection): Database connection object Tables created: 1. scraped_data: Stores successful scraping results with product details 2. error_urls: Tracks failed scraping attempts with error messages """ cursor = conn.cursor() # Create `scraped_data` table if it doesn't exist with separate title and price columns cursor.execute(""" CREATE TABLE IF NOT EXISTS scraped_data ( url TEXT, title TEXT, sale_price TEXT, price TEXT, discount TEXT, brand TEXT, model TEXT, sku TEXT, rating TEXT, reviews TEXT, key_specs TEXT, date TEXT ) """) # Create `error_urls` table if it doesn't exist cursor.execute(""" CREATE TABLE IF NOT EXISTS error_urls ( url TEXT, error_message TEXT, date_scraped TEXT, scraped INTEGER DEFAULT 0 ) """) conn.commit() ``` The create\_tables function acts like the architect of our data storage system. Before we can store anything, we need a solid structure—a blueprint that tells the database exactly what kinds of information we’re going to save, and where each piece should go. That’s exactly what this function sets up. It creates two key tables. The first one, called scraped\_data, is where we’ll keep everything we collect about refrigerators. This table is designed with care, including columns for all the details we might extract—such as the product URL, title, sale price, regular price, discount, brand, model number, SKU, customer rating, review count, and even a slot for detailed specifications. Each column is assigned the appropriate data type, so whether we’re storing numbers, text, or longer descriptions, everything fits neatly in place. The second table, error\_urls, serves a different but equally important purpose. Sometimes, scraping a product page might fail—maybe the page didn’t load, or the data wasn’t in the expected format. Instead of letting these failures vanish, we log them here. This table includes the URL that caused trouble, the error message, the date it happened, and the current status—so we can come back and handle those issues later. To keep things smooth, the function uses the “CREATE TABLE IF NOT EXISTS” command. This means it won’t break or complain if the tables already exist. You can safely run it every time the script starts, and it will only create the tables if they aren’t there yet. Finally, it commits the changes to the database right away, making sure the setup is saved and ready to go—like locking in the foundation of a building before construction begins. ### Database Operation Functions #### fetch\_unscraped\_urls ``` # Fetch unsaved URLs and their dates def fetch_unscraped_urls(conn): """ Retrieve URLs that haven't been successfully scraped yet. Args: conn (sqlite3.Connection): Database connection object Returns: list: Tuples of (url, date_scraped) for unscraped products """ cursor = conn.cursor() cursor.execute("SELECT url, date_scraped FROM products WHERE scraped = 0") return cursor.fetchall() ``` The fetch\_unscraped\_urls function works like a smart task manager for our scraping system. Its job is to figure out which URLs still need attention—basically, the to-do list of pages we haven’t successfully scraped yet. To do this, it looks into the error\_urls table, which is where we keep track of all the pages that caused trouble during previous scraping attempts. It specifically checks for rows where the scraped flag is set to 0, meaning those URLs haven't been successfully processed yet. These are the tasks waiting in line, and this function helps our scraper pick them up and try again. When the function runs, it returns two pieces of information for each URL: the link itself, and the date it was first added to the table. This not only gives us a clear view of which URLs are pending, but also provides historical context—like when we first tried to scrape them. That information helps us make better decisions, like whether to prioritize older tasks or simply track how long something has been in the queue. Behind the scenes, the SQL query used in this function is written for efficiency, which means it runs quickly even when the table has thousands of entries. That’s important, especially for large-scale scraping projects that run for hours or days. In short, fetch\_unscraped\_urls makes sure we don’t miss anything and that we always know what still needs to be done—keeping the whole operation flowing smoothly. #### mark\_as\_scraped ``` # Update 'scraped' status def mark_as_scraped(conn, url): """ Mark a URL as successfully scraped in the database. Args: conn (sqlite3.Connection): Database connection object url (str): URL of the scraped product """ cursor = conn.cursor() cursor.execute("UPDATE products SET scraped = 1 WHERE url = ?", (url,)) conn.commit() ``` The mark\_as\_scraped function plays the role of a progress tracker in our scraping system. Once we've successfully collected data from a URL—without any errors—this function steps in and updates the record in the database to reflect that the job is done. Specifically, it changes the scraped flag to 1 in the error\_urls table for that particular URL. This update is important because it tells our system, “We’ve already taken care of this one—no need to do it again.” Without this step, we might end up scraping the same page over and over, wasting time and resources. What makes this function reliable is that it’s built with error handling in mind. If something unexpected happens while trying to update the database, it doesn’t let that error cause bigger problems. And once the update is successfully made, it commits the change to the database, ensuring the progress is saved right away. Even better, this function performs its update in what's called an atomic way. That means the operation either completes fully or doesn't happen at all—there’s no in-between. This helps protect the accuracy of our tracking system. If updates were only partially saved, we could end up with mixed or corrupted information about what’s been scraped and what hasn’t. So, mark\_as\_scraped is a small function with a big responsibility: it keeps our progress clear, our data collection efficient, and our records accurate. #### save\_scraped\_data ``` # Save scraped data to the `scraped_data` table with separate title and price columns def save_scraped_data(conn, url, title, sale_price,price,discount,brand,model,sku,rating,reviews, key_specs,original_date): """ Store successfully scraped product data in the database. Args: conn (sqlite3.Connection): Database connection object url (str): Product URL title (str): Product title sale_price (str): Current sale price price (str): Regular price discount (str): Discount amount/percentage brand (str): Product brand model (str): Model number sku (str): SKU number rating (str): Product rating reviews (str): Number of reviews key_specs (str): JSON string of product specifications original_date (str): Date when URL was first discovered """ cursor = conn.cursor() cursor.execute("INSERT INTO scraped_data (url, title, sale_price, price,discount, brand,model,sku,rating,reviews,key_specs, date) VALUES (?, ?, ?, ?,?,?,?,?,?,?,?,?)", (url, title, sale_price,price,discount,brand,model,sku,rating,reviews, key_specs,original_date)) conn.commit() ``` The save\_scraped\_data function is like the official archivist of our project. Its job is to take all the details we’ve collected about a refrigerator—like its name, price, brand, ratings, and more—and carefully store them in our database for future use. Each time this function is called, it receives a set of inputs—one for each piece of product information. It then inserts all of that into the corresponding columns of the scraped\_data table. Think of it like neatly filing each product’s info into the right drawer in a well-organized filing cabinet. But this function doesn’t just drop the data in and hope for the best—it’s built to be smart and cautious. It checks that everything is in the correct format, makes sure special characters won’t break the database, and even handles complex fields like technical specifications. This attention to detail helps keep our records clean, consistent, and easy to retrieve later. Another important aspect is the commit step. Once the data is inserted, the function commits the change to the database, locking it in. This means that even if something unexpected happens later—like the script crashes or the internet drops—whatever was already saved stays safe. Lastly, this function always ties each entry back to the original product URL. That way, we always know exactly where the data came from. This traceability is key when double-checking information or fixing issues later on. In short, save\_scraped\_data is what makes sure all our hard-earned data is safely stored, correctly organized, and easy to access. #### save\_error\_url ``` # Save error details to `error_urls` table def save_error_url(conn, url, error_message, original_date): """ Log failed scraping attempts for retry. Args: conn (sqlite3.Connection): Database connection object url (str): Failed product URL error_message (str): Description of the error original_date (str): Date when URL was first discovered """ cursor = conn.cursor() cursor.execute("INSERT OR REPLACE INTO error_urls (url, error_message, date_scraped,scraped) VALUES (?, ?, ?,?)", (url, error_message, original_date,0)) conn.commit() ``` The save\_error\_url function works like the troubleshooter of our scraping system. Whenever something goes wrong while trying to scrape a page—maybe the site didn’t load properly, or the data wasn’t found—this function steps in to log exactly what happened. It records three key pieces of information: The URL that caused the issue, a detailed error message explaining what went wrong, and a timestamp to show when the error occurred. This detailed logging is incredibly helpful later on. It lets us see patterns—for example, if the same URL fails multiple times or if similar errors keep popping up. With that kind of insight, we can better understand where the system needs fixing or adjusting. Instead of blindly adding new entries every time, this function uses an approach called "INSERT OR REPLACE". That means if the same URL has failed before, it simply updates the existing error record with the latest message and time. This keeps our database clean and avoids duplicate entries. It also resets the scraped flag back to 0, which is like saying, “Hey, this one didn’t work—try again later.” That way, the scraper knows to come back to it in future runs. Perhaps the most valuable part? The error messages themselves. These tell us whether the issue was caused by the website changing, a network hiccup, or something wrong in our own scraping code. Over time, this error history becomes a goldmine for improving the system—making it stronger, more accurate, and more reliable with each update. ### Web Scraping Helper Functions #### handle\_country\_selection ``` # Handle country selection if prompt appears async def handle_country_selection(page): """ Handle Best Buy's country selection popup if it appears. Args: page: Playwright page object """ if "Choose a country" in await page.content(): await page.click("body > div.page-container > div > div > div > div:nth-child(1) > div.country-selection > a.us-link") await page.wait_for_load_state("domcontentloaded", timeout=60000) ``` The handle\_country\_selection function acts like our geographic guide, helping the scraper navigate to the correct version of the Best Buy website—specifically, the U.S. version. This step is important because many international websites display a prompt asking visitors to choose their country, and we need to make sure we're seeing the same content that a U.S.-based customer would see. Here's how it works: when the scraper lands on a page, it checks if there’s a country selection pop-up. If it finds one, the function automatically selects “United States.” This ensures we get access to the right products, prices, and layout—just as we expect. But the function doesn’t rush. It patiently waits for the page to fully reload after making the selection. It uses Playwright’s built-in waiting features to confirm that the site is ready for the next steps. This helps avoid problems like trying to scrape a page before it's finished loading, which could lead to missing or broken data. What makes this function especially reliable is its error handling. If the country prompt is missing or appears in an unexpected format, it doesn’t crash or throw off the whole scraping process. Instead, it handles the situation smoothly, helping the scraper move forward without interruption. #### scroll\_to\_bottom ``` # scrolling function async def scroll_to_bottom(page): """ Implement infinite scroll to ensure all content is loaded. Args: page: Playwright page object This function scrolls until no new content is loaded, as determined by comparing page heights before and after scrolling. """ previous_height = await page.evaluate("document.body.scrollHeight") while True: await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await page.wait_for_timeout(2000) # Wait for lazy-loaded content new_height = await page.evaluate("document.body.scrollHeight") if new_height == previous_height: break previous_height = new_height ``` The scroll\_to\_bottom function plays the role of our content explorer. Its job is to make sure we don’t miss any data that's hidden behind endless scrolling—just like when you're on Instagram or Facebook and more posts appear as you scroll down. Many modern websites, including Best Buy, use this technique called lazy loading or infinite scrolling, where new items load only when you scroll further. So, if we only scrape what’s initially visible, we’d miss out on a large part of the product list. That’s where this function comes in. It scrolls the page in small steps, just like a real user would—going down, pausing, waiting for more content to appear, and checking if the page has grown in size. It does this by comparing the height of the page before and after each scroll. If the height stays the same, it knows we’ve likely reached the end. To make the process smooth and responsible, the function also waits between scrolls, giving the page time to load and avoiding too many requests at once. This reduces the risk of being blocked by the website or loading content too fast to capture. And in case something doesn’t go as expected—like a scroll gets interrupted or the page doesn't behave normally—the function has built-in error handling to keep things running safely without crashing the whole process. ### Data Extracting Functions #### extract\_title ``` def extract_title(soup): """Extract product title from the page.""" tittle_tag = soup.select_one('div.sku-title > h1') return tittle_tag.get_text(strip=True) if tittle_tag else "NOT AVAILABLE" ``` This function works like a focused scanner, carefully searching for the main title of the product on a webpage—specifically, the refrigerator’s name. It uses BeautifulSoup’s select\_one method, which allows us to target a very specific part of the page using a CSS selector. In this case, it looks for the text inside the HTML tag 'div.sku-title > h1', which is usually where Best Buy places the product title. You can think of this like using a magnifying glass to examine a very particular section of a document. If the product title is there, the function grabs the text, cleans up any extra spaces, and returns it neatly. But websites aren’t always predictable. Sometimes the layout changes, or maybe the product title is temporarily missing. To prepare for this, the function is designed with a backup plan. If it can’t find the title where it's expected, it simply returns "NOT AVAILABLE" instead of causing the whole scraper to crash. This makes the function both precise and resilient. So in simple terms, this function is like a smart assistant that goes straight to where the product name is supposed to be, checks if it’s there, and either brings it back cleaned up—or politely lets us know it couldn’t find it. Similarly, other parsing functions are used to extract details like sale price, brand, original price, discount, IDs, rating, and reviews. They follow the same pattern as the title parser—targeting specific elements and returning the value or a default if not found. See the code next for how each one works. #### extract\_sale\_price ``` def extract_sale_price(soup): """Extract current sale price if available.""" saleprice_tag = soup.select_one('div.flex.gvpc-price-1-2441-31 > div:nth-child(1) > div:nth-child(1) > div > span:nth-child(1)') return saleprice_tag.get_text(strip=True) if saleprice_tag else "NOT AVAILABLE" ``` #### extract\_brand ``` def extract_brand(soup): """Extract product brand name.""" brand_tag = soup.select_one('div.pb-200 > a') return brand_tag.get_text(strip=True) if brand_tag else "NOT AVAILABLE" ``` #### extract\_price ``` def extract_price(soup): """Extract regular product price.""" price_tag = soup.select_one('div.flex.gvpc-price-1-2441-31 > div:nth-child(1) > div.pricing-price__savings-regular-price > div.pricing-price__regular-price-content--block.pricing-price__regular-price-content--block-mt > div:nth-child(1) > span') return price_tag.get_text(strip=True) if price_tag else "NOT AVAILABLE" ``` #### extract\_discount ``` def extract_discount(soup): """Extract discount information if available.""" discount_tag = soup.select_one('div.flex.gvpc-price-1-2441-31 > div:nth-child(1) > div.pricing-price__savings-regular-price > div.pricing-price__savings.pricing-price__savings--promo-red') return discount_tag.get_text(strip=True) if discount_tag else "NOT AVAILABLE" ``` #### extract\_ids ``` def extract_ids(soup): """Extract model and SKU numbers.""" model_tag = soup.select_one('div.title-data.lv > div > div.model.product-data.pr-100.inline-block.border-box > span.product-data-value.text-info.ml-50.body-copy') sku_tag = soup.select_one('div.title-data.lv > div > div.sku.product-data.pr-100.inline-block.border-box > span.product-data-value.text-info.ml-50.body-copy') return (model_tag.get_text(strip=True) if model_tag else "NOT AVAILABLE", sku_tag.get_text(strip=True) if sku_tag else "NOT AVAILABLE") ``` #### extract\_rating\_reviews ``` def extract_rating_reviews(soup): """Extract product rating and number of reviews.""" rating_tag = soup.select_one('ul > li > a > div > span.ugc-c-review-average.font-weight-medium.order-1') reviews_tag = soup.select_one('ul > li > a > div > span.c-reviews.order-2') return (rating_tag.get_text(strip=True) if rating_tag else "NOT AVAILABLE", reviews_tag.get_text(strip=True) if reviews_tag else "NOT AVAILABLE") ``` #### extract\_key\_specs ``` def extract_key_specs(soup2): """Extract and format product specifications as JSON.""" specs = {} for row in soup2.find_all('div', class_='zebra-row flex p-200 justify-content-between body-copy-lg'): complete_text = row.get_text(separator='\n').strip() lines = [line.strip() for line in complete_text.split('\n') if line.strip()] if len(lines) >= 2: specs[lines[0]] = lines[1] return json.dumps(specs, indent=4) ``` This function is like a super-organized assistant whose job is to carefully go through every line of a refrigerator’s specification sheet and turn it into something we can actually work with—clean, structured data. Think of it as someone reading a product brochure. For each row in the specifications table, the assistant looks at the feature name—like “Capacity” or “Cooling System”—and then notes down the corresponding value, such as “25.5 cubic feet” or “Twin Cooling Plus.” The function collects all these key-value pairs and saves them in a format called JSON, which is great for storing and comparing data in a structured way. This step is especially important because the specs section usually contains the most technical details about the product, and it's where you really start to understand the differences between models. Having this information organized lets us do things like comparing sizes, energy efficiency, or smart features across many refrigerators at once. It’s also designed to be thorough and cautious. The function looks only for rows with certain class names to avoid grabbing unrelated content. This helps make sure we collect the right kind of information while skipping the clutter. ### Main Scraping Functions #### scrape\_data ``` # Scrape data for a single URL async def scrape_data(url): """ Scrape all product details from a single URL. Args: url (str): URL of the product page to scrape Returns: tuple: All scraped product details Raises: RuntimeError: If any error occurs during scraping This function: 1. Launches a browser instance 2. Navigates to the product page 3. Handles country selection if needed 4. Scrolls to load all content 5. Extracts all product details 6. Clicks to view full specifications 7. Extracts detailed specifications """ try: async with async_playwright() as p: browser = await p.chromium.launch(headless=False) page = await browser.new_page() await page.goto(url) # Handle country selection if it appears await handle_country_selection(page) # Scroll to load all products await scroll_to_bottom(page) await page.wait_for_timeout(9000) print(f"scraping :{url}") soup = BeautifulSoup(await page.content(), 'html.parser') # Use the separate parsing functions title = extract_title(soup) sale_price = extract_sale_price(soup) price=extract_price(soup) discount=extract_discount(soup) brand=extract_brand(soup) model=extract_ids(soup)[0] sku=extract_ids(soup)[1] rating=extract_rating_reviews(soup)[0] reviews=extract_rating_reviews(soup)[1] # Click the button (add the actual selector for the button here) await page.click('div.col-xs-7 > div > button.c-button.c-button-outline.c-button-md.show-full-specs-btn.col-xs-6') # Replace 'button_selector' with the actual selector # Wait for content to load after clicking the button await page.wait_for_timeout(3000) # Adjust timeout as needed for the content to load # Retrieve and parse the new page content soup2 = BeautifulSoup(await page.content(), 'html.parser') # additional parsings key_specs=extract_key_specs(soup2) await browser.close() except Exception as e: raise RuntimeError(f"Error scraping {url}: {e}") return title,sale_price, price,discount,brand,model,sku,rating,reviews,key_specs ``` The scrape\_data function is like the main director of our entire scraping process. It’s the one calling the shots—making sure every other part of the script works together smoothly to collect data from a single product page. First, it opens up a browser window using Playwright. But not just any browser—it starts with specific settings that help it act more like a real human browsing the web. This helps reduce the chances of the website blocking us for being a bot. Once the browser lands on the product page, the real coordination begins. The function checks if there's a country selection popup and makes sure the U.S. version of the site is selected. Then, it scrolls down the page, step by step, allowing all the product information to load—just like a human would scroll slowly to read the page. After that, it pulls together all the important data: the product title, prices, ratings, specifications, and more. It does all of this while handling each step with care. If something doesn’t go as planned—maybe the layout changed or an element didn’t load—it catches the error and moves on without crashing the entire process. This kind of error handling is what makes the function reliable even when the website isn’t acting perfectly. And once everything is done, the function doesn’t just walk away. It makes sure to clean up—closing browser tabs and freeing up resources. Think of it as someone turning off the lights and shutting the door after finishing work. Overall, scrape\_data brings together all the smaller parts of our scraper—handling the page, collecting details, storing the data, and managing problems—all in one smooth and efficient flow. It’s the backbone of the entire scraping system. #### main ``` # Main scraper function async def main(): """ Main execution function that coordinates the scraping process. This function: 1. Establishes database connection 2. Ensures necessary tables exist 3. Retrieves unscraped URLs 4. Attempts to scrape each URL 5. Saves successful results and logs failures 6. Handles cleanup """ conn = connect_db() # Ensure tables exist create_tables(conn) urls_dates = fetch_unscraped_urls(conn) for url, original_date in urls_dates: try: title, sale_price, price,discount,brand,model,sku,rating,reviews ,key_specs= await scrape_data(url) # Await scrape_data since it's async save_scraped_data(conn, url, title, sale_price, price,discount, brand,model,sku,rating,reviews,key_specs, original_date) # Pass title and price separately mark_as_scraped(conn, url) # Mark as scraped after saving except Exception as e: error_message = str(e) save_error_url(conn, url, error_message, original_date) # Save error details conn.close() if __name__ == "__main__": asyncio.run(main()) ``` The main function is the supreme orchestrator of our entire scraping operation. It begins by establishing the database connection and ensuring that our data structures—such as tables—are correctly set up. This function creates the workflow backbone, tying together all the components of the scraping system. It manages the overall control flow: fetching the list of URLs that need processing, launching scraping attempts one by one, handling successful scrapes, capturing any errors, and ensuring that the results—whether successful data or failure logs—are stored appropriately. Importantly, it includes top-level error handling, which means that even if individual URLs fail, the entire scraping process continues without interruption. In addition to coordination, the main function handles resource management—cleaning up database connections and freeing up system resources after scraping is done. It is designed to support long-running sessions with stability and integrity, making sure no data is lost and system performance stays reliable. Finally, it serves as the true entry point for the script, initiating the asyncio event loop that powers our asynchronous scraping logic. This ensures our scraper runs efficiently—handling multiple tasks at once—while maintaining tight control over flow and exceptions. ## Conclusion Scraping refrigerator data from Best Buy showcases how automation can streamline large-scale data collection, making it significantly easier to analyze product listings, prices, and availability. By combining Playwright for handling dynamic web elements with BeautifulSoup for precise data extraction, we were able to gather well-structured information efficiently—without manual effort. This method proves especially valuable for price tracking, trend analysis, and market research, empowering both businesses and consumers to make smarter decisions. That said, web scraping must be used responsibly. It’s essential to respect a website’s policies, avoid overwhelming servers with too many requests, and implement proper delays between actions. Understanding the structure of Best Buy’s site, managing dynamic content carefully, and ensuring ethical practices are key to sustainable data extraction. As e-commerce platforms grow more complex and data-driven decisions become the norm, web scraping stands out as a critical skill for analysts, developers, and researchers who want to stay competitive in a fast-paced digital market. ### AUTHOR I’m Shahana, a Data Engineer at Datahut, where I specialize in building smart, scalable data pipelines that convert raw, dynamic web content into clean, actionable insights—fueling better decisions across retail and consumer electronics. At Datahut, we’ve spent more than a decade helping businesses harness the power of automation for product tracking, price intelligence, and competitive analysis. In this blog, I walk you through how we used Playwright and BeautifulSoup to scrape refrigerator product data from Best Buy—capturing detailed specifications, prices, and availability from a JavaScript-heavy site with accuracy and efficiency. If your team is exploring ways to automate product data collection in the electronics sector—or any high-volume e-commerce domain—feel free to reach out through the chat widget on the right. We’d be happy to help you design a robust scraping solution that meets your needs. FAQ SECTION 1.What kind of refrigerator data can be extracted from Best Buy using web scraping? You can extract a wide range of data points such as product names, model numbers, prices (regular and discounted), availability, customer ratings, number of reviews, brand, capacity, energy efficiency ratings, dimensions, features (like smart connectivity, water dispensers), warranty details, and promotional offers. 2\. Is it legal to scrape refrigerator data from [BestBuy.com](https://www.bestbuy.com/?ref=blog.datahut.co)? Scraping public data is generally legal, but Best Buy’s website terms may restrict automated access. To avoid legal issues, it’s best to use scraping techniques responsibly (e.g., rate limiting) and review the site's robots.txt and terms of service. 3\. Why would someone scrape refrigerator data from Best Buy? Common reasons include: - Price comparison with other retailers - Market research on brands and features - Inventory monitoring for stock availability - Competitor analysis - Affiliate marketing or price tracking tools 4\. Can I track price drops of refrigerators using web scraping? Yes, by scheduling regular scrapes (e.g., daily or weekly), you can detect changes in pricing, identify limited-time deals, or track discount trends over time. 5\. What tools are used to scrape Best Buy for refrigerator data? Popular tools include Python libraries such as BeautifulSoup, Scrapy, Selenium, and Playwright. For large-scale projects, you may also use cloud scraping services or scraping APIs. 6\. How often should I scrape Best Buy for accurate refrigerator data? It depends on your use case. For price tracking, daily or hourly scrapes may be ideal. For product listing or feature analysis, a weekly or bi-weekly scrape might suffice. 7\. Will Best Buy block my scraper? Yes, if your scraper sends too many requests too quickly or violates site rules. Use techniques like rotating IPs, proxies, and user-agent headers, and add delay between requests to reduce the risk of being blocked. 8\. Can scraped refrigerator data be used in a product comparison app? Yes, as long as you comply with copyright and terms of use. Many apps use scraped data to display features, reviews, and prices for comparison purposes. 9\. What’s the difference between using Best Buy’s API and scraping the website? Best Buy offers a public API with structured data access, but it may have usage limits or restricted access. Scraping allows you to extract data not available in the API, like promotional banners or dynamically loaded specs. ### Bullwhip Effect: Why You're Losing Money on Your Ecom Store and How to Stop It URL: https://www.blog.datahut.co/post/the-bullwhip-effect-why-you-re-losing-money-on-your-ecom-store-and-how-to-stop-it/ Last updated: 2026-09-07T09:44:21.000Z Ever wonder why your ecommerce store runs out of top-selling products right when they’re in demand—or why you’re stuck with piles of dead inventory a few weeks later? That’s not just bad luck. It’s[ ](https://www.unleashedsoftware.com/blog/bullwhip-effect-in-supply-chain?ref=blog.datahut.co)[the bullwhip effect](https://www.unleashedsoftware.com/blog/bullwhip-effect-in-supply-chain?ref=blog.datahut.co)—one of the most frustrating (and expensive) challenges for online brands like yours. A tiny shift in sales this week—maybe a 5% dip—can trigger massive supply chain overreactions. Before you know it, your suppliers are cutting production, your warehouse shelves are either empty or overflowing, and your revenue is stuck in limbo. You’re not alone. According to the IHL Group,[ ](https://www.retaildive.com/news/out-of-stocks-could-be-costing-retailers-1t/526327/?ref=blog.datahut.co)[stockouts and overstocks cost retailers nearly $1 trillion](https://www.retaildive.com/news/out-of-stocks-could-be-costing-retailers-1t/526327/?ref=blog.datahut.co) a year. But here’s the thing: you can stop it. In this post, we’ll show you how this happens inside your business, how to spot the early warning signs, and what you can do—starting now—to fix it with smarter demand tracking and web scraping. ## What Is the Bullwhip Effect (and Why It Hurts Your Store)? The bullwhip effect happens when a small change in demand at your store causes big swings in decisions up the supply chain. Example: One of your top products dips in sales over the weekend. Maybe it’s weather, maybe it’s just timing. You cut your next order slightly to be cautious. Your supplier sees the drop and pulls back further. And suddenly, when demand rebounds next week, you’re stocked out—and customers move on. Meanwhile, if a product suddenly spikes, you might overreact and over-order—only to be left discounting that same inventory three weeks later. Sound familiar? ## [The Real-World Version](https://www.inboundlogistics.com/articles/bullwhip-effect/?ref=blog.datahut.co) It’s like sneezing in a quiet room… and someone three doors down thinks there’s a fire. In ecommerce, even a short-term sales blip can look like a trend if you're not monitoring demand continuously—and in context. ## How You Might Be Misreading the Signals Here are a few ways ecommerce teams accidentally trigger the bullwhip effect: - A product sells poorly one week due to weather or competitor ads - You assume demand dropped permanently - You reduce your inventory orders - Sales rebound the next week—but your product’s out of stock Now you're spending on ads with nothing to sell. Without real-time insights, these short-term blips look like long-term trends. And your inventory planning starts to lag behind reality. ## Promotions Are a Trap (If You’re Not Prepared) Running a sale or influencer campaign? Promotions should drive sales—not chaos. But if your stock planning isn’t based on current data, here’s what can happen: - Your top products run out mid-promo - You panic-order more stock - Sales drop after the promo ends - You're stuck holding excess inventory By using[ ](https://www.blog.datahut.co/post/web-scraping-asos-data-insights-into-pricing-promotions-and-product-diversity/)[web scraping services](https://www.blog.datahut.co/post/web-scraping-asos-data-insights-into-pricing-promotions-and-product-diversity/) to[ ](https://www.blog.datahut.co/post/web-scraping-asos-data-insights-into-pricing-promotions-and-product-diversity/)[monitor competitor promotions and market trends](https://www.blog.datahut.co/post/web-scraping-asos-data-insights-into-pricing-promotions-and-product-diversity/), you can forecast more accurately and avoid the feast-or-famine trap. ## A Quick Example: How One Brand Turned It Around A mid-sized DTC apparel brand specializing in seasonal wear struggled with constant stockouts during influencer campaigns. They assumed it was a supplier issue, but the real cause was internal: their forecasting relied only on last month’s averages. After partnering with Datahut to track competitor activity and social trends, they began[ ](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/)[spotting early demand shifts](https://www.blog.datahut.co/post/understanding-zara-s-pricing-strategy-analyzing-product-price-distribution/)—like trending styles on competitors. With better forecasting, they: - Improved in-stock rate during promos by 25% - Reduced express shipping costs by 40% Web scraping enabled them to anticipate demand instead of reacting to it. ## Visualizing the Bullwhip: A Data Cascade Here’s how a 5% demand dip can spiral: - Retailer cuts order by 10% - Distributor cuts forecast by 20% - Manufacturer reduces output by 40% All from one quiet weekend. Without real-time data, each player overreacts—and your store pays the price. ## The Hidden Costs You’re Already Paying The bullwhip effect[ ](https://www.retailtouchpoints.com/features/industry-insights/ihl-study-inventory-distortion-will-cost-retailers-1-77-trillion-in-2023?ref=blog.datahut.co)[drains your profits](https://www.retailtouchpoints.com/features/industry-insights/ihl-study-inventory-distortion-will-cost-retailers-1-77-trillion-in-2023?ref=blog.datahut.co) in several ways: ### [Lost revenue from stockouts](https://www.flieber.com/blog/stockout-costs?ref=blog.datahut.co) When your most in-demand products disappear from shelves, customers quickly turn to competitors. Every missed sale also weakens customer loyalty and harms your store’s credibility—especially during high-intent moments like promotions or seasonal peaks. ### Carrying costs from excess inventory: Over-ordering doesn’t just fill up shelf space—it drains capital. You’re paying for warehousing, insurance, handling, and depreciation on products that may not move for weeks or months. This ties up cash you could otherwise use for marketing, new launches, or operational improvements. ### Wasted ad spend on out-of-stock pages When ads drive traffic to unavailable products, every click becomes a sunk cost. Not only do you lose immediate revenue, but poor user experience also reduces ad efficiency over time, harming ROAS and scaling potential. ### Customer churn from unavailable products Shoppers expect consistency. When customers repeatedly encounter out-of-stock issues, they lose trust and switch brands. This creates long-term revenue leakage, because reacquiring churned customers is significantly more expensive than retaining them. ## What Happens Without Web Scraping Without scraping, teams rely on slow or outdated data: - Old spreadsheets - Delayed dashboards - Lagging historical sales - Guess-based promo planning Automated data extraction closes this gap. ## Why Your Team’s Always in Firefighting Mode When your teams operate from misaligned or delayed data, your business becomes reactive. Marketing, ops, and suppliers all end up responding late—and often at odds. ## The Fix: Continuous Monitoring Using Web Scraping The solution to the bullwhip effect is real-time demand sensing. Web scraping provides - Live SKU visibility - Competitor stock and price trends - Daily shifts in category performance - Promo triggers ## Web Scraping for Different Teams Different teams benefit in different ways: ### Marketing Web scraping gives marketing teams real-time intelligence on competitor promotions, trending products, and shifting customer interest. This helps them schedule campaigns at the most profitable moment, avoid promoting out-of-stock items, and align messaging with market demand instead of guesswork. ### Ops Operations teams gain instant visibility into stock movements across competitors and marketplaces. With automated alerts when competitor SKUs go out of stock or categories heat up, ops can adjust procurement before issues escalate—reducing stockouts, unnecessary replenishment, and warehousing waste. ### Finance Finance teams benefit from clearer visibility into where cash gets locked up. Scraped pricing, availability, and demand trends highlight which SKUs may become slow-moving or risk overstocking, helping finance better forecast cash flow, plan budgets, and support smarter working-capital decisions. ### Leadership: Executives get a unified, data-driven view of market shifts, competitive actions, and operational blind spots. Instead of reacting to lagging reports, leadership can make proactive, high-ROI decisions—whether that’s scaling a product line, adjusting assortments, or reallocating budgets across teams. ## Why Web Scraping Works for Ecommerce Brands Using Datahut’s scraping platform, brands can: - Track competitor catalogs - Adjust pricing dynamically - Detect demand spikes - Feed structured data into tools ## A Quick Gut Check—Are You Bullwhipped? - [Stockouts mid-promo?](https://www.zenventory.com/blog/preventing-costly-stockouts-strategies-for-e-commerce-success?ref=blog.datahut.co) - Over-ordering? - Surprise sales swings? - Heavy discounting? If yes—you’re dealing with it. ## How the Bullwhip Effect Affects Your Cash Flow Over-ordering ties up capital in unsold products. Under-ordering leads to missed revenue and unreliable forecasts. Both eat into your margins. ## Are You Ready to Fix Your Demand Signals? (Quick Checklist) Ask yourself: - Do you rely on last month’s sales? - Are teams using different data sources? - Are you tracking competitors in real time? - Is your data weekly or monthly? - Do stock decisions often arrive too late? If yes—you need better demand tracking. ## What to Do Now The bullwhip effect is a data problem. With continuous data, you can: - See real-time demand - Align team decisions - Control inventory, pricing, and cash flow Datahut helps brands capture clean, live product data across the web. ## Don’t Wait for the Next Crisis The bullwhip effect is slow and silent until it becomes expensive. Early adopters gain an edge—acting before crises hit. ## Automate Your Demand Signals with Datahut Stop flying blind. Datahut gives you real-time competitor prices, stock levels, and product trends—delivered as clean, structured data. - Prevent stockouts - Avoid over-ordering - Improve forecasting accuracy by 25–40% 📩 [Book a free 30‑minute consulting](https://www.datahut.co/?ref=blog.datahut.co) to see what signals you’re currently missing. ## Final Thoughts Winning brands listen to real-time demand signals instead of reacting to lagging reports. Start small. Stay consistent. Become a more resilient, profitable brand. ## Frequently Asked Questions (FAQ) 1. What causes the bullwhip effect? Small changes in customer demand get amplified as they move up the supply chain. When retailers, distributors, and suppliers all interpret a small dip or spike differently, the result is overreaction—leading to stockouts, excess inventory, and operational chaos. 2. How can I prevent stockouts during promotions? Use real-time tracking of competitor promotions, category shifts, and fast-moving SKUs. Web scraping feeds your team live insights, letting you adjust inventory before demand hits instead of reacting after a campaign goes live. 3. How does web scraping improve inventory planning? Scraped data reveals pricing trends, availability patterns, and product movements across the market. This context helps your forecasting models become more accurate, reducing over-ordering and missed sales. 4. Can small brands use scraping? Absolutely. Even small and mid-sized brands gain an edge because scraping automates insights that would normally take hours of manual research. It levels the playing field against larger competitors. 5. What’s the first step? Identify where your demand signals are lagging—pricing, availability, competitor promotions, or category trends. Then plug those gaps with automated data feeds from Datahut. 6. How accurate is scraped data compared to internal data? Internal POS data shows what already happened. Scraped data shows what’s happening right now in the market. Combined, they give your team a complete picture for smarter decisions. 7. Is web scraping legal and compliant? Yes—when done responsibly. Datahut focuses only on publicly available information and follows ethical compliance guidelines to ensure your data workflows are fully safe. 8. Can scraped data help with long-term forecasting? Yes. While short-term signals help with immediate stock decisions, long-term scraping uncovers seasonal patterns, category trends, and demand cycles that strengthen your forecasting models. ## Ready to Fix Your Demand Signals? Your competitors are already tracking the market in real time. Don’t wait for the next stockout or inventory write-off. 🔥 Talk to a Datahut specialist and get a customized data plan for your business. ### Web Scraping Services Explained for Business Growth URL: https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/ Last updated: 2026-09-07T09:44:24.000Z Web scraping services offer tools and experience for scraping data from websites and converting unstructured data to useful information. [Web scraping](https://www.blog.datahut.co/post/15-web-scraping-questions/) services can be utilized for different functions, including market analysis, lead generation, and data analysis. ## Introduction: Why Web Scraping Matters A single e-commerce site can have 100,000+ product listings. Manually tracking prices across even five competitors? Nearly impossible. That’s why web scraping isn’t just helpful — it’s essential. Modern businesses operate in data-rich environments where decisions are only as good as the data behind them. Web scraping empowers you to collect competitive pricing, inventory availability, product attributes, and customer sentiment—turning public web content into structured, actionable insights. And with today’s demand for real-time data delivery, you can't afford to wait. Web scraping services fill that gap—offering speed, scale, and adaptability far beyond manual efforts or even traditional APIs. ![How web scraping services work](https://www.blog.datahut.co/content/images/2026/07/img-319.png.webp) ## What Are Web Scraping Services? Web scraping services are platforms or tools that automate the process of extracting data from websites. These solutions parse HTML structure to extract meaningful content—such as pricing, specifications, stock levels, or reviews—and convert it into machine-readable formats like CSV, JSON, or SQL. They simulate user behavior, handle dynamic loading, and even solve CAPTCHA challenges. Most importantly, they scale—allowing businesses to collect data from thousands of pages across multiple domains, frequently and reliably. In high-volume use cases, AI-powered platforms now handle much of this via AI-driven data extraction, adapting to layout changes and reducing manual oversight. Data scraping, often used interchangeably with web scraping, is the broader practice of collecting publicly available online content for business intelligence. ## Key Use Cases: Why Brands Use Web Scraping Here are the most popular ways smart brands apply scraping to gain a competitive edge: ### Competitive Intelligence - [Monitor competitor prices](https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/) and promotions in near real-time. - Track out-of-stock alerts to adjust your bidding or product mix. - Identify gaps in product listings or descriptions. ### Product & Market Insights - Extract reviews for sentiment analysis and R&D. - Track trends in fashion, tech, and consumer electronics. - Enrich datasets for training AI models. - Scraped data can reveal market trends, such as rising demand for specific product categories or shifting consumer sentiment. - By continuously monitoring competitors, brands can respond faster to market trends and optimize their strategy. ### Operational Efficiency - Use automated data pipelines to update product feeds. - Run assortment analysis for merchandising or pricing decisions. - Improve targeting and timing in performance marketing campaigns. Web scraping services help businesses tap into the vast ocean of web data available publicly—from product listings to customer sentiment. Whether you’re optimizing pricing or detecting market shifts, web data is the fuel for modern decision-making. The future of business intelligence lies in transforming web data into usable, real-time insight streams. ## Tools & Technologies in Web Scraping ![Best web scraping tools in 2025](https://www.blog.datahut.co/content/images/2026/07/img-320.png.webp) ### [Web Scraping Tools](https://www.blog.datahut.co/post/python-web-scraping-tutorial/) - BeautifulSoup (Python): Great for basic HTML parsing and static pages. - [Scrapy ](https://www.scrapy.org/?ref=blog.datahut.co)(Python): High-performance framework for scheduled crawls and pipelines. - Selenium: Simulates full browser interaction for JavaScript-heavy content. - Puppeteer / [Playwright](https://playwright.dev/?ref=blog.datahut.co): Best for headless scraping with infinite scroll and modals. - [Datahut](http://www.datahut.co/?ref=blog.datahut.co): Fully managed service with proxy support, IP rotation, and Web Scraping API access for high-scale, enterprise-grade extraction. While many scraping frameworks include an integrated HTML parser, choosing the right one impacts speed and accuracy—especially on pages with inconsistent or nested structures. At the core of every scraper lies an HTML parser, which interprets the structure of a webpage to extract content from elements like tables, lists, and product grids. Tools like BeautifulSoup act as lightweight HTML parsers, making it easier to target specific tags and attributes without loading entire browser sessions. Bonus: With API support and integration-ready delivery formats, services like Datahut make scraped data instantly usable in analytics stacks. ## Scraping Techniques That Actually Scale ### HTML Parsing & Scraping Layers At its core, scraping involves HTML parsing—analyzing the structure of a page and identifying key elements like , , or tags. But modern scraping involves more than just reading code. ### Advanced Techniques - Static vs. Dynamic Scraping: Dynamic pages require JavaScript rendering and simulation of scrolling or clicking. - Headless Browsing: Runs browsers without a UI to speed up tasks. - Proxy & IP Rotation: Essential for avoiding blocks and maintaining uptime. - CAPTCHA Solving: Uses AI-based services to bypass human verification systems. - Data Normalization & Enrichment: Ensures consistency across fields and formats. - AI-powered automation: Adapts scrapers automatically when websites change structure. - Reliable proxy solutions are key to maintaining scraper uptime and accessing region-specific content without getting blocked. - Whether you're scraping thousands of pages or targeting geo-restricted websites, scalable proxy solutions help bypass IP restrictions and ensure consistent delivery. 🔐 With compliance in mind, good scrapers also include governance for data protection regulations like [GDPR ](https://gdpr.eu/what-is-gdpr/?ref=blog.datahut.co)and CCPA. ## Website Change Detection In fast-moving industries, knowing when a competitor changes something can be just as valuable as knowing what they changed. Modern web scrapers can be configured to detect changes in product descriptions, prices, or SEO elements like meta tags and structured data. This helps brands respond faster — adjusting pricing, messaging, or offers in near real time. ## Data Quality & Validation Techniques Scraping at scale means dealing with messy, inconsistent data. Robust web scraping systems include: - Schema validation to ensure data structure matches expectations. - Outlier and anomaly detection. - Duplicate removal and format consistency. This ensures your business doesn’t just collect data — it collects clean, decision-ready data. ## [Is Web Scraping Legal?](https://www.blog.datahut.co/post/is-web-scraping-legal/) ### Legal Boundaries Web scraping operates in a gray area—but there are well-defined lines when it comes to ethical practice: - Stick to publicly accessible data. - Respect robots.txt and Terms of Service. - Don’t republish copyrighted text or images. - Be cautious with personal data—ensure [GDPR/CCPA compliance.](https://www.blog.datahut.co/post/guide-to-legal-and-transparent-data-practices-in-web-scraping-under-gdpr/) - Understand and adhere to data license agreements when applicable. ### Ethical Guidelines - Throttle requests to avoid overloading servers. - Avoid scraping private or sensitive content. - Be transparent in research, academic, or journalistic contexts. As regulations evolve, ethical data scraping practices are becoming a key differentiator for enterprise-grade providers. Companies are moving toward AI-enhanced data scraping systems that adapt dynamically to changes in site structure. Web scraping done right is legal, ethical, and powerful. It’s all about intent and implementation. ## API vs. Web Scraping — Which One Should You Use? ![API vs Web scraping](https://www.blog.datahut.co/content/images/2026/07/img-321.png.webp) APIs are ideal when they exist and offer the data you need. But they often come with limitations like: Hybrid strategies work best—use APIs where possible, but don’t hesitate to use scraping for broader or richer datasets. Unlike static APIs, web data scraping allows you to extract exactly what you see on a live webpage, regardless of how the data is presented. ## Automating Data Collection at Scale ### Why Automation Is Essential Manual scraping can’t scale. That’s where automation in data collection makes a difference: - Scheduled updates - Real-time syncs - Error handling & retries ### Tools for Automation - Scrapy + Cron Jobs: Run spiders on a schedule. - Apache Airflow / Prefect: Manage scraping within full ETL workflows. - n8n / Zapier: Send data to Google Sheets, CRMs, or Slack. - Proxy Managers: Handle IP rotation and ban prevention. - Datahut Platform: Fully automated with quality checks, real-time data delivery, and visual alerts. - Effective proxy management tools help rotate IPs, monitor usage, and detect bans in real-time. - Integrating smart proxy management into your scraping workflow ensures higher uptime and fewer disruptions. Smart scrapers also detect anti-bot mechanisms like honeypots, rate-limit traps, and behavior detection scripts and adjust scraping speed or switch proxies automatically to remain undetected. ## Turning Scraped Data into Insights Data becomes valuable when it's transformed into insight. That’s where visualization comes in. ### Tools to Use - Google Sheets / Excel: For fast dashboards. - Power BI / Tableau / Looker Studio: Enterprise-grade visual reporting. - Python (Plotly, Seaborn, Matplotlib): Custom visuals and exploratory analysis. ### Use Cases - Price trend lines across competitors - Heatmaps of stock availability - Word clouds from customer reviews ## Real-World Examples - A DTC fashion brand scrapes Zara and H&M to track color trends every week. - A cosmetics company uses reviews from Sephora for product development. - An electronics retailer automates Amazon price matching using scraped data and AI. - A SaaS product enriches lead scoring with scraping + data analytics solutions. - A travel aggregator scrapes airline and hotel websites to optimize dynamic pricing. - A legal research firm scrapes court websites for case filings and docket updates. - A fintech startup collects bank rate and fee data across geographies for comparison tools. - A real estate portal scrapes listings to monitor price shifts, availability, and trends. - A hiring platform scrapes job boards for competitive salary insights and skill gaps. - A CPG brand scrapes shelf placement and visibility across online marketplaces. ## Final Thoughts: The Future of Web Scraping ### What You Now Know - What web scraping is, and how it works - When to use APIs vs. scrapers - How AI, automation, and HTML parsing power modern extraction - The importance of legal compliance and ethical use - How scraped data feeds into BI tools and decision-making ### What’s Next - More AI-driven data extraction - Domain-specific Web Scraping APIs - Built-in compliance with global data protection regulations - Tighter integration with cloud-based data analytics solutions ## Work With Datahut At Datahut, we’ve helped hundreds of companies from [e-commerce](https://www.blog.datahut.co/post/web-scraping-asos-data-insights-into-pricing-promotions-and-product-diversity/) to real estate to SaaS—extract reliable, scalable, and ethical data. With built-in automation, visual alerts, proxy handling, and API-ready delivery, our platform is made for teams that rely on data, not just collect it. A Smart Proxy Manager automates IP rotation, detects suspicious patterns, and intelligently routes traffic based on website sensitivity. At Datahut, we use a Smart Proxy Manager as part of our infrastructure to ensure consistent delivery and reduce scraping friction. ### FAQ Section Q1\. What are web scraping services? A1: Web scraping services help businesses automatically collect data from websites in a structured format. These services extract product prices, reviews, market trends, and other valuable information that can be analyzed to make informed business decisions. Q2\. Why do web scraping services matter for businesses today? A2: Web scraping services empower businesses with real-time insights from competitors, customers, and market trends. This enables smarter pricing, better product positioning, and data-driven strategies — essential for staying competitive in digital markets. Q3\. Is web scraping legal?A3: Web scraping is legal when done responsibly and ethically, following website terms of service and data protection laws like GDPR. Partnering with trusted providers such as Datahut ensures compliance with all relevant regulations. Q4\. What industries benefit most from web scraping services? A4: Web scraping is useful across multiple industries including e-commerce, real estate, travel, finance, and retail. It helps track pricing, monitor competitors, understand customer sentiment, and uncover hidden market opportunities. Q5\. Do I need coding knowledge to use web scraping services? A5: Not at all! Professional scraping services like Datahut handle all the technical details for you. You simply specify the data you need, and we deliver it in ready-to-use formats such as CSV or JSON. ### How B2B Manufacturers Use Web Scraping to Convert URL: https://www.blog.datahut.co/post/how-b2b-manufacturers-converts-high-intent-prospects-to-customers-using-web-scraping/ Last updated: 2026-09-07T09:44:26.000Z Imagine this scenario: a maintenance engineer at a food-processing plant needs a specific type of industrial valve—right now. They search online, find your company’s website (you manufacture and sell valves), and click through to your product page. But when they arrive, key information is missing or outdated: the material grade isn’t clear, flow rate specifications are incomplete, and the connection dimensions are only in an old PDF. Frustrated, they click away and never return—before you even had a chance to quote or qualify them. That’s a high-intent B2B buyer lost in seconds, all because your product information wasn’t up to date. Meanwhile, unbeknownst to you, a competitor’s equivalent valve is temporarily sold out. Had your website shown accurate specs and real-time availability, you could have priced your in-stock valve slightly higher—knowing the buyer’s urgency—and captured a profitable sale. Instead, you left revenue on the table and gave your competitor an opening to build trust with that customer for future orders. In the world of low-traffic, high-intent B2B manufacturing websites, you often see only a handful of qualified visitors each day. If any of them leave disappointed, you’ve missed a critical opportunity. To prevent this, every industrial valve manufacturer needs two fundamental systems: Below, we’ll walk through why these two pillars are essential, how they work in practice, and what happens when you get them right (or fail to). ## 1\. High-Intent Visitors Demand Immediate Clarity ### 1.1 Why “Low Traffic” Doesn’t Excuse Outdated Data On a B2B industrial website especially one selling critical components like valves—incoming traffic may be modest: perhaps only a few dozen visits per day. But unlike mass-market consumer sites, B2B visitors arrive with purchasing power and tight requirements: - Precision matters. A plant maintenance engineer needs to know that your valve meets exact specifications—material (e.g., 316 stainless steel), pressure rating (e.g., up to 300 psi), temperature range (e.g., −20 °C to 200 °C), and connection type (e.g., ½″ NPT). - They’re on a deadline. If a valve fails on a production line, downtime can cost tens of thousands of dollars per hour. Every minute counts. - They expect transparency. They’ll compare multiple suppliers, request detailed data sheets, and often need to confirm compatibility with their existing piping. If your website doesn’t clearly provide that information, they’ll move on immediately. Because their intent is high, even a single oversight—like an outdated pressure curve or a missing flow coefficient—can instantly erode trust. They won’t “bookmark and come back later”; they need answers in real time. And once you lose them, recapturing that interest is extremely difficult. ### 1.2 Real-World Cost of a Single Lost Opportunity Let’s say your typical valve sale is 500 units at $25 each—a $12,500 transaction. If you lose just one qualified prospect per week because of outdated product data, that could be over $650,000 in lost revenue each year. Add to that the indirect costs: - Wasted marketing spend. You may run SEO or targeted industry campaigns to drive qualified leads. If even 10 % of those leads leave because your site data is stale, that’s money down the drain. - Brand reputation damage. In tightly-knit industrial sectors, word travels quickly. A single unhappy plant manager might tell their peers that your website “didn’t have the data I needed,” making other potential buyers hesitate. Clearly, accurate, up-to-date product information is non-negotiable. Even if you only see 20 qualified sessions per day, missing data on just 2–3 key product pages can turn away half your prospects. ## 2\. Always-Accurate Product Information: Your First Line of Defense ### 2.1 What “Accurate Product Information” Looks Like For an industrial valve manufacturer, an ideal product page must include: 1. Precise Technical SpecsValve type (e.g., globe, ball, butterfly).Material (e.g., 316 stainless steel, carbon steel, brass).Pressure rating (e.g., 300 psi, Class 150); temperature range (e.g., −20 °C to 200 °C).Connection details (e.g., ½″ NPT female, flange dimensions per ANSI B16.5).Flow coefficient (Cv) or Kv. 2. Current Availability & Lead TimeIn-stock quantity (e.g., “250 units available for immediate shipment”).Typical lead time for custom orders or high-volume runs (e.g., “4–6 weeks for orders over 1,000 units”).Any supply chain notes (e.g., “Due to raw material shortages, brass valves may have a 2-week delay”). 3. Minimum Order Quantity (MOQ) & Volume Pricing TiersClearly state MOQs (e.g., “MOQ: 10 units; volume pricing: 10–49 units @ $25 each, 50–99 @ $22 each, 100+ @ $20 each”). 4. Certifications & ComplianceRelevant standards (e.g., “Meets ANSI B16.34, API 600”; “CE-marked”; “ISO 9001 certified”).Material certifications (e.g., 316 stainless steel traceable to ASTM A351). 5. High-Quality Images & MediaClear photos showing valve body, internal trim, and connection faces.360° or exploded-view diagrams when possible to illustrate internal parts. 6. Downloadable Technical FilesDetailed datasheets (PDF) with dimensional drawings.3D CAD files (STEP, IGES) so engineers can perform fit checks in their CAD software.Installation and maintenance manuals. 7. Usage Notes & Typical ApplicationsExplain common use cases (e.g., “Ideal for chemical processing, water treatment, and steam control”).Provide best practices (e.g., “For high-temperature service, use stainless-steel ball seals”). 8. Customer Testimonials & Case StudiesShort quotes from existing clients (e.g., “We’ve been using these valves on our piping skids for two years with zero failures”).Highlight any notable projects (e.g., “These valves were selected for the new refinery expansion based on reliability in sour service”). When all these elements are accurate and refreshed regularly, you signal to B2B buyers that you understand their exact needs—and you remove every obstacle that might make them click away. ### 2.2 How to Keep Product Information Fresh 1. Implement a Product Data Governance ProcessWhenever engineering or quality updates a valve spec (e.g., a new pressure rating or material change), there must be an immediate handoff to your e-commerce/content team. Use a simple workflow (e.g., a shared spreadsheet or ticketing system) to track “Pending Updates” → “Review” → “Published.”Assign a “Data Steward” who audits 5–10 SKUs weekly to confirm specifications match the latest engineering data and regulatory standards. 2. Leverage AutomationIf you have an ERP or PIM ([Product Information Management](https://pimcore.com/en/?ref=blog.datahut.co)) system, integrate it so that any spec change in the ERP automatically pushes to the site (via API or daily CSV sync).Use version control for datasheets and manuals: display “Last Updated: MM-DD-YYYY” so buyers see you’re current. 3. Schedule Regular AuditsOnce a month, perform a “catalog health check” on 10 randomly selected SKUs: open the product page, compare specs to the latest internal documents, verify images load properly, and confirm pricing tiers match your backend system. Any discrepancies get flagged immediately. 4. Offer Multiple Data FormatsEngineers value both HTML tables and downloadable PDFs/CAD. Providing all common formats prevents friction. If a competitor’s site offers interactive 3D viewers, consider adding that feature as well—especially for high-value, engineered valves. By adopting these practices, you ensure no qualified prospect leaves your site because they can’t find the information they need. ## 3\. A Real-Time Competitor Price & Stock Monitoring System: Your Margin Multiplier Even with perfect product data, you still need to price competitively and know who’s in or out of stock across the market. Here’s why: - B2B buyers compare multiple suppliers before purchasing. In a typical “request-for-quote” (RFQ) process, they’ll benchmark three to five vendors—looking at price, lead time, and specs. If they see your valve is 5 % more expensive than a rival’s, they’ll likely buy from the cheaper source, even if your specs are identical. - Stock-outs create urgency—and opportunity. If a competitor’s valve is marked “temporarily unavailable” on their site, buyers with an urgent need may pay a small premium for your in-stock valve rather than face downtime. - Material cost swings can erode margins quickly. Suppose the price of stainless steel spikes due to supply disruptions. If you’re still listing old, lower prices on your site, you risk selling at a loss or having to retract quotes. Conversely, if competitor costs rise but your real-time monitoring system flags it, you can proactively adjust your price to protect margin before the buyer even notices. ### 3.1 How a Competitor Monitoring System Works 1. Identify Key Competitor SKUsStart by mapping your valve’s part number (e.g., IV-316-½NPT) to each competitor’s equivalent SKU or description (e.g., “316 SS ½″ NPT Ball Valve”). Include flow coefficient, connection type, and material. 2. Set Up Automated Scraping or API FeedsUse a specialized tool or build a custom scraper to pull competitor pricing and stock data at a set cadence—ideally every 4–6 hours for high-demand SKUs, and daily for less critical ones.Monitor both competitor websites and major distributors (e.g., MSC Industrial, Grainger, Amazon Business). Normalize their listings so you see “$24.50 each” vs “$23.75 each” for the same valve. 3. Normalize & CompareConvert all prices to a consistent unit (price per valve) and account for volume breaks. If Competitor A offers a 100-unit pack for $24 each and a 500-unit pack for $22 each, capture both tiers and know which is relevant to the buyer’s potential order volume. 4. Trigger Alerts & ActionsDefine thresholds for action:Price Drop Alert: “Notify me if any competitor dips below our cost + 5 % gross margin.”Stock-Out Alert: “Notify me if competitor inventory < 50 units.”Alerts can go to email, Slack, or a pricing dashboard. More advanced setups can automatically adjust your own site price within guardrails you set (e.g., never go below $23 per valve). ### 3.2 Why You Need Both Price and Stock Signals - Capturing Stock-Out DemandWhen a competitor’s valve goes out of stock, buyers often have an urgent need—if they can’t wait weeks, they’ll accept a small price premium for immediate delivery. A stock-out alert lets you temporarily raise your price (e.g., +5 %), protecting margin while meeting that urgent demand. - Avoiding Unnecessary Under-PricingIf a competitor heavily discounts leftover inventory (e.g., $18 per valve instead of $25), your system flags that as an outlier. Your pricing manager can decide whether to match (at thin margins) or hold firm, preserving profitability. - Maintaining Margin IntegrityBy continuously comparing real costs—especially when raw material costs fluctuate—you align your pricing so that you never sell below your target margin. If your steel cost increases by $2 per valve, you can update your price before a competitor undercuts you, preventing losses. ### 3.3 Quick Implementation Steps 1. Compile Your SKU Master ListList all valve SKUs, including part number, description, material, connection type, and volume tiers. 2. Identify Competitors & Distribution ChannelsDecide which direct competitors (e.g., local valve manufacturers, national suppliers like Parker, Swagelok) and which distributors or marketplaces (e.g., Amazon Business, Grainger, Fastenal) to monitor. 3. Choose a Monitoring Tool or Build a Custom ScraperOff-the-shelf platforms like Price2Spy can track both price and stock. For highly specialized valves, you may need a custom scraper to handle different website structures. 4. Define Alert RulesSet pragmatic rules, for example:Alert if competitor price < cost + $1 (ensuring you never price below cost).Alert of competitor stock < 50 units for a high-turnover valve.These rules generate actionable notifications rather than noise. 5. Integrate With Pricing WorkflowsDecide if adjustments are manual (pricing manager reviews alerts and updates your site) or semi-automated (alerts feed into a pricing dashboard, and you can approve rapid updates). For stock-out opportunities, you might allow automated temporary price hikes within a 5 % margin window. By implementing this system, you ensure you’re always in sync with the market—defending margin when needed and seizing premium windows when competitors run out. ## 4\. Using Rb2b / [Happierleads](https://happierleads.com/?utm%5Fcampaign=PV233UDdMGpO&gad%5Fsource=1&gad%5Fcampaignid=22199563654&gbraid=0AAAAAoVHqWvpxjK2sP1wIF4oS7d6rKdIr&gclid=EAIaIQobChMInN6roe2GjgMV%5FqNmAh2RyCugEAAYASAAEgKcXPD%5FBwE&ref=blog.datahut.co) to [Identify High-Intent Visitors](https://www.blog.datahut.co/post/how-to-leverage-data-to-rank-first-on-amazon-proven-strategies-for-maximum-visibility/) If a maintenance engineer or other B2B buyer lands on your site but doesn’t convert, Rb2b / Happierleads can help you re-engage them by revealing their company and contact details. Here’s how it works: Install the Rb2b / Happierleads Tracking Snippet • Add the JavaScript snippet Rb2b / Happierleads provides into your website header. • Whenever someone visits—especially key product pages for industrial valves—their IP address and behavior (pages viewed, time on page) get matched against Rb2b / Happierleads’s firmographic database. Capture Firmographic & Contact Data • Rb2b / Happierleads surfaces details like company name, industry, employee count, and sometimes even an approximate revenue range. • If Rb2b / Happierleads can map that visitor to a specific business email pattern (e.g., tony@retailvalves.com), you’ll see a probable contact address—no form fill needed. Prioritize High-Intent Signals • Filter Rb2b / Happierleads alerts so that you only get notified for visits to your most critical product pages (e.g., any page under /valves/316-ss-ball-valve). • Set a threshold (e.g., visitors who view ≥3 SKU pages in a single session) and flag those as “high intent.” Enrich Your CRM & Automate Follow-Up • Send Rb2b / Happierleads’s data directly into your CRM (e.g., Salesforce, HubSpot). Tag records based on pages viewed (e.g., “Viewed 316-SS Ball Valve Specs”). • Trigger a personalized outreach: “Hi \[Name\], I noticed your team at \[Company\] was researching our 316 SS ball valves—any questions about Cv ratings or material certifications? Happy to share the latest datasheet or CAD model.” Measure & Refine • Track conversion rates on follow-up emails sent to Rb2b / Happierleads-identified leads. If certain product pages generate more qualified matches, put extra emphasis on keeping those pages perfectly up to date. • Use A/B testing on subject lines and messaging—mentioning “Updated ISO 9001 Datasheet” vs. “Quick Question on 316 SS Valve Specifications” to see what resonates best. By leveraging Rb2b / Happierleads, you never lose sight of a high-intent engineer who bounces before filling out a form. Instead, you get firmographic context and a direct line of contact—turning anonymous visits into actionable sales conversations. ## 5\. Putting It All Together: A Day in the Life Let’s revisit our original maintenance engineer, but now with your “perfect” website and monitoring system in place: 1. Search & Click: They type “316 stainless steel ½″ NPT ball valve” into Google and land on your site. 2. See Accurate Specs: The product page clearly states “316 SS Ball Valve, ½″ NPT Female, 300 psi, Cv = 4.0, ISO 9001 certified, in stock: 150 units.” Exactly what they need. 3. Check Price & Availability: Price is $24 each for orders of 50–99 units, $22 each for 100+ units. Your competitor monitor flagged that their usual supplier is sold out—your stock-out alert triggered a temporary price increase to $24 (instead of your normal $20), knowing the buyer has no other immediate option. 4. Download CAD & Buy: They download the CAD file to verify fit in their piping layout, confirm everything matches, and place an order for 75 units at $24 each. 5. Win the Order: You capture a $1,800 sale. Because your specs were clear, they didn’t bother requesting data from multiple suppliers. Since you were in stock and priced competitively (despite the premium), they clicked “Add to Cart” without hesitation. Sale closed in under five minutes of browsing. Contrast that with the “old” scenario—outdated site, no real-time pricing—where the buyer bounces after 10 seconds of frustration. That single transaction would have been $1,800, plus potential repeat quarterly orders. The difference between “perfect” and “old” is the presence of two systems: 1. Accurate Product Info that answers every technical question upfront. 2. Real-Time Pricing & Stock Monitoring that lets you seize the moment when competitors falter. ## 6\. Key Takeaways & Next Steps 1. Treat Low Traffic as an Asset, Not an Excuse. Even if your B2B valve website only sees 20–30 qualified visits per day, each visitor likely represents a $1,000–$10,000 purchase. Losing just a couple because of stale data or uncompetitive pricing can outweigh any monthly ad spend. 2. Build a Data Governance Process for Product Info. Assign an owner, create a documented workflow, and leverage automation (ERP or PIM integrations) to ensure specs, images, and downloadable assets are always fresh. 3. Invest in Competitor price & Stock Monitoring. Implement a tool or custom scraper that tracks key competitor websites and major distributor listings every few hours for both price and inventory status. Set up alert rules for significant price swings and stock-outs. 4. [Turn Monitoring Alerts into Action](https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/). Decide in advance how to react: match a lower price, hold if margin is too thin, or raise price when competitors run out. Keep your margins healthy while capturing urgent orders. 5. Continuously Measure & Refine. Track KPIs such as “conversion rate on valve product pages,” “time to adjust price after competitor changes,” or “number of orders captured during competitor stock-outs.” Use those metrics to fine-tune alert thresholds and product page content. By putting these systems in place, you transform every high-intent visitor into a potential sale, rather than a missed opportunity. And because B2B traffic tends to be low but extremely qualified, keeping your valve specifications up to date and your prices aligned with the market is the fastest path to higher margins, greater win rates, and sustained growth. Every high-intent visitor represents a significant revenue opportunity. By ensuring your product pages are always up-to-date with accurate specifications and by pairing that with a near real-time competitor price and stock monitoring system, you eliminate guesswork—and turn qualified visits into closed deals. Low traffic is not a limitation; it’s a chance to capture high-value orders one at a time, confidently knowing you’re offering the most reliable data and the most competitive pricing in the market. Ready to Get Started? At [Datahut](http://www.datahut.co/?ref=blog.datahut.co), we specialize in empowering industrial manufacturers with: - Automated Web Scraping & Data Aggregation—We build and maintain custom scrapers that pull precise product specs, pricing, and availability data from multiple sources. - Real-Time Competitive Monitoring—Our tools track price shifts and stock levels across distributors and rival sites, so you can adjust your pricing strategy proactively. - Data Governance & Integration—We help you set up workflows (via PIM/ERP integrations or custom APIs) to keep product information fresh, synced, and audit-ready. - High-Intent Visitor Identification—With integrations like Rb2b and Happierleads, we enrich lead data and alert you the moment a decision-maker browses your site, turning anonymous visits into actionable outreach. Don’t let outdated information and static pricing cost you another high-value industrial order. Start today by auditing your product pages and setting up a simple competitor monitoring pilot. You’ll quickly reclaim lost revenue and lock down those high-intent buyers before your competition even notices they’re gone. Author I’m Tony Paul, founder of Datahut. I’ve been working in the web scraping industry for over 14 years. Through our work with B2B manufacturers across various industries, my team and I have discovered how real-time data can play a pivotal role in converting high-intent leads into loyal customers. If you’re looking to leverage data to optimize your sales pipeline and gain a competitive edge, get in touch with us using the chat widget on the right-hand side. ### Why Web Scraping WebMD Drug Data Is Essential URL: https://www.blog.datahut.co/post/why-is-web-scraping-essential-for-extracting-drug-information-from-webmd/ Last updated: 2026-09-07T09:44:28.000Z [WebMD](https://www.webmd.com/?ref=blog.datahut.co) is the best source of health and medical information, with lots of resources related to health, wellness, and medical problems. The website is also extremely popular due to its well-researched content, expert-approved articles, and beneficial tools for users to find reliable health information. It is one of those trusted sources providing general advice on wellness, in-depth explanations of symptoms, treatments, and other medical procedures. Many people apply it to obtain knowledge relating to various medical issues. Professionals in their field employ it for the latest medical information and insights. This is where web scraping comes into play. If you're new to it, check out our [guide to web scraping with](https://www.blog.datahut.co/post/python-web-scraping-tutorial/) Python to understand the basics. Web scraping is basically getting information from websites automatically through tools and scripts that operate like an individual browsing on the Internet. It really helps individuals gather a considerable amount of information from the web page without having to copy-paste it all manually. Web scraping can collect both organized and unorganized data, such as product details, prices, reviews, or text. This data can be applied for analysis, reporting, or adding to other systems. Many industries, e-commerce, data analysis, and research, use this for easy understanding of online data. However, it needs to follow the rules, legal guidelines of the site while collecting content. Choosing the right [web scraping service](https://www.blog.datahut.co/post/how-to-choose-the-best-web-scraping-service/) is essential to ensure accuracy and compliance when dealing with healthcare data. There are two steps to web scraping in this project in order to make them efficient and accurate in collecting drug information pages from WebMD. The first step is the collection of product URLs from WebMD's drug index. This begins with going through the alphabetical index, wherein each letter (A-Z) and number category opens a new webpage with various drugs. For each one of these pages, a scraping tool takes the URLs that link to the individual drug pages. This creates an entire database of product URLs, ready to be used in the next process of gathering data. SQLite is used as a data storage medium where not only the product links are kept track of, but the processing status is also shown for each of the links. This approach, being systematic, helps us to tackle a tremendous amount of data and revisit the unformatted links whenever required. Once the URLs are gathered, then the next step is to extract the detailed information from each product page. This is basically getting the essential drug details, such as drug names, generic names, uses, side effects, warnings, precautions, interactions, and overdose information. These important fields from each product page are scraped and then saved in a structured format in an SQLite database. This way, the data collected is well organized, which is good for many tasks later on. Also, any pages that cannot be scraped are noted in a different table, so they can be tried again or looked into more, making sure no data is lost. This two-step method of scraping is crucial to collect a large amount of well-organized, detailed data from complex websites such as WebMD. It provides both speed and a robust way to handle any problems that may arise during the process of scraping. ## An Overview of Libraries for Seamless Data Extraction In the case of the WebMD web scraping project, several types of Python libraries are utilised to assist in bringing together product links and obtaining final data from the scraped pages. Here is a detailed breakdown of the libraries that each step uses ie scraping the product links and finally collecting the data. ### Requests This is one of the most widely used libraries while making HTTP requests in Python. Both parts of the code use requests to make GET requests to the WebMD website. The right headers and cookies used with the requests library simulate a browser, helping the scraper get the web pages needed without being blocked by the security of the site. ### BeautifulSoup Parsing the HTML documents to extract needed data is through BeautifulSoup, an implementation from the bs4 library. Once requests have fetched a webpage, it takes over and starts parsing its HTML content. This would enable the script to search and fetch specific elements such as anchor tags or div containers. In product link scraping code, it identifies the section carrying product links. In the final phase of data scraping, it captures more detailed information such as the names of drugs, uses, side effects, and much more on every page of the product. SQLite3 is an SQLite database library. It presents an easy and lightweight way to put scraped data into a database. As SQLite3 creates a SQLite database called webmd\_webscraping.db and a SQLite table called product\_links at which the product links get kept in the process of scraping product links. In the final step of data scraping, the same database is expanded. A new table called final\_product\_data is created to hold the detailed drug information. Another table named failed\_urls is created to log any links that do not load. This helps in storing the data efficiently and reliably in the process of scraping. Time is used to give pauses between web requests. It is helpful in ensuring that the scraper does not make too many requests in a short period, which might make it look suspicious and may get blocked. This product link scrape code in addition, adds a delay between requests by using time.sleep(2) so not to overwhelm the WebMD servers. Adding some random sleep time between requests in the final data scrape code will simulate more like human browsing which avoids getting detected. The final data scraping script uses the random function for selecting one of the user agents from the list at a random. By changing up the user agents, each time the scraper sends in requests, it looks like something different - either browsers or devices. This is difficult for the website to have identified it as some bot, and random choice offers another layer of protection against being flagged for IP bans or hitting limits on requests from a target website. These libraries work together to make a powerful flexible web scraping tool. This tool, in turn, collects data on products from WebMD; organizes the collected data using SQLite; and follows steps to ensure that the act of the scraper mimics a real person by alternating user agents and adding time delays. ## STEP 1 : Product link scraping ### Importing Libraries ``` import time import requests from bs4 import BeautifulSoup import sqlite3 ``` This code imports essential libraries for web scraping and data handling. It includes requests for making HTTP requests, random and time for adding delays, sqlite3 for database interaction. ### Defining the Base URL ``` # Base URL for WebMD BASE_URL = "https://www.webmd.com" ``` The BASE\_URL is the variable where you're going to store the web address of the root website. In this case, BASE\_URL points to "https://www.webmd.com," the core website from which you'll begin scraping drug-related information. Store the base URL separately, so while scraping, you can append different paths to create full URLs for different pages. This makes the code pretty flexible and easy to maintain whenever the base URL changes. ### Configuring HTTP Headers and Cookies for Requests ``` # Headers and cookies for HTTP requests HEADERS = { 'authority': 'www.googletagservices.com', 'method': 'GET', 'scheme': 'https', 'accept': '*/*', 'accept-encoding': 'gzip, deflate, br, zstd', 'accept-language': 'en-IN,en-GB;q=0.9,en-US;q=0.8,en;q=0.7,ml;q=0.6', 'cache-control': 'no-cache', 'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/127.0.0.0 Safari/537.36', } COOKIES = { 'name': 'VisitorId', 'value': '3b80b89d-7007-4be2-84fe-377d67abff64', 'domain': '.webmd.com', 'path': '/' } ``` The HEADERS and COOKIES dictionaries are used to mimic a real user's browser when making HTTP requests to the WebMD website. These components help avoid being blocked by the website and ensure that responses are served correctly. - HEADERS: Defines the metadata sent with the request, such as user-agent (which tells the server what browser and operating system are being used) and accept-encoding (specifying the supported response encoding). This helps in simulating a legitimate browser request and managing caching. - COOKIES: Represents the stored data from previous interactions with the website, such as user session information. These values are important for maintaining state across multiple requests and may be required to access certain pages or resources on WebMD. Together, the headers and cookies make the web scraping process more seamless and efficient by reducing the chances of being detected as a bot. ### Fetching and Parsing Webpage Content with BeautifulSoup ``` def get_soup(url): """ Sends an HTTP GET request to the provided URL, retrieves the HTML content of the page, and parses it into a BeautifulSoup object. Args: url (str): The URL of the webpage to retrieve. Returns: BeautifulSoup: Parsed HTML of the page. """ response = requests.get(url, headers=HEADERS, cookies=COOKIES) return BeautifulSoup(response.text, 'html.parser') ``` The get\_soup function fetches and parses the content of the webpage with the help of requests and BeautifulSoup libraries. It sends an HTTP GET request to the given URL, using predefined headers and cookies for the simulation of a browser session so that proper access to the page is possible. The HTML content of the response gets extracted and parsed into a BeautifulSoup object. The parsed object allows the navigation of the HTML structure, where the extraction of data and searching operations by tags, classes, or attributes can be made possible. It's easier to work with web pages while scraping by returning a ready-to-use BeautifulSoup object. ### Extracting Unique Links from a Specific Web Page Section ``` def scrape_links_from_list(soup): """ Extracts all unique links from a specific div in the parsed HTML. This function looks for anchor (``) tags within a div element that has the class 'drugs-search-list-conditions'. It collects the `href` attributes of those anchor tags, prefixes them with the base URL, and returns a set of full URLs. Using a set ensures uniqueness, so duplicate links are automatically removed. Args: soup (BeautifulSoup): The parsed HTML content of a webpage, represented as a BeautifulSoup object. This object is usually created after fetching and parsing the webpage content. Returns: set: A set containing unique URLs (as strings) extracted from the specified div. Each URL is prefixed with `BASE_URL` to form a complete link. """ div = soup.find('div', class_='drugs-search-list-conditions') if div: return {BASE_URL + a['href'] for a in div.find_all('a')} return set() ``` scrape\_links\_from\_list : The function concept is supposed to find all the unique links that occur in some class called 'drugs-search-list-conditions' of an HTML page. This function locates all the tags which are found inside a div that has class 'drugs-search-list-conditions'. Then it picks up their href attributes which happen to be the relative paths of links. Then, these relative paths are combined with an already defined variable called BASE\_URL to generate an absolute path. And all returns as a set, thus making sure only a link gets stored without repeating themselves. In case a page doesn't contain the targeted div element, the function returns an empty set. Such a method comes in handy when one would want to scrape only meaningful links from structured sections of the webpage. ### Processing Sub-Alpha Links and Storing Unique Product URLs in the Database ``` def process_subalpha_links(alpha_soup, cursor, conn, unique_links): """ Processes sub-alpha links from the current alphabet page and extracts unique product links for database insertion. This function handles the second layer of alphabet-based navigation. It identifies 'sub-alpha' links from the current page and processes each sub-alpha link to extract product links. These links are checked for uniqueness using the `unique_links` set, and new links are inserted into the `product_links` table. Args: alpha_soup (BeautifulSoup): Parsed HTML of the current main alphabet page. cursor (sqlite3.Cursor): SQLite cursor for executing SQL queries. conn (sqlite3.Connection): SQLite connection to commit changes to the database. unique_links (set): Set of unique product links to avoid duplicate entries. Function Workflow: 1. Looks for the 'sub-alpha' section on the alphabet page. 2. Fetches each sub-alpha page and extracts product links. 3. Inserts new, non-duplicate links into the database and updates the set. 4. Commits changes to the database after processing all links. """ subalpha_ul = alpha_soup.select_one( '.alpha-container.subalpha-container .browse-letters.squares.sub-alpha.sub-alpha-letters' ) if subalpha_ul: for subalpha_link in subalpha_ul.select('li.sub-alpha-square a[href]'): subalpha_url = BASE_URL + subalpha_link['href'] subalpha_soup = get_soup(subalpha_url) final_links = scrape_links_from_list(subalpha_soup) # Update the unique links set and insert new unique # links into the database new_links = final_links - unique_links unique_links.update(new_links) for link in new_links: cursor.execute( "INSERT INTO product_links (product_link) VALUES (?)", (link,) ) conn.commit() ``` The process\_subalpha\_links function is meant to handle a deeper level of navigation on an alphabet-page-based webpage. That is after it lands on a page 'A', 'B' etc. it looks for subdivision in that page. Once found, the sub-alpha link directs the user to even deeper pages where the target URLs are found. The extracted URLs of the products is added into the SQLite data but only unique ones whose previous insertion has not happened yet. The function first scans all the sub-alpha sections in the parsed HTML, using BeautifulSoup. For every sub-alpha link, it then fetches and parses the page, getting all the links of the products. Checks these extracted links against the unique\_links set to avoid adding duplicate entries in the SQLite database. Every new link is added to the product\_links table, and a new update is done for the unique links set with regard to the newly added URL. Changes to the database are made at the last stages so that links to products can be saved. This approach effectively manages the hierarchy of links and ensures that product URLs are stored without duplications, making the scraping process better. ### Creating and Initializing the SQLite Database for Product Links ``` def create_database(): """ Initializes the SQLite database and creates the necessary table for storing product links. This function checks whether the SQLite database file (`webmd_webscraping.db`) exists, and if not, it creates it. Within the database, it ensures that the `product_links` table is present, creating it if it does not already exist. The `product_links` table contains the following columns: - `id`: An auto-incremented integer that serves as the primary key. - `product_link`: A unique text field that stores the scraped product URLs. - `status`: An integer field that defaults to 0, indicating whether the link has been processed (e.g., 0 for unprocessed and 1 for processed). After ensuring that the table is set up, the function returns the database connection and cursor objects for use in subsequent database operations. Returns: tuple: A tuple containing: - conn (sqlite3.Connection): A SQLite database connection object, which can be used to manage transactions and commit changes to the database. - cursor (sqlite3.Cursor): A SQLite cursor object, which is used to execute SQL queries. """ conn = sqlite3.connect('webmd_webscraping.db') cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS product_links ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_link TEXT UNIQUE, status INTEGER DEFAULT 0 ) ''') conn.commit() return conn, cursor ``` The create\_database function sets up the foundation for storing scraped product links by creating an SQLite database and establishing a table for managing the URLs. The function checks if the database file (webmd\_webscraping.db) exists, creating it if necessary. Within the database, it ensures the existence of the product\_links table, which is designed to store the URLs of product links collected during the web scraping process. The table structure includes three columns: 1. id - a primary key that auto-increments for each entry. 2. product\_link - a unique field that holds the actual product URLs. 3. status - an integer field (defaulting to 0) used to indicate whether a link has been processed (e.g., 0 for unprocessed and 1 for processed). By returning the database connection (conn) and cursor (cursor) objects, the function enables further database operations, such as inserting or querying data, in subsequent steps of the scraping workflow. Additionally, it uses the IF NOT EXISTS clause in the SQL query to avoid duplicating the table if it already exists, ensuring seamless initialization. The function plays a critical role in organizing and tracking the status of the product links during the scraping process, allowing for efficient data management and ensuring that each link is processed only once. ### Scraping WebMD Drug Links and Storing in an SQLite Database ``` def scrape_webmd_links(): """ Scrapes WebMD drug links for all alphabets and saves unique links into an SQLite database. This function iterates over all alphabet characters ( from 'a' to 'z' and '0' for numeric entries), constructs a URL for each letter on WebMD's drug index page, and retrieves the corresponding HTML content. It parses the page to find sub-alphabet links and scrapes all drug-related links from these sub-pages. Unique links are then inserted into the SQLite database for storage, ensuring no duplicates are added. Workflow: 1. The function initializes the SQLite database by calling `create_database()` and prepares a set to track unique links. 2. It iterates over each alphabet character and retrieves the corresponding WebMD URL (e.g., for 'a', it scrapes "https://www.webmd.com/drugs/2/alpha/a"). 3. For each letter page, it extracts relevant sub-alphabet links and processes those to get the final drug links. 4. Extracted links are stored in the `product_links` table, with uniqueness enforced by both the database and the set. 5. A short delay (`time.sleep(2)`) is added between each request to avoid overwhelming the server. 6. Once scraping is complete, the database connection is closed. """ conn, cursor = create_database() unique_links = set() try: for char in 'abcdefghijklmnopqrstuvwxyz0': url = f"{BASE_URL}/drugs/2/alpha/{char}" alpha_soup = get_soup(url) alpha_text_div = alpha_soup.select_one( '.alpha-text[data-metrics-module="drugs-az"]' ) if alpha_text_div: process_subalpha_links( alpha_text_div, cursor, conn, unique_links ) time.sleep(2) # Short delay between each main URL request finally: conn.close() ``` The scrape\_webmd\_links function does the following: it automatically scrapes drug-related links from WebMD's drug index, covering all alphabet characters from 'a' to 'z', as well as numeric entries ('0'). It systematically constructs URLs for each alphabet character, retrieves the corresponding HTML page, and extracts the sub-alphabet and drug links. These links are then stored in an SQLite database with a mechanism to ensure uniqueness through both a Python set and database constraints. The workflow starts off by invoking the create\_database() function, which should initialize both the SQLite database and the product\_links table. The set (unique\_links) keeps track of those links already processed in this session. Then the function enters a loop for each character fetching the page corresponding to it off WebMD. For each of the pages, it processes links to the sub-alphabets, retrieves the URL of the drug, inserts it into the database not to duplicate it. The function ensures that it does not bombard the server by placing a 2-second delay between every request. After scraping all links, SQLite connection is closed in order to ensure data integrity. This function handles the whole operation of scraping and storage within the database, so an exhaustive list of URLs regarding drugs can be retrieved with no duplication. ### Entry Point for WebMD Drug Link Scraping ``` # Start scraping and save the links to the SQLite database if __name__ == "__main__": """ Entry point for executing the WebMD drug link scraping process. When this script is run directly, the `scrape_webmd_links()` function is invoked to start the scraping process. It scrapes WebMD's drug pages, extracts unique product links for each alphabet character (a-z and 0), and stores them in an SQLite database named 'webmd_webscraping.db'. The scraped links are stored in the `product_links` table with two columns: - `product_link`: Contains the URL of the drug-related product. - `status`: A default status column (set to 0) for potential future use (e.g., marking links as processed). """ scrape_webmd_links() ``` The following block of code, if \_\_name\_\_ == "\_\_main\_\_": is the entry point for executing the process of scraping links on WebMD for drugs. This script calls the scrape\_webmd\_links() function when executed directly to start the scraping of drug pages on WebMD. For each alphabet character, it collects unique drug-related links from a-z and '0' and stores them in an SQLite database called webmd\_webscraping.db. Two major columns are present in the product\_links table. The product\_link column contains the URL of the drug product and the status column, which has been set to 0 as a default column. The status column will come into handy when tracking the processing state for each link in the future extensions of this project. To run the script, it can be run from the command line with Python, and the needed dependencies, requests, beautifulsoup4, and sqlite3 should be installed.This script will open a new SQLite database if it doesn't already exist and begin storing the scraped links in an efficient manner. ## STEP 2 : Final Data Scraping From Product Links ### Importing Libraries ``` import requests from bs4 import BeautifulSoup import time import random import sqlite3 ``` The requests library sends HTTP requests, BeautifulSoup parses HTML/XML, time handles delays, random generates random numbers, and sqlite3 manages SQLite database operations. ### Configuration Variables ``` # Configuration Variables DB_NAME = 'webmd_webscraping.db' USER_AGENTS_FILE = 'data/user_agents.txt' SLEEP_MIN = 2 SLEEP_MAX = 5 ``` The Configuration Variables section defines key parameters for the script: 1. DB\_NAME: Specifies the name of the SQLite database (webmd\_webscraping.db) where the scraped data will be stored. This is the database that will hold the drug information and any failed URLs during the scraping process. 2. USER\_AGENTS\_FILE: Refers to the file path (data/user\_agents.txt) that contains a list of user agent strings. These user agents are used to mimic different browsers when making requests to the website, helping avoid detection and blocking by the server. 3. SLEEP\_MIN and SLEEP\_MAX: These variables control the random sleep interval between consecutive web scraping requests, ranging from 2 to 5 seconds. The random delay helps prevent the server from detecting the scraping activity as automated or overwhelming it with too many requests in a short time. Together, these configuration variables ensure that the script runs efficiently, accesses the website responsibly, and stores the scraped data properly. ### Headers and Cookies to Mimic a Real Browser Request ``` # Headers and cookies to mimic a real browser request headers = { 'authority': 'www.googletagservices.com', 'method': 'GET', 'scheme': 'https', 'accept': '*/*', 'accept-encoding': 'gzip, deflate, br, zstd', 'accept-language': 'en-IN,en-GB;q=0.9,en-US;q=0.8,en;q=0.7,ml;q=0.6', 'cache-control': 'no-cache' } cookies = { 'name': 'VisitorId', 'value': '3b80b89d-7007-4be2-84fe-377d67abff64', 'domain': '.webmd.com', 'path': '/' } ``` This section sets up headers and cookies that mimic a real browser's behavior during web requests. These elements are essential for avoiding detection when scraping websites like WebMD, which may block automated scripts that don’t resemble real users. 1. Headers: The headers dictionary includes parameters like accept, accept-encoding, accept-language, and cache-control, which tell the server what kind of content to return, how it can be encoded, and which languages are acceptable. These headers simulate the browser's settings when making the request, helping the scraper appear as if it's coming from a regular browser. 2. Cookies: The cookies dictionary includes the VisitorId, which tracks the visitor's session on the WebMD site. This cookie is passed along with each request to maintain the browsing session, which might be needed to access certain content. By using headers and cookies, the scraper behaves more like a legitimate user, increasing the likelihood of successful data retrieval without being blocked by the website’s security measures. ### Database Setup and Table Creation ``` # Database setup conn = sqlite3.connect(DB_NAME) cursor = conn.cursor() # Create tables if not exist cursor.execute(''' CREATE TABLE IF NOT EXISTS final_product_data ( id INTEGER PRIMARY KEY, product_link TEXT, drug_name TEXT, generic_name TEXT, uses TEXT, side_effects TEXT, warnings TEXT, precautions TEXT, interactions TEXT, overdose TEXT ) ''') cursor.execute(''' CREATE TABLE IF NOT EXISTS failed_urls ( id INTEGER PRIMARY KEY, url TEXT ) ''') conn.commit() ``` This section connects to SQLite Database and ensures that database tables for storing scraped data along with failed URLs are created in it. The script does connect to the database using sqlite3.connect(DB\_NAME), where the actual name of the DBNAME is the database file or rather webmd\_webscraping.db. If the file didn't exist, it has created automatically. A cursor object is then initialized as SQL commands are to be executed. Two tables are created, in case they do not exist: final\_product\_data and failed\_urls. The table final\_product\_data contains all the relevant information regarding a drug, including product link, drug name, generic name, uses, side effects, warnings, precautions, interactions, and overdose details. In case of a failure to scrape any URL, it gets logged in the table failed\_urls, making it easy to track and retry that URL later. Once these commands are executed to create the tables, conn.commit() commits all of those changes to the database, ensuring that the structure of the database is correctly defined before entering the web scraping phase. ### Loading User Agents from File ``` # Functions def load_user_agents(user_agents_file): """ Loads a list of user agents from a specified text file. Args: user_agents_file (str): The path to the text file containing user agent strings, with each user agent on a new line. Returns: list: A list of user agent strings read from the file, with whitespace stripped. """ with open(user_agents_file, mode='r', encoding='utf-8') as file: return [line.strip() for line in file if line.strip()] ``` This function, load\_user\_agents, reads the user agent strings from a given text file and later used to mimic browser requests while web scraping. The user agent is an identifier used by a website to know what sort of device or browser the request is coming from. This function accepts one argument, namely the file path to a text file containing multiple user agents, each one on a new line. Within the function, this file is opened for reading with UTF-8 encoding to enable support for all languages and characters. This function iterates through every line in the file, removing leading/trailing whitespaces and then aggregates a list of user agents with no empty lines. The function then returns this list of user agents to later pick a random one for every web scraping request, avoiding the detection/blocking of requests by the targeted website. ### Selecting a Random User Agent ``` def get_random_user_agent(): """ Selects and returns a random user agent string from the pre-loaded list of user agents. Returns: str: A randomly selected user agent string. """ return random.choice(user_agents) ``` The get\_random\_user\_agent function is responsible for selecting a random user agent from a pre-loaded list of user agents. This list, generated by the previous function, contains various user agent strings that mimic different browsers or devices. By selecting a random user agent for each request, this function helps avoid detection by websites that block or throttle repeated requests from the same user agent. It returns a single, randomly chosen user agent string, which will be used to simulate a browser during web scraping activities. This strategy helps make the web scraper's behavior appear more natural and less likely to trigger anti-bot measures. ### Retrieving and Parsing Web Page Content ``` def get_soup(url): """ Makes an HTTP GET request to the specified URL and parses the HTML content using BeautifulSoup. Args: url (str): The URL to retrieve and parse. Returns: BeautifulSoup: Parsed HTML content of the requested page. """ headers['user-agent'] = get_random_user_agent() response = requests.get(url, headers=headers, cookies=cookies) return BeautifulSoup(response.text, 'html.parser') ``` The get\_soup function is defined so that it fetches the HTML of a given URL and makes a BeautifulSoup parse out of it. It starts with setting the 'User-Agent' header to a random user agent, chosen by the get\_random\_user\_agent function, to simulate an actual browser request. It then simulates a valid browser session functionally by conducting an HTTP GET request to a specified URL, passing in the headers and cookies. The received response is parsed with the use of BeautifulSoup, returning the content as a BeautifulSoup object that can be used further to extract certain data from the webpage. ### Extracting Text from HTML Div Elements ``` def extract_text_from_div(soup, selectors): """ Extracts text content from a specified HTML div element based on a list of CSS selectors. Args: soup (BeautifulSoup): The parsed HTML content from which to extract text. selectors (list): A list of CSS selectors to identify the target div elements. Returns: str: The extracted text content from the first matching div. If no matching div is found , returns "Not Available". """ for selector in selectors: target_div = soup.select_one(selector) if target_div: return target_div.get_text(separator='\n', strip=True) return "Not Available" ``` This function extract\_text\_from\_div is going to be used for extracting the text content from particular HTML
    elements contained in the parsed HTML document. It accepts two parameters: soup, which is a BeautifulSoup object that holds the parsed HTML; and selectors, which is a list of CSS selectors targeting the desired
    elements. This function iterates over the given selectors and calls the select\_one method on each one to return the first
    element found by the CSS selector. If a
    element is encountered, the function takes its text content; there are newline characters between lines to take out leading and trailing whitespace. If the string has no
    , then this function will return "Not Available" so that it can be explicitly clear whether it found or not in an open manner. This is useful to locate information with much efficiency from the web pages in their HTML structure. ### Extracting the Generic Name of a Drug ``` def extract_generic_name(soup): """ Extracts the generic name of a drug from the provided parsed HTML content. Args: soup (BeautifulSoup): The parsed HTML content of the drug's page. Returns: str: The extracted generic name of the drug. If the generic name is not found, returns "Not Available". """ generic_name = soup.find('h3', class_='drug-generic-name') if generic_name: return generic_name.get_text(strip=True).split(':')[-1].strip() drug_info_holder = soup.find('div', class_='drug-info-holder') if drug_info_holder: generic_name_li = drug_info_holder.find('li', class_='generic-name') if generic_name_li: generic_name_span = generic_name_li.find('span') if generic_name_span: return generic_name_span.get_text(strip=True) return "Not Available" ``` The extract\_generic\_name is defined to fetch the name of a drug in generics from the parsed HTML content in its web page. There is one argument, which accepts the soup-a BeautifulSoup object that contains the parsed HTML. It begins to hunt for the generic name beginning with an

    element that has its class as drug-generic-name. If this element exists, the function pulls and returns the text that comes after a colon(:) after stripping any white space. If

    does not exist, the function looks for a
    class drug-info-holder. Then it looks inside this div for a
  • class generic-name. When available, it looks within this list item for a to recover the text of the generic name. If none of those are found, the function returns "Not Available", which is pretty strong indication that the generic name was not recovered. Thus, it gets the needed information while pretty effectively dealing with the varied types of HTML structures. ### Scraping Drug Information from a Web Page ``` def scrape_drug_info(url): """ Scrapes detailed drug information from the specified URL. Args: url (str): The URL of the drug page to scrape. Returns: dict: A dictionary containing the scraped information, including: - 'product_link' (str): The URL of the drug. - 'drug_name' (str): The name of the drug. Returns "Not Available" if not found. - 'generic_name' (str): The generic name of the drug. Returns "Not Available" if not found. - 'uses' (str): The uses of the drug. Returns "Not Available" if not found. - 'side_effects' (str): The side effects of the drug. Returns "Not Available" if not found. - 'warnings' (str): The warnings associated with the drug. Returns "Not Available" if not found. - 'precautions' (str): The precautions for using the drug. Returns "Not Available" if not found. - 'interactions' (str): The interactions with other drugs. Returns "Not Available" if not found. - 'overdose' (str): Information regarding overdose. Returns "Not Available" if not found. """ soup = get_soup(url) drug_name = soup.find('h1', class_='drug-name') drug_name = drug_name.get_text(strip=True) if drug_name else "Not Available" generic_name = extract_generic_name(soup) selectors = { 'uses': ['.uses-container.center-content', '.uses-container'], 'side_effects': ['.side-effects-container.center-content', '.sideeffects-container'], 'warnings': ['.warnings-container.center-content', '.warnings-container'], 'precautions': ['.precautions-container.center-content', '.precautions-container'], 'interactions': ['.interactions-container.center-content', '.interactions-container'], 'overdose': ['.overdose-container.center-content', '.overdose-container'] } extracted_info = {key: extract_text_from_div(soup, sel) for key, sel in selectors.items()} return { 'product_link': url, 'drug_name': drug_name, 'generic_name': generic_name, 'uses': extracted_info['uses'], 'side_effects': extracted_info['side_effects'], 'warnings': extracted_info['warnings'], 'precautions': extracted_info['precautions'], 'interactions': extracted_info['interactions'], 'overdose': extracted_info['overdose'] } ``` The scrape\_drug\_info function is developed to get detailed information of drugs from a given URL of the web page. It has one argument: url-the address of the page that contains the drug information for scraping. First, the function calls the get\_soup function to fetch and parse the given URL's HTML content. Then, it tries to find the name of the drug by finding an

    element with class drug-name, returning the text content after stripping any white space; if not found, it defaults to "Not Available." Finally, the function calls extract\_generic\_name to obtain the generic name of the drug. The function creates a selectors dictionary to gather additional information like use, side effects, warnings, precautions, and possible interactions along with its overdoses. This function also contains keys as each different piece of information with lists for respective CSS selectors to figure out the relevant HTML elements available in the page. Applying this dictionary comprehension, each selectors list applies the use of extract\_text\_from\_div with the content of a different key to get an assigned text value for extraction of text content from each list. Finally, the function returns a dictionary containing the extracted information that includes the product link, drug name, generic name, and other miscellaneous data like uses, side effects, warnings, precautions, interactions, and overdose data. It ensures that "Not Available" is returned for the unavailable information, thus showing that when specific details could not be found, their actual lack is clearly highlighted. Such a comprehensive design thus efficiently extracts data in clarity with robustness as opposed to handling different kinds of layouts on pages. ### Scraping Drug Data from the Database ``` def scrape_drug_data_from_database(): """ Scrapes drug information from URLs stored in the database and saves the results. This function retrieves URLs from the 'product_links' table with a status of 0 (indicating that they have not been scraped yet). For each URL, it calls the `scrape_drug_info` function to obtain drug information. If the drug name is not available, the URL is recorded in the 'failed_urls' table. Otherwise, the scraped data is inserted into the 'final_product_data' table. After successfully scraping a URL, its status is updated to 1 to indicate it has been processed. The function handles exceptions by logging errors and saving failed URLs to the 'failed_urls' table. It also introduces a random sleep interval between requests to avoid overwhelming the server. Returns: None """ cursor.execute('SELECT product_link FROM product_links WHERE status = 0') urls = cursor.fetchall() for url, in urls: try: drug_data = scrape_drug_info(url) if drug_data['drug_name'] == "Not Available": cursor.execute('INSERT INTO failed_urls (url) VALUES (?)', (url,)) else: cursor.execute(''' INSERT INTO final_product_data (product_link, drug_name, generic_name, uses, side_effects, warnings, precautions, interactions, overdose) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) ''', ( drug_data['product_link'], drug_data['drug_name'], drug_data['generic_name'], drug_data['uses'], drug_data['side_effects'], drug_data['warnings'], drug_data['precautions'], drug_data['interactions'], drug_data['overdose'] )) # Update status to 1 (scraped) cursor.execute('UPDATE product_links SET status = 1 WHERE product_link = ?', (url,)) conn.commit() except Exception as e: print(f"Error processing URL {url}: {e}") cursor.execute('INSERT INTO failed_urls (url) VALUES (?)', (url,)) conn.commit() time.sleep(random.uniform(SLEEP_MIN, SLEEP_MAX)) # Load user agents and start scraping user_agents = load_user_agents(USER_AGENTS_FILE) scrape_drug_data_from_database() # Close database connection conn.close() ``` This function retrieves drug information from URLs in a database. It uses the product\_links table in the database where the status of scraping is marked as 0, meaning the links have not been scraped yet. It will run a SQL query to fetch the relevant URLs first and then iterate over those fetched URLs. For each URL, it tries to scrape information on drugs by using the scrape\_drug\_info function. This script inserts scraped data into final\_product\_data table in case if the drug name was identified; otherwise, the URL of non availability of the drug name gets logged in failed\_urls to reiterate through all other URLs again. This way after processing the URL function goes for update of record in the table called product\_links that was in "0" status for its url changes "1." It also tracks all previously already passed so the same job will not occur with the same URL twice. Another essential feature of the function: error handling. It includes exception control, which sometimes could be faced during web scraping. If an error emerges in a URL, it produces the error message and adds a troublesome url to the failed\_urls table. Thus, the requests did not fall out for the server; it provides including a random sleep between each requests applying time.sleep method with time to sleep randomly defined in constant values SLEEP\_MIN and SLEEP\_MAX. After successful scrapping, it returns conn.close() which closes down all the connections of the database and the resources used are free. It therefore only maintains a fair balance between being efficient and handling errors because this is the very valuable information related to drugs that is gathered gradually. ## Libraries and Versions This code utilizes several key libraries to perform web scraping and data processing. The versions of the libraries used in this project are as follows: BeautifulSoup4 (v4.12.3) for parsing HTML content, Requests (v2.32.3) for making HTTP requests. These versions ensure smooth integration and functionality throughout the scraping workflow. ## Conclusion The intention of this WebMD web scraping is to show an organized way of extracting data from webmd while maintaining ethical scraping principles. This project incorporates structured workflows, intelligent request handling, and database integration to facilitate effective and reliable data collection. The use of SQLite as a storage mechanism allows for easy retrieval and analysis, as well as mechanisms for logging failed URLs that can be re-scraped to maximize data completeness. This project is useful in automating the retrieval of medical data and offering insights into how the data can be further analyzed or integrated into health care applications. AUTHOR I’m Ambily, Data Analyst at Datahut. I specialize in building automated data pipelines that convert scattered medical and pharmaceutical information into structured, usable formats—empowering healthcare researchers, analysts, and data-driven platforms. At Datahut, we’ve helped businesses across industries harness the power of web scraping to streamline research, monitor updates, and make informed decisions. In this blog, I’ll explain why web scraping is essential for extracting accurate and timely drug information from WebMD. If you're looking to build a custom scraping solution for healthcare data, feel free to reach out using the chat widget on the right. We’re here to help you unlock the full potential of structured medical insights. ### Hidden Cash Flow Problems in Online Fashion Retail URL: https://www.blog.datahut.co/post/the-hidden-cash-crisis-in-profitable-online-fashion-stores/ Last updated: 2026-07-23T07:48:32.000Z As a founder of Datahut, I've spent the last 14 years working with a lot of [fashion stores, both small and large](https://www.datahut.co/retail-solutions?ref=blog.datahut.co). We;ve had a front row seat as we saw real data powering decisions at successful fashion retailers. One of the problems I've seen with struggling stores is that they’re profitable on paper but their cash is tied to their inventory. Am I talking about you? Lets find out Open your P&L and you might see enviable numbers: - Gross margins comfortably above 55 % - A conversion rate nudging 3 % - Double-digit month-on-month growth Yet your current account feels anaemic. Why? Because as much as 20 %–30 % of the value in every garment evaporates into storage, insurance, capital interest, shrinkage and markdown risk—the classic inventory-carrying cost range for fashion retail. Layer on the industry’s average sell-through rate of just 40 %–80 % and you realise you’re effectively paying rent on your own cash. Every stagnant SKU also hurts perception: stale collections, broken size runs and perpetual clearance banners scream “bargain bin,” not “premium brand.” Bottom line: inventory is not an asset until it moves. Until then it is a cash-eating liability. The antidote is radical visibility—turning messy data into simple, actionable numbers you can trust. ## What is happening in the market [Sell-through rate (STR)](https://www.shopify.com/in/blog/sell-through-rate?ref=blog.datahut.co) – the percentage of inventory sold within a period versus received – is a critical gauge of how well fashion e-commerce brands balance supply with demand. Industry benchmarks consider \~70–80% sell-through as the “ideal” range for healthy performance. In practice, most retailers aim to sell about 70% of a season’s stock at full price before markdowns. Recently, however, achieving that has become tougher: initial sell-through by the start of clearance often hovers closer to \~60% in fashion, as consumers’ purchasing shifts and seasons blur. Top-quartile performers (e.g. successful fast-fashion and lean DTC brands) consistently hit 80%+ sell-through in-season, meaning the vast majority of their styles sell without discounting. Lagging brands or marketplace sellers with poor alignment might see sell-through rates down in the 40–50% range, indicating a lot of slow-moving stock. (Industry experts note that a sell-through below \~40% is problematic, leading to heavy markdowns and write-offs. By the end of the full cycle (including clearance sales), retailers generally target around 90–95% total sell-through so that only a minimal fraction of inventory remains unsold. For instance, luxury and premium brands – which traditionally had very high full-price sell-through – have also felt pressure; many now resort to outlet channels or online flash sales to clear stock and reach close to 90% sell-through by season’s end. Notable case:[ ASOS’s recent “Test & React” initiative](https://asos-12954-s3.s3.eu-west-2.amazonaws.com/files/7217/0065/7934/ASOS%5FAnnual%5FReport%5F2023.pdf?ref=blog.datahut.co#:~:text=through%20ASOS.com%20and%20third,3x%20faster%20than%20average) (a rapid small-batch program) achieved \~60% sell-through within just 7 days of launch for new products, with those items turning 3× faster than its average stock. This underscores how data-driven assortment and nimble restocking can dramatically lift [sell-through performance,](https://www.opensend.com/post/sell-through-rate-statistics-ecommerce?ref=blog.datahut.co) approaching the levels of an Amazon marketplace seller (many marketplace sellers consider \~80% STR as excellent, and may discontinue or liquidate items that perform far below that. ## Working Capital & Inventory Carrying Costs Efficient inventory management is crucial for freeing up cash in fashion e-commerce. Inventory ties up a substantial share of working capital – in fact, in 2022 the inventory held by global fashion companies averaged about 20.7% of annual revenue. This ratio has risen from roughly 18–19% pre-2020 to over one-fifth post-pandemic, reflecting larger stock buffers and excess goods in recent years. In dollar terms, this means a fashion e-tailer with $1 billion revenue might have over $200 million sitting in inventory on its balance sheet. Such capital could otherwise be invested in new products, marketing, or technology. The scale of the issue is evident in aggregate: the fashion industry over-produced an estimated 2.5–5 billion garments in 2023, resulting in $70–$140 billion worth of unsold stock globally– a massive amount of cash effectively locked in inventory. Even top brands were not immune: luxury leaders [LVMH and Kering together held about €5 billion ($5.4B)](https://www.businessoffashion.com/articles/retail/the-state-of-fashion-2025-report-inventory-excess-stock-supply-chain/?ref=blog.datahut.co) in excess fashion inventory in 2024, and Nike saw the portion of its products needing discounting jump to 44% in 2024 (from just 19% in 2022) due to inventory overhang. These examples illustrate how excess stock directly impacts working capital and profitability. ## Inventory carrying cost [Inventory carrying cost](https://www.opensend.com/post/inventory-carrying-cost-statistics-ecommerce?ref=blog.datahut.co) is another key metric: it encompasses storage, insurance, depreciation, and the cost of capital for held inventory. Retail benchmarks show that [holding inventory typically costs about 20–30% of the inventory’s value per year](https://www.netsuite.com/portal/resource/articles/inventory-management/inventory-carrying-costs.shtml?ref=blog.datahut.co). In other words, a company with $100 million of inventory may incur $20–30 million in annual carrying costs (warehouse space, financing, obsolescence risk, etc.). This is why improving turnover and sell-through has outsized financial benefits – every week reduction in average stock levels saves money and releases cash. High-performing e-commerce fashion players leverage techniques like demand forecasting AI, just-in-time replenishment, and drop-shipping to minimize on-hand stock (thus reducing carrying cost). For instance, Shein’s \~40-day inventory model not only boosts sell-through but also sharply cuts storage time and markdown expense Similarly, many marketplace-based sellers (who often use fulfillment services like FBA) are incentivized to keep stock lean, since overstocking can raise holding costs by 25%+ and even incur extra fees The most recent data and reports from firms like McKinsey, BCG, and BoF indicate that improving inventory metrics remains a critical priority for fashion e-commerce. Those that succeed in boosting turnover and sell-through (through better demand analytics, agile supply chains, and tighter buying) will not only free up cash and cut costs but also outperform peers in responsiveness and profitability ## A Five-Step Rescue Plan You Can Run in Google Sheets ### Step 1 – Export the Basics 1. Pull 12–24 months of data from Shopify, WooCommerce or your ERP:Orders (with dates & quantities)Stock receipts (PO dates & units)Returns (date, SKU, qty) 2. Standardize your CSV:Ensure consistent SKU formatting (no stray spaces or dashes)Normalize date formats (YYYY-MM-DD)Filter out cancelled or test orders 3. Set up your sheet tabs:Raw DataCleaned Data (with formulas applied)Metrics Dashboard Pro Tip: Use Google Sheets’ “Import” functions (e.g. IMPORTDATA) or a simple n8n/Make scenario to auto-refresh this CSV each month. Grab 12–24 months of order, stock and return data from Shopify, WooCommerce or your ERP. Save as CSV. Minimum columns are given below SKU | Style | Colour | Size | Qty Received | Qty Sold | Date Received | Current Stock | Returns ### Step 2 – Add Three Killer Metrics ![three killer metrics](https://www.blog.datahut.co/content/images/2026/07/img-234.png.webp) - Automate your formulas so they update when you paste new data. - Color-code high/low performers with conditional formatting. - Add dynamic date filters to compare rolling 4-week vs. 12-month trends. ### Step 3 – Find Heroes & Laggards - Sort by ST % (desc) and WOH (asc) in your cleaned Data tab. - Define your groups:Heroes: Top 20 % (ST % ≥ 80 % & WOH < 8)Watch-List: Middle 60 % (all others)Laggards: Bottom 20 % (ST % ≤ 40 % & WOH > 12) - Create pivot tables to break out heroes/laggards by:StyleColourSizeSeasonality tag - Visualize with a simple bar chart—heroes vs. laggards—so you can instantly see where cash is trapped. Insight: Look for patterns—are certain colours or sizes consistently lagging? That’s often a sizing-curve or styling issue, not just “bad demand.” ### Step 4 – Rebalance the Next PO 1. Heroes first:Increase reorder qty by 10–30% above last cycle’s sell rateConsider testing a small new colourway or size on 1–2 % of total qty 2. Laggards last:Pause fresh POs until you’ve liquidated ≥ 50 % of existing stockBundle or cross-sell via targeted promotions (e.g., “Buy a hoodie, get a laggard tee 50 % off”) 3. Size-curve alignment:Mirror your actual sales distribution vs. vendor’s suggestionUse your pivot “Size” tab to calculate exact ratios 4. Lead-time buffer:Factor in supplier lead time variability—if your average lead time is 21 days ± 5 days, order 10 % extra heroes to cover delays. ### Step 5 – Rinse & Repeat Monthly (Analyse → Rebalance → Test → Analyse again.) 1. Set a calendar reminder on Day 1 of each month:Export fresh dataRe-run your metrics tab 2. Track your progress:Maintain a “Month-Over-Month” tab showing changes in overall ST %, average WOH, cash released 3. Continuous testing:If a new tactic (e.g., bundling) frees up > 5 % extra cash, roll it out to similar SKUs 4. Scale with minimal effort:Once your Sheets workflow is bullet-proof, automate the data pull (via Datahut feeds or an integration tool) so you spend 30 minutes reviewing, not copying & pasting. Bonus: Share your dashboard with key stakeholders (e.g., merchandising, finance) so everyone sees the impact of freed-up cash—and stays aligned on the plan. ## How Web Scraping Supercharges Each Step Here’s a more detailed look at how web scraping supercharges each step—packed with extra boosters and the tangible edges you gain: ![How webscraping supercharges each step](https://www.blog.datahut.co/content/images/2026/07/img-235.png.webp) Extra Tips: - Cross-Market Signals: Scrape international sites to see which styles are trending globally before they hit your region. - Supplier Health Check: Monitor vendor websites or marketplaces (e.g., Alibaba, IndiaMART) for lead-time changes or stock shortages—critical when POs are 60+ days out. - Pricing Psychology: Track competitor price thresholds (e.g. the “.99” effect) and adjust your pricing slice to maximize perceived value. With these data-driven boosters, your simple Sheets workflow becomes a full-blown market radar—delivering competitive edge at every turn. ## Putting ChatGPT on Autopilot (Zero Data Team Needed) Even a solo founder can mimic a BI team by pasting a CSV slice (2 000 rows tops) into ChatGPT with Advanced Data ### Analysis: Prompt: “Load the file and: 1\. Compute sell-through %, WOH, and seasonality tag by SKU | size | colour. 2\. Flag SKUs with WOH > 12 or ST % < 40. 3\. Cluster SKUs into heroes (top 20 %), middle (60 %), laggards (20 %). 4\. Recommend reorder quantities to keep total stock value < $ 1Million.” ChatGPT will: - Generate pivot tables in-memory. - Visualise a heat-map of seasonality spikes. - Output a cash-release forecast—e.g., “Cutting laggard buys by 30 % frees ₹18 lakh over 90 days.” - Produce Python code you can copy into Google Colab for repeat runs. Supercharge with Datahut: Feed ChatGPT a second CSV scraped by Datahut—competitor pricing and stock snapshots. Ask: “Overlay my hero SKUs against competitor stock-out frequency. Which items should I push in ads this month for maximum margin?” Within minutes you have campaign ideas with built-in pricing leverage. ## Turning Laggards Into Revenue—Without Discounting Your Brand Deep discounts erode perception. Instead, treat laggards with: Hero-Anchored Bundles: Pair a best-selling jogger (ASP ₹2 199) with a slow-moving tee (COGS ₹200, ASP ₹999). Frame it: “Free tee worth ₹799 when you grab our bestselling joggers.” AOV stays high, margins intact, tee exits quietly. Gift-with-Purchase (GWP): Offer stale scarves as GWP on orders ≥₹3 000\. Perceived value rises; shelf clears. Tiered Loyalty Rewards: Let VIPs redeem points for laggard SKUs—moving stock while buying goodwill. Flash Bundles via WhatsApp: Send a private drop to VIPs: “24-hour lookbook bundle—30 % off.” Urgency without public markdowns. ### Metrics That Matter Going Forward ![Metrics that matter going forward](https://www.blog.datahut.co/content/images/2026/07/img-236.png.webp) Build a lightweight dashboard in Google Sheets or pipe Datahut Feeds into Looker Studio for efficient tracking. 1. What causes a cash crisis in profitable fashion stores?Even profitable brands can run into cash flow problems when too much capital is tied up in slow-moving or stale inventory. 2. What is the impact of stale SKUs on cash flow?Stale SKUs accumulate carrying costs, reduce sell-through, and often require heavy discounts to clear — all of which drain cash. 3. How can web scraping help identify inventory issues?Web scraping lets brands track real-time pricing, competitor stockouts, and demand trends — helping them avoid overbuying and detect laggards early. 4. What’s a good sell-through benchmark in fashion?A sell-through rate above 70% is considered strong. Anything consistently below 40% over weeks may indicate poor product-market fit or pricing issues. 5. How can I unlock trapped cash in my inventory?Analyze sell-through %, cluster SKUs by performance, reduce reorders of laggards, and use scraping tools to make proactive stocking decisions. Author I’m Tony Paul ,founder of Datahut . I’ve been working in the web scraping industry for over 14 years. Working with top fashion brands and their data teams - Me and my team at Datahut learned a lot of invaluable lessons on how to use data to free up cash tied to the inventory. If you're someone facing this trouble and want to solve these problems. Get in touch with us using the chat widget on the right hand side. ### Automate Trulia Real Estate Scraping Using Python URL: https://www.blog.datahut.co/post/how-to-automate-trulia-real-estate-data-scraping-with-python/ Last updated: 2026-07-23T07:48:32.000Z ## Introduction Did you know that most home buyers start by searching online? With Trulia, users can browse property listings, compare prices, and analyze neighborhood trends before making a decision. However, manually tracking real estate data across multiple listings is time-consuming and inefficient. This blog will be a tutorial for you on how to automatically scrape Trulia's real estate listings. You'll learn how, using Python and scraping techniques, you can grab valuable property information such as price, location, amenities, and mortgage rates right from Trulia's listings. Whether you're a property investor, data analyst, or market researcher, this guide will help you collect structured data at scale, leading to better decision-making and deeper insights into the markets. So let's dive into how this two-step automated web scraping process works! ## What is Web Scraping ? It is the process by which websites are automatically inquired through the use of programming techniques for extracting some form of data. Web scraping retrieves information quickly using a request that is sent to web pages and retrieves content as well as draws out details to be reused again. Here, in Trulia, structured data on real estate, such as price, location, amenities, and mortgage rates, is aggregated without manual effort using web scraping. It turns out to be an essential tool for investors, analysts, and researchers requiring current market information in terms of new trends. ## Brief Overview of the Trulia Web Scraping Process The most crucial role of web scraping plays is in the scraping of data from Trulia, which helps to extract property-related data directly from the website's network APIs and HTML content.This project performs a two-step automated web scraping process to collect detailed information about homes listed on the Trulia website in San Francisco. The first code scrapes unique product links by dynamically generating URLs from a base API endpoint and transaction IDs, then sending POST requests and recursively parsing JSON responses to identify and store links in an SQLite database. The next code takes all those links and scrapes detailed property information by sending GET requests with a randomized user agent and with proxy support to hide IP. Then it parses every property page through BeautifulSoup in terms of fetching home name, location, price, mortgage details, specifications, description, highlights, amenities, and tax data. Extracted data is also stored in a structured SQLite database and logs failed requests for review. This highly robust, end-to-end solution allows comprehensive data gathering for further analysis. ## Libraries Behind the Scenes in Trulia Web Scraping A set of Python libraries serves as the backbone of this project, ensuring smoother execution and robust performance. Let's explore the libraries used, their purposes, and how they contribute to this project. ### Requests Requests is the core of the web communication process in this project. It makes handling HTTP requests, such as GET and POST, a seamless affair with Trulia's servers. In the first phase of the project, requests is used to send POST requests to Trulia's GraphQL API, fetching JSON data that contains product links. In the second phase, the library retrieves HTML content of pages for each property. The ease and reliability make it indispensable for web interactions. ### SQLite3 The sqlite3 library provides the project with a powerful yet lightweight database solution. It ensures data is properly organized and persisted throughout the two phases of scraping. In the first phase, it is used in database creation and management that keeps unique product links along with their scraping status. This avoids duplicates as sqlite3 optimizes the scraping process. In the second phase, it captures detailed property information, home names, locations, prices, and logs failed attempts and enables retries for incomplete or unsuccessful scraping operations. ### BeautifulSoup The BeautifulSoup library, taken from bs4, will be a core component of this project in the second stage. It will take in HTML content scraped off property pages, process raw HTML to create a navigable structure such that specific information about names, prices, location, and amenities can be gathered without error. This would rely on the parsing power of BeautifulSoup in processing unstructured web data to structured and usable data. ### Random and Time For natural-looking and bot-detect-resistant scraping, random and time libraries are used. Each request uses random strings for user-agents by reading from a text file. The user-agent string will simulate many different kinds of browsers and versions as they send the requests. The random library also makes it vary the amount of delay between requests that can be used with time library. This intentional slowdown in scraping will keep the project away from being noticed and throttled by web servers. ### urllib3 The urllib3 library extends the HTTP capabilities of the project. Specifically, it handles secure HTTPS connections and disables SSL warnings when making connections over proxy servers. This capability is critical to maintaining uninterrupted communication with Trulia's servers, particularly when using a proxy to anonymize requests. ## Understanding Proxies and Their Role in Web Scraping A proxy is an intermediate server between the web scraping script and the target server . The target server never sees the original IP address of the client when the request is relayed through a proxy. It sees the proxy's IP. Masking the original IP address is essential in web scraping to prevent IP blocking and bypass geographical restrictions. Proxies distribute requests across multiple IPs to avoid detection and allow access to region-specific content, ensuring seamless and unrestricted data collection. This project uses proxies mainly because of IP blocking. Indeed, most of the sites would use some sort of an anti-scraping tactic such as rate limiting or even IP blocking once they determine queries coming from the same IP address are going too frequently or repetitively. A scraping script can spread its traffic across different IP addresses by sending requests through a proxy, or even through a pool of proxies. It also reduces the risk of getting caught and flagged or blocked by a scraper. Other than preventing IP blocking, proxies are used to bypass geographical restrictions. A website might limit its content only to the location of a user. The proxies from other regions make the scraping script appear to be coming from a variety of places, so it is very helpful in all-around data gathering and evades such restrictions. Proxies also contribute to anonymity as they mask the actual identity and location of the client. This is really helpful in concealing from tracking systems that flag suspicious activity for attention. This, however is not the end; with proxies scraping is now more reliable and scalable. If one proxy gets flagged or blocked, the scraper can switch to another proxy to make sure that data extraction isn't interrupted. Proxies are smoothly integrated in the request mechanism of this project. The proxy settings define the urllib3 library. This will make all the requests in an HTTPS secure connection go to a proxy server, which should provide anonymity, accessibility along with traffic distribution that creates quite a robust process of web scraping. Proxies, with no doubt, are a bit of a luxury and a necessity in any web scraping project, especially when dealing with serious measures put in place against scraping.Users can either choose Datahut's proxy services or opt for any other free or paid proxy services based on their preferences and requirements. ## STEP 1 : Product URLs Scraping ### Importing Required Libraries ``` import requests import sqlite3 ``` The code imports the requests library for sending HTTP requests to retrieve web content and the sqlite3 library for interacting with an SQLite database to store and manage scraped data. ### SQLite Database Configuration ``` # SQLite database configuration DB_NAME = "Trulia_Webscraping.db" TABLE_NAME = "product_links" ``` This section defines the configuration for the SQLite database used in the project. The DB\_NAME variable specifies the name of the database file which serves as the storage location for the scraped data. The TABLE\_NAME variable sets the name of the table within the database. This table is used to store the links of the products . By defining these constants, the script ensures consistency and simplifies database operations such as table creation, data insertion, and querying. ### URL and Transaction IDs Configuration ``` # Constant part of the URL BASE_URL="https://www.trulia.com/graphql?operation_name=WEB_searchResultsMapQuery&transactionId=" # List of transaction IDs TRANSACTION_IDS = [ "30a21e3f-5b7e-49b6-8deb-01f0e5f8cb07","bf8443bc-991e-4e80-85bb-fd47c1d3e615", "d441c1be-0201-487e-b5cc-79a93795d9c4","5b006fc1-0a85-49ef-a3ae-985ec3c320ae", "1f4b143a-71f2-45d7-b7ec-dd78dcc69cd1","82d9baa6-ca36-444b-b97f-ac6fa65ee79f", "794ae677-5a6d-4e44-8dc0-7146a6ef574e","fd7eb450-99f2-49c4-8bbf-7ca2180d5214", "fddcc9ce-cd47-4b9b-9d70-1cc1a1886a6b","fd7e7e4a-cee9-4205-b685-01f66b593c60", "32db8852-62df-49c3-942c-46c1526b049c","ac86cff6-2d9a-4502-9588-a179ec9a5fbd", "8033dce0-b699-47d4-abcd-8be1b85741c3","cd904320-f180-43b2-95b7-6cdc864b9ab1", "0ff9c18e-6e09-4658-838c-571fcf9a5971","bb34872c-2a95-4806-8c91-c18df7228497", "eae572f1-d55e-440f-88ca-b3f97a0e1732","7f325934-9ca3-4743-b941-4c003b25abbe", "fead1864-4f52-4008-8f10-3de9e1220567","ea8693cf-8b20-4fca-87d4-664372db7b11", "f48b7f88-a9df-4399-af9e-31bbedbf85ce","7856e515-3c00-4145-9f86-8554c53370cc", "e92d4d95-f803-4787-a26f-2475b134d0fb","3ea3f7df-4c91-4ad3-be01-ac5356de0454", "d6aaaf47-b5fb-4c87-b447-ebbcea97f349" ] ``` It forms a list of requests toward Trulia using the base URLs. BASE\_URL refers to the constant part of each network API url, and TRANSACTION\_IDS are unique IDs that are combined with the base URL to give each request its full URL form. By combining the BASE\_URL with these transaction IDs, the program can create a list of URLs to get data from the website. This method manages many requests well while keeping the code clear and easy to expand. ### Setting Up Proxy Configuration ``` PROXIES = { "http": "http://datahutapi:password@proxy-server.datahutapi.com:8001", "https": "http://datahutapi:password@proxy-server.datahutapi.com:8001" } ``` Here defines a dictionary called PROXIES. This dictionary contains configurations with regard to connecting to the Internet using a proxy server. The dictionary has two keys: "http" and "https". They are HTTP and HTTPS protocols respectively. For both of them, the value will be a URL that shows an address of the proxy server plus username and password with regards to authentication. The proxy server is set at proxy-server.datahutapi.com and running on port 8001\. Now, any kind of request made using such protocols pass through the particular proxy server set and which can then hide the client's IP address or also bypass various restrictions. ### Defining HTTP Headers ``` HEADERS = { 'accept': '*/*', 'accept-language': 'en-IN,en-GB;q=0.9,en-US;q=0.8,en;q=0.7,ml;q=0.6', 'cache-control': 'no-cache', 'content-type': 'application/json', 'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36', } ``` The HEADERS dictionary in the code defines HTTP headers being sent with a web request. These are very important while performing web scraping as they make requests look as if they are coming from a real web browser. Hence the server treats this request as valid and not coming from a bot or some automated tool. Headers like User-Agent help the scraping script determine it to be some special browser that helps the server not to reject the request considering it suspicious or unknown. Accept and Accept-Language headers also specify kinds of content a response ought to have and the favored language, so the scraper will be more adaptable to the different structure of the website and localization. This code does create headers like a regular browser request without cache enforced on them . The content type is set to JSON, which is very commonly used in API-based data fetching. These headers help the scraper not get detected, smooth extraction of data, and return content in a structured and readable format like JSON. ### Function to Create SQLite Database and Table ``` def create_database(db_name, table_name): """ Create an SQLite database and a table if it doesn't exist. This function initializes a database connection and creates a table with the specified name. The table includes the following columns: - `product_link`: A unique identifier for each product (primary key). - `status`: An integer field to track the scraping status of the link (default value is 0, indicating not scraped). Parameters: db_name (str): The name of the SQLite database file. table_name (str): The name of the table to be created. Returns: None """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(f""" CREATE TABLE IF NOT EXISTS {table_name} ( product_link TEXT PRIMARY KEY, status INTEGER DEFAULT 0 ) """) conn.commit() conn.close() ``` This function, create\_database, is meant to initialize an SQLite database and create a table if it does not already exist. It first establishes a connection to the specified SQLite database and then defines a table with the given name. The table contains two columns. The first column, which is product\_link, uniquely determines each product and therefore takes the role of the primary key in the table. The status of each product link can now be tracked using an integer field whose default value is 0-the link has not yet been scraped. The function does that by ensuring that the table creation only happens if the table does not already exist hence preventing redundant table creations. After executing the needed SQL commands, the database connection is committed and closed. It is often used in a web scraping project to maintain and track the status of all scraped product links efficiently. ### Function to Send POST Request with Transaction ID ``` def send_request(transaction_id): """ Send a POST request to the specified URL with transaction ID. This function sends a POST request to the Trulia website's GraphQL endpoint with the provided transaction ID . It handles the request using the specified headers, proxies, and disables SSL verification. If an error occurs during the request, it logs the error and returns None. Parameters: transaction_id (str): The unique identifier used in the URL to fetch specific data from the server. Returns: response (Response or None): The response object if the request is successful, or None if an error occurs during the request. """ url = BASE_URL + transaction_id try: response = requests.post( url, headers=HEADERS, proxies=PROXIES, verify=False ) return response except Exception as e: print(f"Error sending request for \ transaction ID {transaction_id}: {e}") return None ``` This function will send a POST request in send\_request(). It gets an ID for the specified URL, sending a possible data request it may call for on the server and uses it in building full URLs with an append, then takes all the appended transactions, constructing its very content. It sends the request using the requests.post method with all the pre-set headers (HEADERS) and proxied setting up (PROXIES), and at the same time, utilize verify=False for disabling the SSL verification; this ensures the request moves on even if the SSL certificates can't be verified. And then the function will return the response object in case the request was a success. This is highly useful in the context of web scraping or APIs, where dynamic fetching of information using such identifiers as transaction IDs has to be done. But when there was some error, an exception is caught by this function, logs an error with information about transaction\_id, and returns None. ### Extracting Product Links from Trulia's API Response ``` def extract_links_from_json(data): """ Extract unique product links recursively from JSON data. This function recursively traverses the given JSON data, which can be in the form of nested dictionaries or lists, to find all product links that start with "/home/". These links are then formatted with the base URL and added to a set to ensure uniqueness. Parameters: data (dict or list): The JSON data structure to be searched, which could contain nested dictionaries or lists. Returns: set: A set of unique product links, each starting with "/home/". """ links = set() def recursive_search(obj): if isinstance(obj, dict): for value in obj.values(): recursive_search(value) elif isinstance(obj, list): for item in obj: recursive_search(item) elif isinstance(obj, str) and obj.startswith("/home/"): links.add("https://www.trulia.com" + obj) recursive_search(data) return links ``` The function extract\_links\_from\_json is designed to scan complex JSON data that could possibly contain nested dictionaries or lists and retrieve unique product links starting with the path "/home/'. These are links to individual product pages, so the base URL https://www.trulia.com is added to any link found. There's a useful function named recursive\_search that performs a search on every node of the JSON structure. It runs over all the values in the dictionary. If it identifies a list, it runs over that list and checks each item on that list. It then recognizes that a string starts with "/home/" and attaches the base URL to that link, stores it in a set so that it will not be counted as a duplicate, and continues this pattern for each item. During its process, the function at its very end returns all uniquely formed and fully formatted links found to products by being picked up from JSON. ### Saving Extracted Product Links to a Database ``` def save_links_to_database(db_name, table_name, links): """ Inserts unique product links into the SQLite database table. This function connects to the SQLite database specified by the `db_name` parameter, and inserts product links into the table specified by `table_name`. If a link already exists in the table, it will be ignored to ensure that only unique links are stored. Args: db_name (str): The name of the SQLite database to connect to. table_name (str): The name of the table where the links will be inserted. links (iterable): A collection of unique product links to be saved into the database. Raises: sqlite3.Error: If there is an error during the database operation (e.g., failure to insert a link into the database). """ conn = sqlite3.connect(db_name) cursor = conn.cursor() for link in links: try: cursor.execute(f"INSERT OR IGNORE INTO {table_name} " "(product_link) VALUES (?)", (link,)) except sqlite3.Error as e: print(f"Database error while saving link " f"{link}: {e}") conn.commit() conn.close() ``` The function save\_links\_to\_database saves all the links for the products it finds in a database. The database used here is SQLite. Its name is determined by the parameter db\_name and the table name given as table\_name. This function accepts a list of product links as an argument and tries to add each link to the database. The function utilizes the SQL command INSERT OR IGNORE, which ignores any links that are already in the table. This ensures no data is copied. A try block is used for each link so that any errors during the database operation will be caught. If it encounters some sort of error, then the error will be posted; however, all the rest links work properly. Subsequent to that, if there are links remaining, all links' changes will be stored within the database and connections will close up after guaranteeing the fact that the safe link of the product has moved toward the database. ### Processing Transaction IDs to Extract Product Links ``` def process_transaction_ids(transaction_ids): """ Process a list of transaction IDs, send requests for each, and extract unique product links. This function iterates over a list of transaction IDs, sends a request for each, and processes the response to extract product links. These links are collected into a set of unique links, which is returned at the end. Args: transaction_ids (list): A list of transaction IDs to be processed. Returns: set: A set containing unique product links extracted from the responses. """ unique_links = set() for transaction_id in transaction_ids: response = send_request(transaction_id) if response and response.status_code == 200: data = response.json() links = extract_links_from_json(data) unique_links.update(links) print(f"Extracted links from transaction ID {transaction_id}") else: print(f"Request failed for transaction ID {transaction_id} " f"with status code {response.status_code if response else 'N/A'}") return unique_links ``` The function process\_transaction\_ids accepts a list of transaction IDs, sends a request for each one, and returns a set of unique product links from the responses. For each transaction ID in the list, the function calls the send\_request function. If the response is successful with a status code of 200, it takes the response data, converts it into JSON format, and passes the result to extract\_links\_from\_json, which then parses out the links to the products from the JSON and puts them in a set called unique\_links. Since sets are ordered, the duplicates will be automatically eliminated. If the request is not successful for any given transaction ID, the function prints an error message associated with the status code it received. Finally, it returns the set of all unique product links collected across all the transaction IDs for a good, efficient product link gathering from multiple API responses while handling possible errors in requests gracefully. ### Main Execution Flow for Extracting and Saving Product Links ``` # Main Execution if __name__ == "__main__": """ Main execution flow for processing transaction IDs, extracting product links, and saving them to an SQLite database. This script performs the following steps: 1. Creates a database and table for storing product links. 2. Processes a list of transaction IDs to extract unique product links. 3. Saves the unique product links into the database. The script prints a confirmation message when all links have been saved successfully. """ # Step 1: Create database and table create_database(DB_NAME, TABLE_NAME) # Step 2: Process transaction IDs and extract unique links unique_links = process_transaction_ids(TRANSACTION_IDS) # Step 3: Save unique links to the database save_links_to_database(DB_NAME, TABLE_NAME, unique_links) print(f"All unique product links have been saved to the database {DB_NAME}" f"in table {TABLE_NAME}") ``` This is the body of execution part.In that section of the script that has to do with fetching links from Trulia's API and storing them in a database. First, declare the database and a table in which the links are going to be stored using the function create\_database. Then, it processes each item in a list of predetermined transaction IDs by passing this item to the process\_transaction\_ids function, which goes through unique product links contained within each API response to the specific transaction ID and finally saves the extracted links in the database using the function save\_links\_to\_database in a way that will exclude saving any duplicate entries. Last it gives a confirmation message after print, stating that all links are saved into a given database and table. This flow merges operations on the database, interaction with API and data processing so that the entire process becomes efficient and reusable. ## STEP 2 : Product Data Scraping From Product Links ### Importing Necessary Libraries ``` import sqlite3 import requests from bs4 import BeautifulSoup import random import time import urllib3 urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning) ``` Importing necessary libraries. sqlite3 for database interaction, requests for making HTTP requests to fetch webpage content, and BeautifulSoup from the bs4 library to parse and extract data from the HTML of each page. Additionally, the random and time libraries are used to introduce random delays between requests. The script also disables SSL warnings using urllib3 to handle scenarios where HTTPS certificates might not be fully trusted. ### Setting Up Proxy Configuration ``` PROXIES = { "http": "http://datahutapi:password@proxy-server.datahutapi.com:8001", "https": "http://datahutapi:password@proxy-server.datahutapi.com:8001" } ``` Here defines a dictionary called PROXIES. This dictionary contains configurations with regard to connecting to the Internet using a proxy server. The dictionary has two keys: "http" and "https". They are HTTP and HTTPS protocols respectively. For both of them, the value will be a URL that shows an address of the proxy server plus username and password with regards to authentication. The proxy server is set at proxy-server.datahutapi.com and running on port 8001\. Now, any kind of request made using such protocols pass through the particular proxy server set and which can then hide the client's IP address or also bypass various restrictions. ### Setting Up the Database for Scraping ``` DATABASE = "Trulia_Webscraping.db" ``` The variable DATABASE is used to define the name of the SQLite database file, which in this case is "Trulia\_Webscraping.db". This database serves as the central storage for managing data throughout the scraping process. It is used to store information such as the list of product links extracted from the Trulia API and any detailed data scraped from individual product pages. By using a database, the script ensures that data is organized, persistent, and easily retrievable for further analysis, even if the scraping process is interrupted or restarted. SQLite, being lightweight and easy to use, is a suitable choice for this task, allowing efficient handling of structured data without requiring additional setup. ### Loading User Agents for Scraping ``` # Load user agents from file def load_user_agents(file_path): """ Load a list of user agent strings from a file. Args: file_path (str): Path to the file containing user agent strings. Each line in the file should represent one user agent. Returns: list: A list of user agent strings, with leading and trailing whitespace removed from each line. """ with open(file_path, "r") as f: return [line.strip() for line in f.readlines()] ``` ``` USER_AGENTS = load_user_agents("data/user_agents.txt") ``` A load\_user\_agents function loads in a list of user agent strings from a file. User agents are text strings that one of the web browsers or tool sends to the website as identity. Using different user agents when scraping, the script is actually mimicking requests coming from various browsers or devices so that the whole process seems a bit more anonymous and may not be easily blocked by the website. This function reads a file that's specified by the file\_path argument. In this file, every line contains one user agent string. It removes any extra spaces or newline characters from each line and returns them as a list. These user agents are applied at the time of scraping to randomly choose different identifiers for each request with the simulation of natural browsing habits, which enhances the reliability of a data scraping process. ### Creating Database Tables for Scraping ``` # Function to create tables if they don't exist def create_tables(): """ Create necessary tables in the SQLite database if they do not already exist. Tables: 1. failed_urls: - id (INTEGER): Primary key, auto-incremented. - url (TEXT): The URL that failed to process. - reason (TEXT): Reason for the failure (optional). - timestamp (TIMESTAMP): Timestamp of the failure (defaults to the current time). 2. product_data: - id (INTEGER): Primary key, auto-incremented. - product_link (TEXT): URL of the product. - home_name (TEXT): Name of the home or product. - location (TEXT): Location details of the product. - price (TEXT): Price of the product. - mortgage (TEXT): Mortgage information (if available). - specification (TEXT): Specifications of the product. - description (TEXT): Product description. - highlights (TEXT): Key highlights of the product. - amenities (TEXT): Amenities included with the product. - tax (TEXT): Tax-related information. - timestamp (TIMESTAMP): Timestamp of when the entry was created (defaults to the current time). The function establishes a connection to the SQLite database specified by the `DATABASE` constant,creates the tables if they do not exist, commits the changes, and then closes the connection. Returns: None """ conn = sqlite3.connect(DATABASE) cursor = conn.cursor() # Create the failed_urls table cursor.execute(""" CREATE TABLE IF NOT EXISTS failed_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT NOT NULL, reason TEXT, timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP ); """) # Create the product_data table cursor.execute(""" CREATE TABLE IF NOT EXISTS product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_link TEXT NOT NULL, home_name TEXT, location TEXT, price TEXT, mortgage TEXT, specification TEXT, description TEXT, highlights TEXT, amenities TEXT, tax TEXT, timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP ); """) conn.commit() conn.close() ``` The create\_tables function is there to make sure the SQLite database has all the tables necessary for the actual process of scraping. It connects to the database pointed to by the constant DATABASE and creates two tables: failed\_urls and product\_data. The failed\_urls table is there to keep track of URLs that failed in the scraping process. This includes the URL itself and an optional reason for failure together with the timestamp of failure. This will track problems that might have occurred while scraping. It will permit retrying or debugging later. Product details are placed in the table product\_data; it will be used to store information gathered from the product page like the name of the product, location, price, mortgage details, specifications, description, highlights, amenities, and tax information. There is a timestamp showing when the data was added to the table. The function returns nothing if the tables are already in the database and does not overwrite data. Finally, it saves changes and closes the connection to the database in order to make sure everything is set up before actually scraping. ### Fetching the Next URL to Scrape ``` # Fetch the next URL to scrape def get_next_url(): """ Fetch the next URL to scrape from the database. The function retrieves a single URL from the `product_links` table where the `status` is 0, indicating that the URL has not yet been scraped. It returns the first available URL or `None` if no such URL exists. Returns: str or None: The next product URL to scrape, or `None` if no URLs are pending. Notes: - The function assumes the existence of a `product_links` table with the following structure: - product_link (TEXT): The URL of the product. - status (INTEGER): A flag indicating the scrape status (0 for pending, 1 for completed). """ # Connect to the database conn = sqlite3.connect( DATABASE ) # Create a cursor object to execute SQL queries cursor = conn.cursor() # Execute the SQL SELECT query to fetch the next URL # with a pending status (0) cursor.execute( """ SELECT product_link FROM product_links WHERE status = 0 LIMIT 1 """ ) # Fetch the result of the query result = cursor.fetchone() # Close the database connection conn.close() # Return the URL if a result is found, else return None return result[0] if result else None ``` This get\_next\_url function retrieves the next product URL to scrape from the database. This works by connecting to the SQLite database and querying the product\_links table, which contains the list of URLs to scrape. The function searches for a URL where the status is set to 0, meaning that it hasn't been scraped yet. The LIMIT 1 in the query will limit the result to the first available URL. If there's a URL, the function will return it; if there's no URL with the pending status, then it will return None. This function assists in tracking which URLs are supposed to be processed so that the script scrapes new links in an orderly fashion. Finally, it closes the database connections opened to avoid the formation of open connections when once a query is executed fetching results and returns a returned result containing the URL that goes in further processing. ### Updating the Status of a URL ``` # Update the status of a URL def update_url_status(url, status): """ Update the status of a URL in the database. This function sets the `status` of a given URL in the `product_links` table.The status indicates whether the URL has been processed or not. Args: url (str): The URL whose status needs to be updated. status (int): The new status to set. Common values might include: - 0: Pending - 1: Completed - Other values as defined by the application. Returns: None Notes: - The function assumes the existence of a `product_links` table with the following structure: - product_link (TEXT): The URL of the product. - status (INTEGER): A flag indicating the scrape status. """ # Connect to the database conn = sqlite3.connect( DATABASE ) # Create a cursor object to execute SQL queries cursor = conn.cursor() # Define the SQL UPDATE query to modify the status of the URL cursor.execute( """ UPDATE product_links SET status = ? WHERE product_link = ? """, ( status, url ) ) # Commit the changes to the database conn.commit() # Close the database connection conn.close() ``` The update\_url\_status function is used to update the status of a URL in the database after it has been processed. It takes two inputs: the url that needs to be updated and the new status to be set. The status helps track whether the URL has been scraped or not, with common values being 0 for pending and 1 for completed. This function will establish the connection with the SQLite database. It creates a cursor then it sends the following UPDATE query in relation to the product\_links table of new status to that given URL. Once this update operation has been done, its results are committed into the database which will ensure those modifications permanent and closes the database. This function makes it possible for the system to keep track of which URLs were processed, thereby preventing rescraping of the same URLs. ### Saving Scraped Data to the Database ``` # Save scraped data to the database def save_data(product_link, data): """ Save scraped product data to the database. This function inserts the scraped product data into the `product_data` table in the database.The table should include fields such as product link, home name, location, price, and other details. Args: product_link (str): The URL of the product being saved. data (tuple): A tuple containing the product details in the following order: - home_name (str): Name of the product or home. - location (str): Location details. - price (str): Price of the product. - mortgage (str): Mortgage information (if available). - specification (str): Specifications of the product. - description (str): Description of the product. - highlights (str): Key highlights of the product. - amenities (str): Amenities included with the product. - tax (str): Tax-related information. Returns: None Notes: - The function assumes the existence of a `product_data` table with the following structure: - product_link (TEXT): The URL of the product. - home_name (TEXT): Name of the home or product. - location (TEXT): Location details. - price (TEXT): Price of the product. - mortgage (TEXT): Mortgage information. - specification (TEXT): Specifications of the product. - description (TEXT): Product description. - highlights (TEXT): Key highlights of the product. - amenities (TEXT): Amenities included with the product. - tax (TEXT): Tax-related information. """ # Connect to the database conn = sqlite3.connect( DATABASE ) # Create a cursor object to execute SQL queries cursor = conn.cursor() # Define the SQL INSERT query for the product_data table cursor.execute( """ INSERT INTO product_data ( product_link, home_name, location, price, mortgage, specification, description, highlights, amenities, tax ) VALUES ( ?, ?, ?, ?, ?, ?, ?, ?, ?, ? ) """, ( product_link, *data ) ) # Commit the changes to the database conn.commit() # Close the database connection conn.close() ``` The save\_data function saves the product data in details, scraped from a certain URL, to the database. It takes two parameters: the product link, which is the URL of the product being scraped, and data, which is a tuple containing a variety of information regarding the product including its name, location, price, mortgage information, specifications, description, highlights, amenities, and tax details. This function links the SQLite database, allowing a cursor to execute SQL queries. It prepares and executes an INSERT INTO SQL query, placing the data scraped into the product\_data table of the database. The table has been structured to store all the necessary information for products in columns: product\_link, home\_name, location, price, and many more. After entering all the data, the changes that have been made are commited to the database so as to save them permanently; then the connection to the database is closed. Thus, this function ensures proper organization and structuring of all scraped data for the future use or analysis. ### Saving Failed URLs to the Database ``` # Save failed URLs to the database def save_failed_url(url, reason): """ Save a failed URL and the reason for failure to the database. This function logs URLs that could not be successfully processed into the `failed_urls` table, along with the reason for the failure. Args: url (str): The URL that failed to be processed. reason (str): A brief description of the reason for the failure. Returns: None Notes: - The function assumes the existence of a `failed_urls` table with the following structure: - url (TEXT): The URL that failed. - reason (TEXT): Reason for the failure. - timestamp (TIMESTAMP): Timestamp of when the failure occurred (defaults to the current time). """ conn = sqlite3.connect(DATABASE) cursor = conn.cursor() cursor.execute(""" INSERT INTO failed_urls (url, reason) VALUES (?, ?) """, (url, reason)) conn.commit() conn.close() ``` The save\_failed\_url() function is used to save all the links that could not be scraped with the corresponding reason for this inability. There are two parameters: one is url, that is to say, the URL that cannot be scraped. The reason is a very short phrase of why this is so in the first place-maybe a network issue, or it is because of a 404 error. The function connects to the SQLite database, and then it creates a cursor to be able to execute SQL commands. This function executes an INSERT INTO query, which stores the failed URL and the reason in the failed\_urls table. It maintains URLs that failed by having columns for the URL itself, the failure reason, and a timestamp to indicate when the failure took place. Once the information has been inputted, it commits the change and thereby saves it in the database while also closing the connection. This is helpful to be able to monitor and debug scraping errors while making available the possibility of re-processing failed URLs later. ### Making a Request with Random User-Agent and Delay ``` # Request wrapper with random User-Agent and delay def make_request(url): """ Make an HTTP GET request to a specified URL with a random User-Agent and a random delay. This function sends a GET request using the `requests` library, selecting a random User-Agent from a predefined list to simulate realistic browsing behavior. A random delay is added between requests to reduce the likelihood of being flagged as a bot. Args: url (str): The URL to send the GET request to. Returns: Response: The HTTP response object returned by the `requests.get` method. Notes: - Assumes the existence of: - `USER_AGENTS`: A list of User-Agent strings. - `PROXIES`: A dictionary of proxy settings for the request. - Uses `time.sleep` to introduce a random delay between 20 and 30 seconds. - SSL verification is disabled (`verify=False`). Be cautious when using this in production. """ headers = {"User-Agent": random.choice(USER_AGENTS)} response = requests.get( url, headers=headers, proxies=PROXIES, verify=False ) time.sleep(random.uniform(20, 30)) # Random delay return response ``` The make\_request function makes a GET request to a particular URL, simulating the presence of a real user navigating through the web. First, it picks a random User-Agent from a set of User-Agent strings so that the script doesn't appear to be originating from the same browser or the same device, thereby raising fewer suspicions as a bot for the website it visits. In addition, the time.sleep is used to introduce a random delay of between 20 and 30 seconds to simulate natural browsing behavior and not to over-load the website with too many requests in a very short time. The function makes the request by calling the requests.get method using the randomly selected User-Agent and proxy settings if available. With security in mind, this function turns off SSL verification with verify=False; however, it should be used very cautiously in production environments. The function waits for the random delay after sending the request and returns the HTTP response, which may then be used to pull out data from the page. ### Extracting the Home Name from HTML ``` def extract_home_name_from_html(soup): """ Extract the home name from the HTML content. This function searches for a `` element with the attribute `data-testid='home-details-summary-headline'` in the provided BeautifulSoup object and retrieves its text content. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted home name as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if the specified `` element is not found in the HTML. """ # Search for the span containing the home name home_name_span = soup.find( 'span', {'data-testid': 'home-details-summary-headline'} ) # Extract the text content if the home name span exists home_name = ( home_name_span.get_text( strip=True ) if home_name_span else None ) # Return the extracted home name or None return home_name ``` The extract\_home\_name\_from\_html function is designed to extract the name of a home or property from a webpage by parsing the HTML content using BeautifulSoup. It looks for a specific tag in the HTML that contains the attribute data-testid='home-details-summary-headline'. This attribute is unique to the home name on the page. The function proceeds to find the tag. Inside the tag, it fetches the text using the get\_text() method. It removes leading and trailing extra spaces as it does that. In a case where the tag doesn't appear on the page, the function returns as None. Otherwise, the function returns the home name extracted as a string; this will be useful to use later in the scraper to extract information regarding the property. ### Extracting Location from HTML ``` def extract_location_from_html(soup): """ Extract the location details from the HTML content. This function searches for a `` element with the attribute `data-testid='home-details-summary-city-state'` in the provided BeautifulSoup object and retrieves its text content. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted location as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if the specified `` element is not found in the HTML. """ # Search for the span containing the location details location_span = soup.find( 'span', {'data-testid': 'home-details-summary-city-state'} ) # Extract the text content if the location span exists location = ( location_span.get_text( strip=True ) if location_span else None ) # Return the extracted location or None return location ``` The extract\_location\_from\_html function is specifically used to get the property location details from the HTML contents of a Trulia website. It uses BeautifulSoup, scanning the HTML for existence of a tag with attribute data-testid='home-details-summary-city-state', and then extracts information contained within this tag inside using the get\_text() function. This method also removes excess white spaces from the text before returning it. It returns None if the tag cannot find the tag in the HTML otherwise, it returns a location detail as a string. The string can then be used in the process of collecting data. ### Extracting Price from HTML ``` def extract_price_from_html(soup): """ Extract the price details from the HTML content. This function searches for a `
    ` element with specific classes that likely contain the price information in the provided BeautifulSoup object and retrieves its text content. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted price as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if the specified `
    ` element is not found in the HTML. - The class names used are based on a specific structure and may need adjustment if the HTML structure changes. """ # Search for the div containing the price information price_div = soup.find( 'div', class_='Text__TextBase-sc-13iydfs-0-div ' 'Text__TextContainerBase-sc-13iydfs-1 hObzVe icHjbr' ) # Extract the text content if the price div exists price = ( price_div.get_text( strip=True ) if price_div else None ) # Return the extracted price or None return price ``` This function is designed to extract a property's price information from the HTML content of the Trulia webpage. This function is searching for an
    tag containing details about the price, which carries class names that are unique for its identification. This function uses BeautifulSoup to look for the specified
    and pulls in the text with get\_text(), which gets rid of leading and trailing whitespaces around the price, then returns the price. If it finds the specified
    , the function does return the price as a string; otherwise, it simply returns None. ### Extracting Estimated Mortgage from HTML ``` def extract_estimated_mortgage_from_html(soup): """ Extract the estimated mortgage details from the HTML content. This function searches for a `
    ` element with the attribute `data-testid='summary-mortgage-estimate-details'` in the provided BeautifulSoup object and retrieves its text content. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted estimated mortgage information as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if the specified `
    ` element is not found in the HTML. """ # Search for the div containing the estimated mortgage details mortgage_div = soup.find( 'div', {'data-testid': 'summary-mortgage-estimate-details'} ) # Extract the text content if the mortgage div exists estimated_mortgage = ( mortgage_div.get_text( strip=True ) if mortgage_div else None ) # Return the extracted estimated mortgage details or None return estimated_mortgage ``` This extract\_estimated\_mortgage\_from\_html function is used to scrape the estimated mortgage information from the HTML of a Trulia property page. It searches for a specific
    element holding the mortgage estimate details identified by a unique attribute, data-testid='summary-mortgage-estimate-details'. If this element exists on the page's HTML, the function pulls its text content that contains mortgage information and strips out extra spaces using the get\_text() method with the argument strip=True. If no such element exists, the function returns None. ### Extracting Property Specifications from Trulia ``` def extract_specifications(soup): """ Extract the specifications from the HTML content. This function searches for a `
    ` element with the attribute `data-testid='facts-list'` in the provided BeautifulSoup object and retrieves its text content. The `separator=' | '` parameter is used to join the text content with a pipe character for better readability. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted specifications as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - A custom separator `|` is used to join the text parts together for better display. - Returns `None` if the specified `
    ` element is not found in the HTML. """ # Search for the parent div containing the specifications parent_div = soup.find( 'div', {'data-testid': 'facts-list'} ) # Extract the text content with a custom separator # if the parent div exists specifications = ( parent_div.get_text( strip=True, separator=' | ' ) if parent_div else None ) # Return the extracted specifications or None return specifications ``` Extract-specifications is a function, taking an input of a property's webpage. It searches for a div
    element with an attribute of data-testid='facts-list.' These might have data containing facts like the number of rooms, square footage, etc. After finding the mentioned
    element, it uses the get\_text() method to get the extracted text content. To make the extracted information more readable, the function uses the separator=' | ' argument to join the individual pieces of text with a pipe character (|). This helps in organizing the specifications in a clear and structured format. If the specified element is not found, the function returns None. ### Extracting Property Description from Trulia ``` def extract_description_from_html(soup): """ Extract the description details from the HTML content. This function searches for a `
    ` element with the attribute `data-testid='home-description-text-description-text'` in the provided BeautifulSoup object and retrieves its text content. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: str or None: The extracted description as a string if the element is found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if the specified `
    ` element is not found in the HTML. """ # Search for the description div in the soup object description_div = soup.find( 'div', {'data-testid': 'home-description-text-description-text'} ) # Extract the text content if the description div exists description = ( description_div.get_text(strip=True) if description_div else None ) # Return the extracted description or None return description ``` This is used to extract the description of a property from a Trulia property page by looking through the HTML for a
    element containing the description, identified by the attribute data-testid='home-description-text-description-text'. Once it encounters that element, it returns text contained within it using the get\_text() function, which also removes any whitespace that could be present at the start and end of the text, as the strip=True argument was passed to the method. It returns None meaning that it did not find a description for this property, in case such an element does not exist on the page. This function captures all relevant information regarding the description of property but depends on the page structure. ### Extracting Home Highlights from Trulia ``` def extract_home_highlights_from_html(soup): """ Extract the highlights details from the HTML content. This function searches for all `
    ` elements with the class `Grid__GridContainer-sc-144isrp-1 iXzkWe` within the provided BeautifulSoup object. It then iterates over these containers, extracting the key-value pairs where the key is represented by a `
    ` with the class `Text__TextBase-sc-13iydfs-0-div Text__TextContainerBase-sc-13iydfs-1 cwsXtm icHjbr` and the value by a `
    ` with the class `Text__TextBase-sc-13iydfs-0-div Text__TextContainerBase-sc-13iydfs-1 IETTU icHjbr`. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: dict or None: A dictionary of highlights with keys and their corresponding values if any highlights are found; otherwise, `None`. Notes: - The function uses the `get_text` method with `strip=True` to remove leading and trailing whitespace. - Returns `None` if no relevant highlights are found in the HTML. """ # Find all highlight containers in the soup object highlights_container = soup.find_all( 'div', class_='Grid__GridContainer-sc-144isrp-1 iXzkWe' ) # Initialize an empty dictionary to store highlights highlights = {} # Iterate over each highlight container for container in highlights_container: # Find the key element in the current container key_tag = container.find( 'div', class_='Text__TextBase-sc-13iydfs-0-div ' 'Text__TextContainerBase-sc-13iydfs-1 cwsXtm icHjbr' ) # If the key element exists, extract the key text if key_tag: key = key_tag.get_text(strip=True) # Find the value element(s) in the current container value_tag = container.find_all( 'div', class_='Text__TextBase-sc-13iydfs-0-div ' 'Text__TextContainerBase-sc-13iydfs-1 IETTU icHjbr' ) # Extract the value text if the value element exists value = ( value_tag[0].get_text(strip=True) if value_tag else None ) # Add the key-value pair to the highlights dictionary highlights[key] = value # Return the highlights dictionary if not empty, otherwise return None return highlights if highlights else None ``` The extract\_home\_highlights\_from\_html function is to scan through the highlights of a real estate property posted on the website of Trulia by scanning for multiple
    elements on the webpage's HTML containing highlight information. It does this by class, which is Grid\_\_GridContainer-sc-144isrp-1 iXzkWe. After the function finds these containers, it then looks to key-value pairs within them. The key is found in a
    with a specific class, and the value associated with that key is located in another
    with a different class. The function extracts the text for both the key and the value, removing any unnecessary spaces around the text. These key-value pairs are then stored in a dictionary, which is returned by the function. If no highlights are discovered, the function returns None. This method allows the function to retrieve all relevant features of property attributes, including amenities, conditions, or special features, all of which are stored neatly as a dictionary for further processing. ### Extracting Structured Amenities Tables ``` def extract_all_amenities_tables(soup): """ Extract all structured amenities tables from the HTML content. This function finds all `div` elements with the attribute `data-testid='structured-amenities-table-category'` within the provided BeautifulSoup object. For each table, it retrieves rows that contain subcategories and their associated details. The function then builds a dictionary where each key is a subcategory and each value is a list of details for that subcategory. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: list or None: A list of dictionaries, where each dictionary represents an amenities table with subcategories and their corresponding details as key-value pairs; otherwise, `None`. Notes: - The function assumes the structure of the HTML remains consistent. - Each `tr` element with class `Table__TableRow-sc-latbb5-0` represents a row in the table. - Each `div` with class `Text__TextBase-sc-13iydfs-0-div Text__TextContainerBase-sc-13iydfs-1 htNosZ icHjbr` represents a subcategory. - Each `span` with class `sc-9be18632-0 bEFlof` represents the details for that subcategory. """ # Find all amenities tables in the soup object amenities_tables = soup.find_all( 'div', {'data-testid': 'structured-amenities-table-category'} ) # Initialize an empty list to store all table data all_tables_data = [] # Iterate over each amenities table found for table in amenities_tables: # Find all rows in the current table rows = table.find_all( 'tr', class_='Table__TableRow-sc-latbb5-0' ) # Initialize a dictionary to store data for the current table table_data = {} # Iterate over each row in the table for row in rows: # Find the subcategory element in the row subcategory_tag = row.find( 'div', class_='Text__TextBase-sc-13iydfs-0-div ' 'Text__TextContainerBase-sc-13iydfs-1 htNosZ icHjbr' ) # Extract the subcategory text, if the tag exists subcategory = ( subcategory_tag.get_text(strip=True) if subcategory_tag else None ) # Find all details elements in the row details_tags = row.find_all( 'span', class_='sc-9be18632-0 bEFlof' ) # Extract text from all details tags details = [ detail.get_text(strip=True) for detail in details_tags ] # Add the subcategory and its details to the table data table_data[subcategory] = details # Append the current table data to the list of all tables all_tables_data.append(table_data) # Return the list of all tables if not empty, otherwise return None return all_tables_data if all_tables_data else None ``` In this segment of web scraping, the intended function of this code is supposed to grab the structured information about any amenities attached to a property listing placed on Trulia. To do so, it scans HTML content in search of sections that comprise amenity tables using div elements that have a set data-testid attribute equal to 'structured-amenities-table-category'. For each of these divisions, the function then hunts for individual rows. Individual rows are represented by tr elements with a certain class assigned to them. In any row, there is a subcategory and its details. The function hunts for subcategories by getting div elements that possess specified classes. It retrieves corresponding details for each of the retrieved subcategories. The details were stored in span elements tagged with a particular class. All these details are stored inside a dictionary with the name of the subcategory being the key and the associated details being the value inside a list. This is done for all the amenities table found on the page. The result is a list of dictionaries, where every dictionary represents a table of amenities. If no tables have been found, the function will return None. This structured approach allows the scraper to extract all information related to the amenities for each property. ### Extracting Tax Details ``` def extract_tax_details(soup): """ Extract tax details from the HTML content. This function searches for specific rows in the HTML structure that contain information about the year, tax, and assessment. It retrieves the text content of the corresponding cells and stores them in a dictionary. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content. Returns: dict or None: A dictionary containing the tax details with keys "Year", "Tax", and "Assessment" and their corresponding values as strings; otherwise, `None` if the details are not found. Notes: - The function looks for specific headers ("Year", "Tax", "Assessment") in the HTML table to find the relevant data. - Uses `find_next('td')` to get the text from the cell immediately following the header. - Returns `None` if no relevant tax details are found in the HTML. """ tax_details = {} year_row = soup.find('th', string="Year") if year_row: tax_details["Year"] = year_row.find_next('td').get_text(strip=True) tax_row = soup.find('th', string="Tax") if tax_row: tax_details["Tax"] = tax_row.find_next('td').get_text(strip=True) assessment_row = soup.find('th', string="Assessment") if assessment_row: tax_details["Assessment"] = ( assessment_row.find_next('td').get_text(strip=True) ) return tax_details if tax_details else None ``` This segment of web scraping identifies the tax-related information associated with a property listing from Trulia. The function checks within the HTML content to determine specific rows of the HTML table that carry major details including year, amount in taxes, and assessment. It first searches for the header cell that contains "Year", and upon finding it, it reads the text in the following cell-that contains the year-using find\_next('td'). Similarly, it uses the same method to search for the "Tax" and "Assessment" headers, reading the values of the tax from the adjacent cells. All these details are stored in a dictionary where keys are "Year", "Tax", and "Assessment" and the values are their corresponding values. If none of the tax details are found, the function returns None. This enables the scraper to collect and structure tax-related information for each property listing. ### Main Scraping Process ``` # Main scraping process def scrape(): """ Main function to manage the scraping process. This function first ensures that the necessary tables in the database are created using `create_tables()`. It then enters a loop to repeatedly get URLs to scrape, make requests to these URLs, and extract data from the HTML content. For each URL: - Makes a GET request to fetch the HTML content. - Uses BeautifulSoup to parse the HTML. - Extracts relevant data such as home name, location, price, mortgage, specifications, description, highlights, amenities, and tax details using the appropriate extraction functions. - Saves the extracted data to the database using `save_data()`. - Updates the URL status to indicate successful scraping using `update_url_status()`. - If an error occurs during the process, it saves the URL and the error message to `failed_urls` table and logs the failure. Args: None Returns: None Notes: - The function operates in a loop until there are no more URLs to scrape. - Each step is wrapped in a try-except block to handle errors gracefully. - Extracted data is saved as a tuple to the `product_data` table in the database. - Failed URLs and reasons are recorded in the `failed_urls` table for further analysis or retry. """ # Ensure tables are created before scraping create_tables() while True: url = get_next_url() if not url: print("No more URLs to scrape.") break try: response = make_request(url) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # Extract data home_name = extract_home_name_from_html(soup) location = extract_location_from_html(soup) price = extract_price_from_html(soup) mortgage = extract_estimated_mortgage_from_html(soup) specification = extract_specifications(soup) description = extract_description_from_html(soup) highlights = str(extract_home_highlights_from_html(soup)) # Save as string amenities = str(extract_all_amenities_tables(soup)) # Save as string tax = str(extract_tax_details(soup)) # Save as string # Prepare data to be saved data = ( home_name, location, price, mortgage, specification, description, highlights, amenities, tax ) # Save data save_data(url, data) # Pass URL along with the data update_url_status(url, 1) # Mark as scraped except Exception as e: save_failed_url(url, str(e)) print(f"Failed to scrape {url}: {e}") ``` The main scraping function coordinates the entire process of gathering property data from the Trulia website. It starts by ensuring that the necessary database tables are created to store the data. Then, it enters an infinite loop, where it repeatedly fetches URLs that need to be scraped. For each URL, it sends an HTTP GET request to fetch the web page’s HTML content. Once the HTML is retrieved, it uses BeautifulSoup to parse and extract specific property details, such as the home name, location, price, mortgage estimate, specifications, description, highlights, amenities, and tax information. Each of these details is extracted using the relevant functions that target specific HTML elements. After extracting the data, it is packaged into a tuple and saved into the database. The status of the URL is updated to indicate that the scraping was successful. If any errors occur during the scraping of a URL (for example, if a request fails or the data cannot be extracted), the URL and the error message are logged in a "failed\_urls" table for further review or retry. The process continues until there are no more URLs to scrape. Each step is wrapped in a try-except block to handle errors and ensure the process runs smoothly even if some URLs fail. ### Starting the Scraper ``` # Start the scraper if __name__ == "__main__": """ Entry point for the scraper script. This script initializes the scraping process by calling the `scrape()` function. It is the starting point when running the script directly. The `scrape()` function handles the entire scraping workflow, including checking URLs, making requests, extracting data, saving to the database, and handling errors. Args: None Returns: None Notes: - Ensure that all necessary configurations (e.g., database connection, user agents, proxies) are correctly set up before running this script. - The script runs indefinitely until all URLs to be scraped are processed. """ scrape() ``` The entry point of the script is located within the if name == "\_\_main\_\_": block, which ensures that the script runs only when executed directly, rather than being imported as a module. When the script is executed, it calls the scrape() function to begin the web scraping process. This function controls the entire workflow of scraping data from Trulia, including checking for URLs to scrape, making requests to fetch the web pages, extracting the relevant data from those pages, saving the data to the database, and handling any errors that may occur. The script runs continuously, processing URLs one after another until all URLs have been scraped. Before running the script, it is important to ensure that all configurations, such as database connections, user agents, and proxies, are correctly set up to ensure smooth operation. The script will keep running indefinitely until all tasks are completed. ## Libraries and Versions This code utilizes several key libraries to perform web scraping and data processing. The versions of the libraries used in this project are as follows: BeautifulSoup4 (v4.12.3) for parsing HTML content, Requests (v2.32.3) for making HTTP requests and Playwright (v1.47.0) for browser automation. These versions ensure smooth integration and functionality throughout the scraping workflow. ## Conclusion With this comprehensive documentation, collecting Trulia data relating to real estate is robust yet systematic. With a two-step process such as gathering all the product URLs before extracting extensive details of each property, large scale data gathering comes out properly. Key strengths of this implementation include: - A resilient scraping mechanism with built-in error handling and retry capabilities - Database-backed storage ensuring data persistence and scraping progress tracking - Anti-detection measures including proxy support, randomized user agents, and request delays - Comprehensive data extraction covering property details from prices to amenities - Modular design with separate functions for different aspects of data collection This solution balances efficiency with responsible scraping practices. It introduces delays between requests, as well as using rotating user agents. The integration of the SQLite database allows it to continue progressing even in the event that the scraping process has been halted. It can log failed URLs easily and allow the retrying of problematic cases. This can be available to developers and researchers to fetch data for real estate markets. Its adaptability to certain specifications depends on requirements; however, proper usage within guidelines from the house rules regarding scrapings for site performance. AUTHOR I’m Ambily, Data Analyst at Datahut. I specialize in developing automated data workflows that transform scattered web data into structured, decision-ready insights—particularly in real estate, e-commerce, and pricing intelligence. At Datahut, we’ve spent over a decade helping businesses harness the power of web scraping to uncover market trends, track competitor listings, and streamline research. In this blog, I’ll walk you through how to automate real estate data extraction from Trulia using Python, saving hours of manual effort while giving you real-time property insights. If you’re looking to build a scalable data pipeline for your real estate research, reach out via the chat widget on the right. We’d love to support your journey. ### Free Web Scraping & Price Tracking with n8n - Competitor Tool URL: https://www.blog.datahut.co/post/free-n8n-web-scraping-competitor-price-tracking/ Last updated: 2026-09-07T09:44:30.000Z In the fast-paced world of e-commerce, staying on top of your competitors’ pricing can be the difference between winning and losing sales. Yet, many founders and marketers lack the resources to build a full-fledged price-monitoring system. We've published a video explaining this automation here if you'd like to check out : If you’re an [Amazon](https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/), Shopify, or WooCommerce seller without a dedicated data team or hefty IT budget, this guide is for you. In it, we’ll walk you through how to automatically track competitor prices—for free—using n8n, a powerful, open-source workflow automation tool. Along the way, you’ll learn basic [web scraping ](https://www.blog.datahut.co/post/how-to-leverage-web-scraping-to-create-a-competitor-price-monitoring-strategy/)concepts, how to structure your data, and ways to receive instant alerts when prices change. Ready to take the guesswork out of your pricing strategy? Let’s dive in. ## Download and Import the n8n Workflow JSON Before diving into the details, we’ve made the complete n8n workflow JSON available for you. Please download the JSON file (link below) and import it into your n8n instance so you can follow along step-by-step as we build this automation. Having the workflow imported will let you see each node’s configuration in real time and test it as you go. 🔽 Download the n8n Workflow JSON: [Download “competitor\_price\_tracker.json”](https://www.dropbox.com/scl/fi/na6j0le9zhmqz5vwz8fhh/monitoring%5Fcompetitor%5Fprices%5F%5F%5Ffor%5Ffree.json?rlkey=16qe93baybrnjqw5f4q9wjonj&st=u6fwnohf&dl=1&ref=blog.datahut.co) 🔽 Download the Google Sheet For Running the automation: [Download the Google Sheet”](https://www.dropbox.com/scl/fi/zaif45nt9uzgz8mrhht9n/competitor%5Fprice%5Fdrop%5Fautomation.xlsx?rlkey=4vj3wz6gqp16dc2hh9hw5o49z&st=i23h513p&dl=1&ref=blog.datahut.co) ## Why Competitor Price Tracking Matters 1. Stay Competitive Online retailers adjust prices daily—or even hourly—based on demand spikes, promotions, or flash sales. Monitoring competitor prices lets you react quickly, ensuring your product remains priced right without manual effort. 2. Protect Your MarginsIf a competitor drops their price unexpectedly, you need to know before your inventory stagnates. Automated [competitor price tracking](https://www.blog.datahut.co/post/how-to-leverage-web-scraping-to-create-a-competitor-price-monitoring-strategy/) ensures you can either match the price or highlight unique value propositions (e.g., free shipping, faster delivery) to maintain your profit margins. 3. Data-Driven DecisionsInstead of relying on hunches or periodic manual checks, you’ll have a clear view of price trends over time. This allows you to spot seasonal dips, bundle opportunities, and knock-out promotions that work in your niche. 4. No IT Team? No ProblemTraditional solutions can involve custom scripting, proxy rotations, and cloud servers—often beyond the budget of small-to-midsize online stores. By leveraging n8n, you can build a free, no-code/low-code price-tracking workflow that runs on a local machine, cloud VM, or even a Raspberry Pi. ## Introducing n8n: The No-Code/Low-Code Automation Platform n8n (pronounced “n-eight-n”) is an extendable workflow automation tool that connects with over 200 services, including Google Sheets, HTTP Requests, Telegram, Slack, and many more. It supports both drag-and-drop node connections and custom JavaScript code for situations where you need extra flexibility—like parsing HTML for price data. For price monitoring, we’ll use these core features: - Schedule Trigger: Automatically kick off your workflow on a set interval (e.g., daily at 8 AM IST). - Google Sheets Integration: Maintain a master list of products and last-known prices in a Sheet. - HTTP Request + HTML Parsing: Scrape the competitor’s product page, extract the current price, and convert it into a numeric format. - Custom JavaScript (Code) Nodes: Clean and normalize data, calculate percentage changes, and format messages. - Conditional Logic (If Node): Only send alerts when the price actually changes. - Telegram Node: Push real-time price-change alerts directly to your phone or team chat. - Google Sheets Update/Append: Update the master data file to reflect new prices and append historical records to a “Price Tracking” sheet. By chaining these nodes together, you can build a fully automated competitor price-monitoring system without writing a full backend. Let’s explore each step. ## Step 1: Prepare Your Google Sheet (“Master Data”) Before configuring n8n, set up a Google Sheet with at least these columns: - row\_number: A unique numeric ID for each product (e.g., 1, 2, 3…). - product\_url: The full URL of the competitor’s product page. - price: The last recorded price (filled manually for the first run, or left blank). You can download the Google Sheet Sample from here and import to your google sheets Download google Sheets Save this Sheet and get its spreadsheet ID (the long alphanumeric string in the URL). In n8n, you’ll reference this ID to read and update your master data. ## Step 2: Schedule Trigger (Run Every Day at 8 AM IST) 1. Add a “Schedule Trigger” node in n8n.Set it to run daily at 08:00 hours.If your n8n instance is set to UTC, adjust the timezone to Asia/Kolkata.This will kick off the workflow automatically each morning, ensuring you always check prices early in the day. 2. Connect the Schedule Trigger’s main output to the next node (Google Sheets). ## Step 3: Read Product List from Google Sheets 1. Add a “Google Sheets” node and connect it to the Schedule Trigger.Choose the “Read” operation.Paste in your spreadsheet ID and sheet name (e.g., Sheet1).This node will pull in all rows of your master sheet: each item will have row\_number, product\_url, and price. 2. Test the connection to ensure n8n replicates the rows correctly. You should see an array of items like: ``` [ { "row_number": 2, "product_url": "https://coffeebros.com/products/cold-brew-coffee-blend", "price": 40 }, { "row_number": 3, "product_url": "https://coffeebros.com/products/medium-roast-coffee-blend?", "price": 30 }, { "row_number": 4, "product_url": "https://coffeebros.com/products/dark-espresso-roast-coffee", "price": 10 } ] ``` ## Step 4: Loop Over Each Product (Split In Batches) 1. Add a “Split In Batches” node and connect it to the Google Sheets node.This ensures that each product is processed individually, so if one item fails or takes longer, it doesn’t block the rest. 2. Keep the batch size as 1 (default). n8n will now pass one item at a time to the next nodes, effectively treating each product as its own mini-workflow. ## Step 5: Delay to Avoid Request Blocking 1. Add a “Wait” node (label it “Delay to Avoid Request Blocking”).Set it to wait for 20 seconds before sending the next request.Many e-commerce sites have rate limits or IP-based filters; spacing out your HTTP requests reduces the chance of being temporarily blocked. 2. Connect the “Split In Batches” output to this “Wait” node. Now, for each product, n8n waits 20 seconds before performing the next action. ## Step 6: Fetch the Competitor’s Product Page (HTTP Request) 1. Add an “HTTP Request” node and connect it to the “Wait” node.URL: ={{ $json.product\_url }} (this pulls the URL from the current batch item).Method: GET.Headers:User-Agent: A realistic browser string (e.g., Chrome on Windows or Firefox on Mac).Accept: text/html,application/xhtml+xml,application/xml;q=0.9,\*/\*;q=0.8Accept-Language: en-US,en;q=0.9Options: Enable “Follow Redirects” if the site uses permanent/temporary redirects. 2. Save and test: If successful, n8n should fetch the raw HTML of the product page. ## Step 7: Parse HTML and Extract Current Price 1. Add an “HTML” node (named “Parse Data From The HTML Page”) and connect it to the “HTTP Request” node. 2. Operation: extractHtmlContent. 3. Extraction Values: 4. Key: current\_price 5. CSS Selector: Targets the price element on the competitor’s site. For CoffeeBros, for example, it might be: ``` .price__regular > span.price-item--regular ``` - Adjust this selector based on your competitor’s page structure (right-click → Inspect their price element to confirm). - An easier way is to right click, open google chrome ai assistant and then ask “share the css selector for extracting price “ share the value you see as pricing along with it. - Output: The node will return a JSON field current\_price containing something like "$9.79". ## Step 8: Data Cleaning (Convert to Numeric, Preserve Original) 1. Add a “Code” node (named “Data Cleaning”) connected to the “HTML” node. 2. Purpose: Clean and parse the scraped price string into a pure number, then merge it with the original data pulled from Google Sheets. 3. Snippet (JavaScript) (this matches what’s in your downloaded JSON): ``` // Get scraped result const scraped = items[0].json; // Retrieve the original data from the “Delay to Avoid Request Blocking” node const original = $items("Delay to Avoid Request Blocking")[0].json; // Convert "$9.79" → 9.79 const priceNum = parseFloat((scraped.current_price || "").replace(/[^0-9.]+/g, "")); // Original price from Google Sheet const lastPrice = parseFloat(original.price); return [{ json: { product_url: original.product_url, row_number: original.row_number, last_price: lastPrice, current_price: priceNum } }]; ``` 1. What This Does: 2. Takes the current\_price (e.g., "$9.79") and removes non-numeric characters to get 9.79. 3. Pulls last\_price (e.g., 9.99) from the original Google Sheets row. 4. Returns a unified JSON object with product\_url, row\_number, last\_price, and current\_price. 5. In your case the code will be different, but don't worry - just paste the input and expected output to the chatgpt instance and it will get you the correct answer. 6. In this case the input is ``` [ { "current_price": "$9.79" } ] ``` - The expected output is ``` [ { "product_url": "https://coffeebros.com/products/dark-espresso-roast-coffee", "row_number": 4, "last_price": 10, "current_price": 9.79 } ] ``` modify this with yours and chatgpt will help you generate the code ## Step 9: Compare and Normalize Prices 1. Add a “Code” node (named “Compare and Normalize”) connected to the Data Cleaning step. 2. Purpose: Check whether the price changed, calculate the percentage difference, and attach a timestamp. 3. Snippet (JavaScript): ``` // Incoming item const item = items[0].json; // Did the price change? const priceChanged = item.last_price !== item.current_price; // Calculate percent difference let priceDiffPct = null; if (item.last_price && item.last_price !== 0) { priceDiffPct = ( ((item.current_price - item.last_price) / item.last_price) * 100 ).toFixed(2); } return [{ json: { product_url: item.product_url, last_price: item.last_price, current_price: item.current_price, price_changed: priceChanged, price_diff_pct: priceDiffPct !== null ? parseFloat(priceDiffPct) : null, timestamp: new Date().toISOString(), } }]; ``` - What This Does:Compares last\_price vs. current\_price.Calculates price\_diff\_pct (e.g., -1.99 if price dropped).Tags a UTC timestamp for logging/alerts. ## Step 10: Fix Tab Issues (Optional Sanitization) 1. Add a “Code” node (named “Fixing the Broken Tab”) connected to the Compare and Normalize step. 2. Purpose: Clean any trailing tabs or whitespace from keys in the JSON (sometimes Google Sheets columns come with hidden tabs). 3. Snippet (JavaScript): ``` if (!items.length) return []; const input = items[0].json; const cleaned = {}; for (const key in input) { const cleanKey = key.replace(/\t/g, "").trim(); cleaned[cleanKey] = input[key]; } return [{ json: { product_url: cleaned.product_url, last_price: cleaned.last_price, current_price: cleaned.current_price, price_changed: cleaned.price_changed, price_diff_pct: cleaned.price_diff_pct, timestamp: cleaned.timestamp, } }]; ``` 1. What This Does:Strips any accidental \\t characters from the JSON keys.Ensures your price\_changed or price\_diff\_pct fields aren’t malformed. ## Step 11: Conditional Logic (If Price Changed → Send Alert) 1. Add an “If” node connected to the “Fixing the Broken Tab” step.Condition: Check if price\_changed === true.Logic:If true, route data to the “Format for Output” node and append to the “Price Tracking” sheet.If false, skip notification but still update the master sheet. 2. Why This Matters: You don’t want to spam yourself or your team with notifications every single day if prices haven’t changed. This node ensures you only act when there’s a real update. ## Step 12: Format Alert for Telegram 1. Add a “Code” node (named “Format for Output”) connected to the “If” node’s true branch. 2. Purpose: Build a user-friendly message with emojis, price details, and IST timestamp. ``` const itemsFormatted = []; for (const data of items) { // Determine alert type const isDrop = data.json.price_diff_pct < 0; const alertType = isDrop ? "📉 *Price Drop Alert*" : "📈 *Price Hike Alert*"; // Convert UTC → IST (+5:30) const date = new Date(data.json.timestamp); const istOffset = 5.5 * 60 * 60 * 1000; const istDate = new Date(date.getTime() + istOffset); const formattedDate = istDate.toLocaleString("en-IN", { day: "2-digit", month: "short", year: "numeric", hour: "2-digit", minute: "2-digit", hour12: true }); // Build message const message = `${alertType} ``` 🛍️ Product: (${data.json.product\_url})💸 Last Price: data.json.lastprice💰 ∗CurrentPrice:∗{data.json.last\_price} 💰 Current Price: data.json.lastp​rice💰 ∗CurrentPrice:∗{data.json.current\_price}📊 Change: ${data.json.price\_diff\_pct > 0 ? "+" : ""}${data.json.price\_diff\_pct}%🕒 Time: ${formattedDate}\`; - What This Does:Checks if the price dropped or rose (using emojis 📉 for drop, 📈 for hike).Converts the UTC timestamp into IST (your local time).Assembles a Markdown-formatted message with product URL, last/current prices, percentage change, and time. ## Step 13: Send Alert via Telegram 1. Add a “Telegram” node connected to the “Format for Output” node.Chat ID: Your personal (or team) Telegram chat ID (e.g., 7739958732).Text: ={{ $json\["message"\] }}.Ensure you’ve already created a Telegram Bot, copied its token into n8n’s credentials, and started a conversation with the bot so you can capture your chat ID. 2. Result: Whenever a price changes, you’ll get a real-time Telegram alert like: 📉 \*Price Drop Alert\* ``` 🛍️ *Product:* (https://coffeebros.com/products/cold-brew-coffee) 💸 *Last Price:* $9.99 💰 *Current Price:* $9.49 📊 *Change:* -5.01% 🕒 *Time:* 02 Jun 2025, 11:30 AM ``` ## Step 14: Update Master Data File (Google Sheets) 1. After the “If” node (both branches), update the original row in the master Google Sheet so that the next run has the latest price.Add a “Google Sheets” node (named “Update Master Data File”).Operation: update.Spreadsheet ID & Sheet Name: Same as your original.Columns Mapping:product\_url → ={{ $json.product\_url }}price → ={{ $json.current\_price }}Matching Column: product\_url (so the node knows which row to overwrite). 2. Why: This replaces the old price with the newly scraped price. When tomorrow’s run begins, n8n will see the updated “last\_price” and calculate changes relative to that value. ## Step 15: Append to “Price Tracking” Sheet (Historical Log) 1. Still on the “If” node’s true branch, after sending the Telegram alert, append a row to a separate sheet (tab) called “Price Tracking.” 2. Add a “Google Sheets” node (named Price Track Sheet1). 3. Operation: append. 4. Spreadsheet ID: Same file. 5. Sheet Name: Tab ID or name for “Price Tracking” (e.g., Price Tracking). 6. Columns Mapping (example): ``` timestamp: ={{ new Date($json.timestamp).toLocaleString('en-IN', { timeZone: 'Asia/Kolkata', day: '2-digit', month: 'short', year: 'numeric', hour: '2-digit', minute: '2-digit', hour12: true }) }} product_url: ={{ $json.product_url }} current_price: ={{ $json.current_price }} last_price: ={{ $json.last_price }} price_diff_pct: ={{ $json.price_diff_pct }} price_changed: ={{ $json.price_changed }} ``` - Check “Use Append” so each new price-change event is logged below the last. - Result: You’ll build a date-stamped history of every price change, allowing you to chart trends and compute metrics like average discount periods, seasonal dips, or high-demand spikes. ## Step 16: Final Wait Before Next Batch 1. After appending to the “Price Tracking” sheet (and/or updating master data), you may add a short “Wait” (e.g., 1 minute) before n8n moves on to the next item in the “Split In Batches” loop.This extra delay can be useful if your competitor site has very strict rate limits or if you want to minimize load spikes. 2. Add a “Wait” node (named “Wait1”) connected to the “Price Track Sheet1” node, set to 1 minute. Then connect it back to the “Update Master Data File” node so the cycle completes properly. ## Putting It All Together Below is a high-level summary of the node flow: 1. Schedule Trigger (08:00 IST daily) 2. → Google Sheets (Read master data) 3. → Split In Batches (Process one product at a time) 4. → Wait (Delay to avoid Request Blocking) (20 s) 5. → HTTP Request (Scrape competitor page) 6. → HTML Extract (Parse current\_price via CSS selector) 7. → Data Cleaning (Code Node) (Extract numeric price + merge with original) 8. → Compare and Normalize (Code Node) (Compute price\_changed, price\_diff\_pct, and timestamp) 9. → Fixing the Broken Tab (Code Node) (Remove stray tabs from keys) 10. → If (price\_changed === true)True →Format for Output (Code Node) (Build Telegram message)→ Telegram (Send alert)→ Price Track Sheet1 (Append historical row)→ Wait1 (1 minute)→ Update Master Data File (Overwrite last price)→ loops back (next item)False →Update Master Data File (Overwrite last price without alert)→ loops back By following these steps, you’ll have a fully automated, self-running competitor price tracker that runs every morning, scrapes your competitor’s product pages, compares prices, sends alerts to Telegram only when prices change, and builds a historical log for analysis. ## Best Practices & Tips 1. Rotate User-Agents or Proxies (If Needed)If you’re scraping a large number of products or more restrictive websites, consider adding random User-Agent strings or lightweight proxy rotations. You can store a list of proxies in a separate Google Sheet and rotate them in your HTTP Request node. 2. Monitor for “No Data” or 404Sometimes product pages get removed (e.g., discontinued items). Add a second “If” node after the HTTP Request to check if current\_price is empty or the HTTP status is 404\. If so, send an alert that the product might have been delisted. 3. Error HandlingUse n8n’s built-in error workflow or add a “Catch Error” node to notify you if any part of the workflow breaks (e.g., CSS selectors changed). This way, you’ll know if scraping fails and can update your selectors or fix permission issues. 4. Use Descriptive Node NamesAs your automation grows, naming nodes like “Parse Data From The HTML Page” or “Format for Output” makes troubleshooting easier. Always label custom code nodes clearly (e.g., “Calculate Price Diff” vs. “Code [#3](https://www.blog.datahut.co/blog/hashtags/3)”). 5. Clean Up Google Sheet FormattingAvoid hidden tabs/spaces in column headers—this can cause mismatches when matching columns in n8n. Always trim whitespace or special characters from header names. 6. Back Up Your WorkflowExport your n8n workflow as a JSON backup. That way, if you ever need to move servers or replicate the setup for a team member, you can import it quickly. 7. Use ChatGPT to Generate or Customize Code SnippetsEvery site’s HTML structure is different, and your needs may vary (e.g., selecting a different CSS class, handling currency symbols, or adjusting the date format). If you ever need to tweak or rewrite a JavaScript snippet, simply ask ChatGPT for help. For example, you might prompt:“I need a custom n8n Function node code to parse the price from my competitor’s HTML. Their price is inside ₹1234. Generate a JavaScript snippet that:Reads the inner text of .our-price.Strips out the currency symbol (₹) and any commas.Returns a number field called current\_price.”ChatGPT will then provide you with a ready-to-paste code snippet. If you need further adjustments—say, changing the date format or adding error handling—just continue the conversation:“Now modify that snippet so it also sets price\_changed to true/false by comparing with a field called last\_price, and calculates a price\_diff\_pct.”This way, you can avoid spending hours debugging JavaScript and instead get a clean, tested snippet you can drop into your “Code” node. Think of ChatGPT as an on-demand coding assistant: describe your desired inputs and outputs, and let the model craft the logic. Always review the generated code for your own site’s selectors and data types, but in most cases you’ll be able to paste-and-go. ## Final Thoughts Automated competitor price tracking levels the playing field for e-commerce founders and marketers who lack big budgets or technical teams. With just a handful of n8n nodes—and the JSON workflow you downloaded—you can build a robust system that: - [Web scrapes](https://www.blog.datahut.co/post/what-are-web-scraping-services-and-why-do-they-matter/) competitor product pages for price data - Cleans, normalizes, and compares prices to your historic values - Sends alerts only when prices actually change - Logs each price update into a historical database for trend analysis All without any paid subscriptions to commercial competitive intelligence platforms. And because n8n is open-source, you maintain full control over your data and workflows. Watch our YouTube Video for Step by Step Guide on this : [https://www.youtube.com/watch?si=7Jr2rVID3HUwumVR&v=a9esT732mmE&feature=youtu.be](https://www.youtube.com/watch?si=7Jr2rVID3HUwumVR&v=a9esT732mmE&feature=youtu.be&ref=blog.datahut.co) If you’re ready to take your pricing strategy to the next level—or want to explore more advanced competitor monitoring solutions—reach out to the Datahut team today. We specialize in ethical, large-scale data extraction and can help you scale this workflow, add richer data points (e.g., dynamic promotions, stock status, review counts), or integrate with your existing BI dashboards. ## ❓ Frequently Asked Questions ### 1\. Do I need to know coding to set up this n8n price tracking workflow? Not at all. This guide is designed for non-developers. With n8n’s drag-and-drop interface, you can build and run the automation visually. Any code snippets provided (like price cleanup) are plug-and-play—you just copy and paste them. ### 2\. Can this method work for any e-commerce website? Yes, mostly. It works best with websites where product prices are visible in the raw HTML. You’ll just need to inspect the page and copy the right CSS selector for the price. For JavaScript-heavy sites, you might need extra setup like Puppeteer or a headless browser. ### 3\. How often will the price tracking run? Once per day by default. In this guide, the workflow is scheduled to run every day at 8 AM IST. You can change this timing in the “Schedule Trigger” node inside n8n to suit your needs—hourly, weekly, etc. ### 4\. Is this solution completely free? Yes. n8n is open-source and free to self-host. Google Sheets and Telegram (used for alerts) are also free services. The entire workflow can be built and deployed without spending anything. ### 5\. Can I get alerts if the product page is removed or fails to load? Yes, absolutely. You can add a simple “If” condition to check if the price is missing or if the HTTP status code is 404\. If so, you can trigger a Telegram alert indicating that the product might be out of stock or discontinued. ### About the Author I'm Tony Paul, founder of Datahut,. We help e-commerce brands and retailers use data for smarter pricing and competitive insights. With over 14 years of experience working alongside large retail chains and major consumer brands, I've has guided countless businesses through the complexities of web scraping, data normalization, and pricing strategy. My passion lies in making sophisticated data tools accessible to everyone—from bootstrapped startups to established enterprises—so they can make faster, data-driven decisions without a huge IT budget. I also enjoy sharing actionable tips and best practices to empower fellow founders and marketers. Have questions or want to scale up your pricing strategy? Contact[ us.](https://www.datahut.co/?ref=blog.datahut.co#contact) [Connect with me on LinkedIn](https://www.linkedin.com/in/tonypauldh/?ref=blog.datahut.co) ### Why Retailers Need Web Scraping, Product Matching, BI URL: https://www.blog.datahut.co/post/why-retailers-should-invest-in-web-scraping-product-matching-and-bi/ Last updated: 2026-07-23T07:48:32.000Z Success in modern retail hinges on strategic insights as much as product quality. Leading retailers gain a competitive edge by integrating web scraping, product matching, and business intelligence (BI) to analyze market trends, benchmark against competitors, and anticipate customer needs. In an industry driven by data, these tools empower businesses to make informed decisions and maintain a leadership position. ## The Synergistic Power of Web Scraping, Product Matching, and BI Web scraping serves as the foundation, automatically extracting vast amounts of publicly available data from the internet. This automated process gathers information from competitor websites, online marketplaces, social media platforms, and review sites. The types of data collected are vast, ranging from product titles, prices, and promotional details to customer reviews, social media sentiment, and competitor strategies. Automating this data collection process saves human effort and time. This raw, scraped data is then transformed through product matching techniques. Often enhanced and automated by Machine Learning (ML) and Artificial Intelligence (AI) algorithms, product matching identifies identical or highly similar products across different online platforms. This capability is crucial for making accurate comparisons on key factors such as current and discounted pricing, product attributes like size or color, and stock availability, turning disparate data points into structured, comparable insights. Finally, Business Intelligence (BI) tools integrate and analyze this matched product data, combining it with internal sales figures, customer information, and other relevant datasets. BI solutions translate complex datasets into actionable insights, typically presented through intuitive reports, dashboards, and predictive analytics. The integration of web scraping, product matching, and business intelligence yields tangible benefits. For instance, a multinational retail brand leveraged these services to overcome data extraction challenges, allowing them to identify market trends, optimize pricing strategies, improve inventory management, and enhance customer targeting and personalization. This approach led to significant results, including a 67% increase in revenue through the use of optimized pricing strategies. They were able to optimize stock levels and avoid stockouts by gaining real-time insights into product demand and availability. Web scraping also informs dynamic pricing models that adjust prices based on market factors, predicts product demand based on online patterns, and helps create precise customer profiles for targeted promotions. Furthermore, analyzing customer reviews obtained via web scraping directly contributes to product optimization and development by highlighting customer expectations and feedback. ## Key Benefits for Retailers Investing in a web scraping, product matching, and BI solution offers numerous benefits for retailers: ### 1\. Competitive Pricing and Profit Optimization Real-time Price Monitoring: Web scraping serves as a fundamental tool for retailers to track competitor prices continuously. This capability allows businesses to react quickly to price changes in the market, helping them maintain a competitive edge. Access to this real-time data is essential for implementing dynamic pricing strategies, which involve adjusting prices based on demand, competition, and other factors. Dynamic pricing aims to maximize profits while ensuring offerings remain attractive to customers. Tools that utilize AI algorithms can implement dynamic pricing models that adapt to market fluctuations. Identifying Pricing Opportunities: By analyzing competitor pricing and market trends, retailers can identify opportunities to adjust their prices. This allows for optimizing profit margins and increasing revenue. For instance, a multinational retail brand used scraped data to optimize its pricing strategies, which contributed to an increase in revenue. Tools can analyze historical pricing data to forecast future movements, helping to set competitive prices without sacrificing profit. Monitoring Minimum Advertised Price (MAP) Compliance: For brands, web scraping helps monitor online platforms to ensure that resellers adhere to predefined MAP policies. This is important for protecting the brand's image and maintaining pricing integrity across various sales channels. Manually monitoring prices across numerous resellers is impossible, making web scraping a necessary tool for MAP monitoring. ### 2\. Optimizing Product Assortment and Visibility Identifying Best-Selling Products and Market Gaps: Analyzing competitor product offerings, customer reviews, and market trends provides retailers with insights into popular products and helps uncover gaps in the market that they can fill. Web scraping can extract data such as product descriptions, specifications, prices, and availability from competitor websites, which can inform product development and strategies. Improving Product Listings: Insights gleaned from competitor listings, including elements like descriptions and keywords, can be invaluable for optimizing a retailer's product listings. Aligning product descriptions and categorizations with sales performance and historical search can enhance search engine performance. Benchmarking Product Features and Quality: Product matching, often enhanced by ML/AI, allows for direct comparison of identical or similar products across platforms based on features and customer reviews. This enables retailers to benchmark their offerings against those of their competitors and identify areas for improvement, directly contributing to product optimization and development. Analyzing customer feedback from various sites can help highlight recurring issues or frequently praised features, allowing for the fine-tuning of product offerings. ### 3\. Enhancing Inventory Management Real-time Stock Level Monitoring: While the source explicitly mentioning tracking competitor stock levels via web scraping is (though this source seems to refer to MAP compliance, not stock levels), the general capability of web scraping to extract data from competitor websites implicitly supports the idea of monitoring their inventory status if that data is publicly available on their sites. This real-time visibility into competitor product availability could provide valuable insights. Furthermore, alternative data sources like geospatial and foot traffic data can provide insights into customer behavior in physical stores, which can inform inventory decisions for brick-and-mortar locations. Effective inventory management is crucial for ensuring products are available when needed while minimizing costs. AI-driven inventory management solutions can enhance stock accuracy and reduce waste. Demand Forecasting: Accurate demand forecasting is vital for planning inventory levels and reducing stockouts or excess inventory. Businesses can utilize historical data and market trends to predict future demand. By analyzing market trends, competitor inventory levels (inferred capability from scraping competitor sites), and historical sales data, retailers can improve the accuracy of their demand forecasting. Machine learning models can provide precise demand forecasts, enabling optimization of supply chain operations. Predictive analytics models use historical data to forecast future actions and can estimate future sales based on historical data and market trends, helping businesses manage inventory and allocate resources effectively. Predictive analytics can enhance customer retention and acquisition, and businesses can use customer lifecycle prediction to optimize resource allocation. Product Popularity Prediction focuses on forecasting which products will gain market traction, which is vital for allocating resources effectively and maximizing sales. It also helps with inventory management by predicting demand. AI enables effective resource allocation for staffing and inventory, particularly during peak buying seasons. Category affinity modeling can help forecast demand for related categories and reduce stockouts or overstock situations by aligning inventory with consumer buying behavior. Optimizing Inventory Across Platforms: For retailers operating on multiple online marketplaces, web scraping can help aggregate and unify product data from various sources, allowing for a more centralized view of product information and potentially informing strategies for optimizing inventory distribution. Although the sources don't explicitly detail unifying inventory data across platforms using scraping, they highlight the use of AI-driven inventory management solutions for enhanced stock accuracy and reduced waste, and the general application of scraped data for efficient inventory management. Accurate predictions resulting from data analysis help businesses maintain optimal inventory levels, reducing stockouts and overstock situations. ### 4\. Understanding Customer Behavior and Sentiment Analyzing Customer Reviews and Feedback: Web scraping is essential for analyzing customer reviews and sentiment. It enables the collection of customer reviews from various online sources, including review sites, social media, and competitor websites. This provides valuable insights into customer opinions, preferences, pain points, and concerns. Analyzing customer reviews helps optimize and develop products. Tools that utilize natural language processing (NLP) can analyze customer feedback at scale and measure customer satisfaction. Actively seeking customer feedback through surveys or reviews helps identify areas for improvement, and addressing concerns promptly can prevent churn. Identifying Emerging Trends: Monitoring various online sources, such as social media, forums, and review sites, allows retailers to identify emerging trends and shifts in consumer behavior. Analyzing data from multiple channels helps identify effective acquisition sources. Web scraping allows businesses to monitor their competitors’ activities on a granular level, which can unravel key insights into market trends. AI enables the analysis of vast amounts of data quickly and accurately, helping identify patterns and trends that may not be visible through traditional analytics. AI-driven analytics can detect style trends. Emerging pattern detection focuses on identifying new and evolving patterns within data, which is crucial for businesses seeking to innovate and adapt to changing consumer needs. Retailers can predict product demand based on online search patterns, social media buzz, or even news events, allowing them to stay ahead of shifting consumer preferences. Market trend analysis can help in forecasting future demand and aid in product development and innovation. Personalizing Marketing Efforts: Understanding customer behavior and preferences allows retailers to create more targeted and personalized marketing campaigns, leading to higher engagement and conversion rates. Businesses can create personalized experiences based on individual customer data. AI-driven solutions can analyze customer data to create highly personalized marketing strategies. Personalized marketing messages can significantly increase engagement rates, fostering a deeper connection with customers. AI-driven analytics can analyze consumer behavior, enabling retailers to deliver personalized recommendations that enhance customer satisfaction and drive sales. Utilizing data to understand customer behavior and preferences is key at each stage of the customer journey. Alternative data can be used to improve customer segmentation and target promotions more effectively. Tailoring communication based on customer preferences and past interactions can enhance engagement. Predicting which products customers are likely to buy together can enhance sales through recommendations based on previous purchases. E-commerce platforms can leverage basket analysis to provide personalized product recommendations based on past purchases. Customer experience personalization refers to tailoring interactions and offerings to meet the individual preferences and needs of customers, enhancing satisfaction and loyalty. Understanding customer behavior enables businesses to personalize marketing messages and product recommendations. ### 5\. Data-Driven Decision Making and Strategic Planning Comprehensive Market Analysis: The combination of web scraping, product matching, and business intelligence (BI) tools provides retailers with a holistic view of the market. Web scraping serves as a crucial tool for extracting vast amounts of data from various online sources. This includes monitoring competitor websites for product offerings, pricing strategies, and marketing tactics. It also involves collecting customer reviews and sentiment from review sites, social media, and forums, as well as tracking broader market trends and technologies. Product matching is essential for comparing specific products across different online sources, especially for competitive pricing and assortment analysis. This process involves identifying identical products to understand competitor pricing, attributes, and availability. Complementary data sources like social media trends, search queries, and geospatial data (foot traffic) also provide valuable alternative data for market insights. Once collected, this raw data needs to be cleaned, structured, and integrated. Business intelligence tools, such as Tableau or Power BI, and analytics software visualize this integrated data, helping to identify patterns, trends, and relationships that may not be visible through traditional analysis alone. This comprehensive market analysis empowers retailers to make informed, data-driven strategic decisions and adapt effectively. Identifying New Opportunities: By analyzing data gathered through web scraping and integrating it into business intelligence (BI) systems, retailers can effectively identify new business opportunities and potential areas for growth. Analyzing customer reviews and feedback collected via scraping provides insights into customer sentiment, preferences, and pain points, which can inform product optimization and development. Monitoring market trends, including shifts in consumer preferences and the emergence of new styles or technologies (often identified through emerging pattern detection and product trend analysis), allows retailers to anticipate demand and adapt their product offerings accordingly. Analyzing competitor activities helps identify market gaps that a retailer can fill. Geographic trend mapping helps businesses understand regional preferences and identify untapped markets for expansion. Understanding digital and purchase behavior patterns through data analysis allows for identifying opportunities for personalization in marketing and making recommendations. Alternative data signals, such as job posting data or patent filings, can even provide early indications of market shifts and emerging opportunities. This allows businesses to pivot or innovate before a trend starts to decline. Improving Operational Efficiency: Automating data collection and analysis through these solutions saves valuable time and resources, allowing retailers to focus on other critical aspects of their business. Web scraping automates the tedious process of extracting data from websites, eliminating the need for manual gathering. AI and Machine Learning models analyze vast amounts of data quickly and accurately, streamlining the analysis process. Real-time data collection and processing are essential for timely insights and enabling quick adjustments to strategies. Accurate demand forecasting, leveraging historical data, market trends, and potentially competitor inventory levels, is significantly improved. Precise forecasts lead to better inventory management, which reduces costly stockouts and overstocking issues. AI-driven inventory management solutions specifically enhance stock accuracy and reduce waste. Data analysis also supports optimizing supply chain operations and logistics management, which can reduce costs and improve delivery times. Automating marketing processes, such as abandoned cart recovery or dynamic pricing based on analyzed patterns, also enhances efficiency. The integration of data from multiple sources streamlines data access for analysis. Data cleaning and transformation processes ensure data quality, which is crucial for accurate analysis and avoiding wasted effort. ### Return on Investment (ROI) The ROI of investing in a web scraping, product matching, and business intelligence (BI) solution can be significant for retailers. Studies indicate that businesses that leverage automated data extraction experience improvements in pricing accuracy, reductions in stockouts, and increased operational efficiency. By automating data collection and analysis, businesses optimize pricing strategies—leading to revenue increases of up to 67% for some brands. Dynamic pricing, demand forecasting, and inventory management reduce stockouts and overstocking, cutting costs while improving efficiency. Additionally, AI-driven insights enhance customer targeting, boosting conversion rates and loyalty. Companies leveraging these tools are 19x more likely to be profitable and 23x more likely to acquire customers, proving the tangible value of data-driven decision-making. Beyond revenue growth, these solutions drive operational efficiency. Automated web scraping eliminates manual data gathering, reducing labor costs and errors. AI-powered analytics streamline inventory and supply chain management, minimizing waste and improving logistics. With high-quality, real-time insights, retailers can act faster, predict trends, and personalize customer experiences—turning data into a sustainable advantage. The result? Lower costs, higher profits, and long-term scalability. ## Addressing Challenges and Ethical Considerations While the benefits are clear, it’s important to approach web scraping and data analysis responsibly. Retailers should: - Respect Data Privacy: Only collect publicly available data and comply with relevant laws and platform terms of service. - Ensure Data Quality: Invest in data cleaning and validation to avoid errors in decision-making. - Balance Automation with Human Insight: Use technology to inform, not replace, expert judgment. ## Conclusion In the data-driven era, a web scraping, product matching, and business intelligence solution is no longer a luxury but a necessity for retailers seeking to thrive. By providing a comprehensive understanding of the market, competitors, and customers, these tools empower retailers to make informed decisions, optimize their operations, and achieve sustainable growth and profitability. At Datahut, we specialize in delivering reliable web scraping, advanced product matching, and powerful business intelligence solutions tailored for the retail industry. Our expertise helps you unlock actionable insights, optimize pricing, streamline inventory, and stay ahead of the competition. Contact [Datahut](https://www.datahut.co/?ref=blog.datahut.co) today for a free consultation or to learn how our solutions can accelerate your retail growth. AUTHOR I’m Ashmi, part of the Marketing Team at Datahut. I work closely with our data analysts and engineers to translate complex data solutions into real-world value for retailers and e-commerce brands. At Datahut, we’ve spent over 10 years building custom web scraping and product intelligence solutions for businesses across the globe. From product matching to market tracking and competitive benchmarking, we’ve seen firsthand how data can empower retailers to make faster, smarter, and more profitable decisions. If you're a retailer looking to gain an edge through automation, pricing insights, or competitive intelligence, reach out using the chat widget on the right—let’s explore how we can help. ### Pricing Fatigue: Key Signs and How Businesses Can Fix It URL: https://www.blog.datahut.co/post/6-signs-your-business-has-pricing-fatigue-and-how-to-fix-it/ Last updated: 2026-07-23T07:48:32.000Z You've got a wonderful product, a starving market, and a fantastic team but pricing continues to feel like tightrope walking. One minute you're doubting your figures, the next minute you're following the lead of competitors with discounts, and before you know it, your margins are eroding and your team's confidence is dented. Are You Suffering from Pricing Fatigue? Here’s How to Find Out We recently ran a short pricing fatigue assessment survey with a group of e-commerce founders and leaders. Each response was scored on a scale of 0 to 100, helping us identify how well their pricing strategies were holding up in today’s fast-moving landscape. The results were eye-opening: Average Pricing Fatigue Score: 32 (well below the benchmark pass mark of 40) 82% didn’t have a clear Standard Operating Procedure (SOP) for pricing decisions. After the survey, a few founders reached out to dig deeper and fix the cracks in their pricing strategy. We started with two core actions: - Define a clear SOP for pricing updates - Incorporate external data sources into pricing and ad planning We piloted this approach with two of them, and the early results speak volumes: - Average selling price increased by 7% - Return on Ad Spend (ROAS) jumped by 19% We’re still in the early stages, but the momentum is promising. If you’re curious about where your pricing strategy stands and want to spot fatigue before it hits your margins, take our free pricing fatigue assessment here! ![If you’re curious about where your pricing strategy stands and want to spot fatigue before it hits your margins, take our...](https://www.blog.datahut.co/content/images/2026/07/img-453.png.webp) Welcome to the universe of pricing fatigue- a stealthy business growth killer and one of the most common business pricing problems today. It appears quietly at first: late pricing decisions, fragmented data, internal disagreements. Left unchecked, it balloons into pricing mistakes, missed opportunities, and organizational disarray. Research cited by [Forbes](https://www.forbes.com/2010/03/25/profit-gain-value-mckinsey-sears-whirlpool-cmo-network-rafi-mohammed.html?ref=blog.datahut.co) highlights that even a modest 1% increase in price can boost operating profits by up to 11%, underscoring the powerful impact of effective pricing strategies on the bottom line If you're a founder, revenue leader, or scaling team member, you’ve likely felt the pressure of trying to figure out how to price products correctly. Without a strong, consistent pricing strategy, you risk more than just confusion, you risk growth. In this blog, we’ll explore 6 clear signs that your business is facing pricing fatigue, what they mean, and how to overcome them before they derail your growth. From inconsistent pricing across teams to avoiding pricing experiments, these red flags can reveal deep-rooted problems in how your business makes pricing decisions. Let's get started and regain control over your pricing future. ## 1\. Inconsistent Pricing Across Teams Imagine: Marketing launches a promotion with bold discounts, Sales negotiates a different deal to close a lead, and Finance questions the margin hit after the fact. Sound familiar? ![Inconsistent Pricing Across Teams](https://www.blog.datahut.co/content/images/2026/07/img-454.png.webp) Why This Happens: This kind of inconsistency often stems from siloed decision-making and the absence of a unified pricing strategy. When teams operate without shared rules or visibility into each other's decisions, it creates confusion, undermines trust with customers, and erodes profit margins. The result? A pricing tug-of-war that benefits no one. The Fix: Establish a centralized pricing playbook that clearly defines discount thresholds, promotional guidelines, and approval workflows. Tools like Configure Price Quote (CPQ) software or dedicated pricing committees can help ensure every team from Marketing to Sales to Finance operates from the same set of pricing rules. The goal is simple: build cross-functional alignment that protects your margins while delivering a consistent customer experience. ## 2\. Delayed Responses to Market Changes Imagine: A competitor slashes prices or a new trend shifts customer demand but your team takes weeks to react. By the time you adjust, the moment (and the market) has moved on. ![Delayed Responses to Market Changes](https://www.blog.datahut.co/content/images/2026/07/img-455.png.webp) Why This Happens: Slow response times often come from bureaucratic decision-making, lack of real-time data, or fear of making the wrong move. Without a dynamic pricing framework in place, businesses become reactive instead of proactive. In fast-moving markets, hesitation doesn’t just cost you customers—it hands them over to more agile competitors. The Fix: Adopt real-time pricing intelligence tools that track market trends, competitor moves, and customer behavior as they happen. Combine this with a streamlined internal process that empowers your team to act fast when conditions change. Set pre-approved pricing triggers or guardrails so you’re not reinventing the wheel every time the market shifts. ## 3\. Fragmented Pricing Data Imagine: Your pricing data lives in spreadsheets, CRM notes, and random Slack threads. Sales can’t find the latest rates, Marketing uses outdated promo rules, and Finance isn’t even looped in. ![Fragmented Pricing Data](https://www.blog.datahut.co/content/images/2026/07/img-456.png.webp) Why This Happens: When pricing information is scattered across multiple platforms with no single source of truth, teams end up making decisions based on outdated, inconsistent, or incomplete data. This data fragmentation leads to errors, revenue leakage, and missed opportunities to optimize pricing. The Fix: Centralize all pricing-related data in a single, accessible platform whether it’s a dedicated pricing database, a unified dashboard, or integrated analytics tools. Ensure all departments have access to real-time updates and insights. Bonus: combine this with automated alerts for price changes, discount limits, or margin thresholds to keep everyone on the same page. ## 4\. Constant Second Guessing of Pricing Decisions Imagine: Your team launches a new pricing tier then spends weeks worrying if it’s too high, too low, or just plain wrong. Instead of focusing on results, everyone’s stuck in a loop of doubt and indecision. ![Constant Second Guessing of Pricing Decisions](https://www.blog.datahut.co/content/images/2026/07/img-457.png.webp) Why This Happens: Pricing uncertainty often stems from a lack of confidence in the underlying data or strategy. Without clear benchmarks, feedback loops, or historical context, teams are left guessing. This pricing anxiety creates hesitation, slows down execution, and weakens your position in the market. The Fix: Build confidence through a data-backed pricing framework. Use past performance, customer behavior insights, competitor benchmarking, and elasticity modeling to inform your pricing. Create a regular pricing review cadence—monthly or quarterly—to evaluate results and make informed adjustments. When decisions are grounded in real data, your team can stop guessing and start growing. ## 5\. Reactive Rather Than Strategic Pricing Your competitor drops prices, and your immediate response is to match or beat it—no questions asked. It becomes a cycle of chasing others rather than leading the market. ![Reactive Rather Than Strategic Pricing](https://www.blog.datahut.co/content/images/2026/07/img-458.png.webp) Why This Happens: When pricing is driven by fear of losing business instead of a clear strategy, it puts your brand in constant defense mode. This reactive pricing approach erodes profit margins, devalues your product, and signals uncertainty to both your team and your customers. The Fix: Shift from reactive to proactive by building a strategic pricing roadmap aligned with your business goals, customer segments, and value proposition. Use competitor insights as one of many inputs not your pricing bible. Define your pricing strategy around differentiation, not imitation. Include pricing goals in quarterly planning and invest in competitive analysis that goes beyond just price comparisons. ## 6\. Avoidance of Pricing Experiments Imagine: Your team avoids A/B testing prices or bundling options because “what if it backfires?” You stick to safe, familiar pricing even if it’s underperforming. ![Avoidance of Pricing Experiments](https://www.blog.datahut.co/content/images/2026/07/img-459.png.webp) Why This Happens: Fear of negative outcomes, lack of testing infrastructure, or cultural resistance to change often keep companies from experimenting. But avoiding pricing experiments leads to stagnation, you never discover what your customers are truly willing to pay or which offers convert best. The Fix: Foster a test-and-learn pricing culture. Start small with controlled experiments—like testing price sensitivity for a specific product or audience segment. Use data analytics to track outcomes and iterate fast. Over time, these experiments will uncover hidden revenue opportunities and give your team the confidence to innovate with pricing ## Conclusion: Don’t Let Pricing Fatigue Stall Your Growth Pricing isn't just about numbers, it’s about confidence, clarity, and control. If your business shows even one of these six signs of pricing fatigue from inconsistent strategies to fear of experimentation, it’s time to act. Each of these challenges reflects broader pricing strategy gaps and the longer you ignore them, the more costly they become. The right tools, processes, and mindset can help you stop making pricing mistakes and start building pricing confidence. Ignoring these red flags leads to lost revenue, misaligned teams, and missed opportunities. But the good news? Each challenge is fixable with the right mindset, tools, and strategy. - Centralize your pricing data - Align teams with a unified pricing playbook - Empower faster, data-backed decisions - Test, learn, and optimize with confidence At Datahut, we help businesses unlock smarter pricing through clean, actionable data. Whether you’re building a new pricing model or need insights to fine-tune your existing one, pricing intelligence starts with better data. ![At Datahut, we help businesses unlock smarter pricing through clean, actionable data. Whether you’re building a new pricing...](https://www.blog.datahut.co/content/images/2026/07/img-460.png.webp) Ready to take control of your pricing? Let’s talk. Contact us to see how [Datahut](https://www.datahut.co/?ref=blog.datahut.co)’s data solutions can support your next pricing move. AUTHOR I’m Aarathi J, Marketing Manager at Datahut. With over 5 years of experience in data-driven marketing, I’ve worked alongside fashion, retail, and tech brands to help them uncover hidden insights through data. At Datahut, we’ve spent over a decade helping businesses solve complex pricing, inventory, and market positioning challenges using web scraping and custom data solutions. From working with top fashion houses to fast-moving consumer brands, we’ve seen how data can fix pricing fatigue and unlock growth. If you're noticing signs of pricing fatigue in your business or are unsure how your pricing compares in a competitive market, please chat with us using the widget on the right. We’ll show you how the right data strategy can make a measurable difference. ### How to Scrape Eyewear Pricing Data from Noon Using Python URL: https://www.blog.datahut.co/post/how-to-scrape-eyewear-pricing-data-from-noon-using-python/ Last updated: 2026-07-23T07:48:32.000Z ## Introduction Envision being able to obtain the information related to the latest eyewear prices, discounts, and sellers on Noon through web scrapping. This can benefit data analysts trying to track pricing shifts, or a company trying to get ahead of other competitors by exploiting scraping Noon’s eyewear listings. It is almost like having real time access to structured information. This blog will teach you how to scrape eyeglasses products lists on Noon, the leading online marketplaces in the Middle Eastern region. This lesson will encompass all the necessary steps like retrieving product links, and the full details such as the current price, ratings, specifications, and seller information. By the end of the tutorial, you will understand how to effectively retrieve and properly store data using Playwright and BeautifulSoup for further analysis. ## What is Web Scraping? Web scraping is the extraction of information from websites using automated tools or techniques. It is frequently utilized in e-commerce, finance, research, and real estate to manage a company's prices, competitors, and other relevant information. Advanced facilities like Playwright and BeautifulSoup allow web scraping to extract structured data from complex websites, making it possible to store the information in a database for further analysis. For our Noon Eyewear project, we executed a web scraping technique in order to retrieve relevant products from [Noon.com](https://www.noon.com/uae-en/?ref=blog.datahut.co). The project involved: 1. Product Link Scraping – Using Playwright to dynamically load category pages, scroll, and extract product URLs efficiently. 2. Data Extraction – Using BeautifulSoup to parse product pages and retrieve key details such as model number, ratings, seller information, and pricing. 3. Database Management – Storing extracted links and product data in SQLite, ensuring data integrity and easy access. 4. Error Handling & Automation – Implementing retry mechanisms, managing failed URLs, and automating the entire process for continuous data collection. By automating this workflow, we successfully built a scalable and structured approach to gathering product insights, which can be useful for e-commerce analytics, pricing strategies, and business intelligence. ## Libraries and Tools Used in Noon Web Scraping In the Noon web scraping project, different Python tools and libraries were used to extract structured information from web pages that are loaded dynamically. Each of them helped in getting the data, processing it, and even storing it. For automation of browsers, there was usage of Playwright and Selenium. This allows the script to work with webpages that utilize JavaScript. Unlike conventional static web scrapers, Playwright manages to scroll automatically, wait for elements to be loaded, and manage dynamic content. During this project Playwright helped retrieve product links and details by simulating clicking and scrolling, which guaranteed that all required information was available before extraction. BeautifulSoup, which is part of the bs4 library, is used to parse the HTML content. It was able to change the captured raw webpage data into an easily parseable format so that the script is able to retrieve product links, ratings, specifications, and other elements through tag-based selectors. Extracting relevant detail from the Noon website was facilitated by BeautifulSoup within the html structure. To manage and store the scraped data in an SQLite database, the sqlite3 module was used. Tracking product links and maintaining a processed URLs flag was done effectively so that duplicate entries are prevented. In addition to that, separate tables were created to store scrapy items. ## STEP 1 : Product Link Scraping ### Importing Necessary Libraries ``` import asyncio import sqlite3 import random from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` Before we begin scraping data from the Noon website, we need to import several important libraries that will help us with different tasks in our script. ### Database Setup ``` # Database setup db_filename = "noon_eyeglasses.db" ``` To store the collected product links in an organized way, we need a database. In this script, we define a database file named "noon\_eyeglasses.db" using the variable db\_filename. This database will act as a storage system where we can save product links so that we don’t lose any data during the scraping process. Instead of keeping the links in a temporary file or a list in memory, using a database ensures that our data is safe even if the script stops running unexpectedly. A database is especially useful for avoiding duplicate entries. If we run the script multiple times, we can check whether a link already exists before adding it again. This prevents unnecessary repetitions and keeps our data clean. By storing the links in a structured way, we can easily access and use them later for further processing, such as extracting detailed product information from each link. ### Initializing the Database ``` def init_db(): """ Initialize the SQLite database and create the 'product_links' table if it does not exist. The table contains: - 'id' (INTEGER, PRIMARY KEY, AUTOINCREMENT) as a unique identifier. - 'product_link' (TEXT, UNIQUE) to store product URLs without duplicates. This function: - Connects to the SQLite database. - Executes the table creation query. - Commits the changes. - Closes the database connection. Returns: - None """ conn = sqlite3.connect(db_filename) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_links ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_link TEXT UNIQUE ) """) conn.commit() conn.close() ``` Before we start collecting product links, we need a structured way to store them. The init\_db() function is responsible for setting up the database. It ensures that our database file, "noon\_eyeglasses.db", has a table named "product\_links", where we can store the URLs of products we scrape. In this function, we first establish a connection to the SQLite database using sqlite3.connect(db\_filename). If the database file does not exist, SQLite automatically creates it. Then, we create a cursor object, which allows us to execute SQL commands. We use the CREATE TABLE IF NOT EXISTS statement to create the product\_links table. This table contains two columns: 1. id – A unique identifier for each entry, which is automatically incremented when a new link is added. 2. product\_link – A text field where each product URL is stored. The UNIQUE constraint ensures that duplicate links are not added. After executing the table creation command, we commit the changes to save them in the database and close the connection. This function ensures that the database structure is always in place before we start adding product links, preventing errors and maintaining data consistency. ### Retrieving Category URLs ``` def get_category_urls(): """ Read category URLs from a text file and return them as a list after stripping whitespace. Returns: - list[str]: A list of cleaned category URLs. The function: - Opens 'data/category_urls.txt' in read mode. - Strips leading/trailing spaces from each line. - Filters out empty lines. - Returns the cleaned URLs as a list. """ with open("data/category_urls.txt", "r") as f: return [ line.strip() for line in f if line.strip() ] ``` The get\_category\_urls() function is responsible for reading category page URLs from a text file and returning them as a list. These category URLs act as starting points for our scraping process, as each category page contains multiple product listings that we need to extract. Inside the function, we open the file "data/category\_urls.txt" in read mode. This file contains a list of category page links, each written on a new line. We then read each line, remove any extra spaces at the beginning or end using the strip() function, and ensure that empty lines are ignored. The cleaned URLs are collected into a list and returned. This approach keeps the category URLs separate from our code, making it easier to update them without modifying the script. If we need to scrape different categories, we can simply edit the "category\_urls.txt" file instead of making changes in the Python code. ### Saving Product Links to the Database ```      def save_links_to_db(cursor, links): """ Insert new product links into the 'product_links' table. Args: - cursor (sqlite3.Cursor): Database cursor to execute SQL commands. - links (list[tuple[str]]): A list of tuples, where each tuple contains a single product link. Returns: - None The function: - Uses executemany() to insert multiple links efficiently. - Ignores duplicate entries using a UNIQUE constraint. - Prints a message if a duplicate entry is encountered. """ try: cursor.executemany( "INSERT INTO product_links (product_link) VALUES (?)", links) except sqlite3.IntegrityError: print("Skipping duplicate entry...") ``` The save\_links\_to\_db() function is responsible for storing the collected product links in the database. Since we are dealing with multiple links at once, this function efficiently inserts them into the "product\_links" table using a batch operation. The function takes two inputs: cursor, which is a database cursor used to execute SQL commands, and links, which is a list of tuples, where each tuple contains a single product link. Using executemany(), the function attempts to insert all the links into the database at once. This method is faster and more efficient than inserting links one by one. Since the "product\_links" table has a UNIQUE constraint on the product\_link column, the database will not allow duplicate entries. If an attempt is made to insert a link that already exists, SQLite raises an IntegrityError. The function catches this error and prints "Skipping duplicate entry..." instead of stopping the script. This ensures that the scraper continues running smoothly without interruptions, even if some links are already present in the database. ### Extracting Product Links from the Webpage ``` def extract_product_links(soup): """ Extract product links from the page's HTML content using BeautifulSoup. Args: - soup (BeautifulSoup): Parsed HTML content of the webpage. Returns: - list[str]: A list of product URLs extracted from the page. The function: - Finds all product containers using their CSS class. - Extracts the 'href' attribute from anchor tags. - Appends the full product URL to a list. - Returns a list of extracted product links. """ product_divs = soup.find_all( "div", class_="sc-57fe1f38-0 eSrvHE" ) links = [] for div in product_divs: anchor = div.find("a", href=True) if anchor: link = "https://www.noon.com" + anchor["href"] links.append(link) return links ``` The extract\_product\_links() function is responsible for finding and extracting product links from a webpage's HTML content. Since the product links are embedded within the webpage structure, we use BeautifulSoup to process the HTML and locate the relevant information. First, the function looks for all
    elements that match a specific CSS class (sc-57fe1f38-0 eSrvHE). These
    elements contain product details, including links to individual product pages. Inside each
    , the function searches for an (anchor) tag with an href attribute, which holds the actual link to the product. Once the link is found, it is combined with "https://www.noon.com" to create a complete product URL. This is necessary because the href attribute usually contains only a partial link, and we need to add the main website domain to make it a valid URL. The extracted links are stored in a list and returned. This function ensures that we collect all product links displayed on a given category page so that they can be further processed to extract detailed product information. ### Scraping Product Links from a Category ``` async def scrape_category(page, cursor, base_url): """ Scrape product links from a given category, iterating through multiple pages. Args: - page (playwright.async_api.Page): The Playwright page instance used for browsing. - cursor (sqlite3.Cursor): Database cursor to execute SQL queries and store product links. - base_url (str): The base URL of the category page to be scraped. Returns: - None: The function saves the extracted links directly into the database. The function: - Visits each page of the category sequentially. - Waits for the page to fully load before scraping. - Scrolls down multiple times to ensure all products load. - Extracts product links using BeautifulSoup. - Checks for duplicates before saving to the database. - Stops scraping if no new links are found. """ page_number = 1 while True: url = f"{base_url}&page={page_number}" print(f"\nVisiting Page {page_number}: {url}") await page.goto( url, wait_until="load", timeout=200000 ) await page.wait_for_load_state( "domcontentloaded" ) # Scroll down to load all products for _ in range(10): await page.evaluate( "window.scrollTo(0, document.body.scrollHeight)" ) await asyncio.sleep(2) content = await page.content() soup = BeautifulSoup( content, "html.parser" ) new_links = [ (link,) for link in extract_product_links(soup) if not cursor.execute( "SELECT COUNT(*) FROM product_links WHERE product_link=?", (link,) ).fetchone()[0] ] if new_links: save_links_to_db(cursor, new_links) cursor.connection.commit() print( f"Scraped {len(new_links)} new URLs " f"from Page {page_number}" ) else: print( f"No new URLs found on Page {page_number}, " "moving to next category..." ) break # Stop if no new URLs are found page_number += 1 await asyncio.sleep( random.uniform(5, 7) # Randomized delay (5 to 7 seconds) ) ``` The purpose of the scrape\_category() function is to get product links from a certain category on the Noon website while going through several pages. Since products are listed across a variety of pages, this function makes sure that we collect links from every page possible. The function starts with a specific page\_number set to 1, meaning the first page of the category. Thereafter, a loop is created to go through each page in order. The page URL is set by adding the page parameter for the base category url. The script goes to the page with Playwright and waits for it to load. Noon, just like many other modern websites, has an infinite scrolling feature where products will load upon scrolling down. To mimic this, the function repeatedly calls window.scrollTo(0, document.body.scrollHeight). This technique ensures that all products on the page are loaded before starting the scraping. The scrolls also have a low time duration between them to simulate real browsing and prevent detection. When the page is completely loaded, product links are extracted through linking BeautifulSoup with HTML parsing. The links are only saved to the database after the system checks for duplicates. Any new links detected will be saved and the script will continue with the next page. In case of no new links being found, ### Fetching Product Links from All Categories ``` async def fetch_product_links(): """ Main function to scrape product links from all categories. This function: - Reads category URLs from a file. - Launches a Chromium browser instance using Playwright. - Sets custom HTTP headers for better request handling. - Iterates through each category URL and scrapes product links. - Stores extracted product links in an SQLite database. - Ensures proper cleanup by closing database connections and the browser after execution. Returns: - None (Results are stored in the database). """ category_urls = get_category_urls() async with async_playwright() as p: browser = await p.chromium.launch(headless=False) page = await browser.new_page() await page.set_extra_http_headers({ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Connection": "keep-alive", }) conn = sqlite3.connect(db_filename) cursor = conn.cursor() for base_url in category_urls: print(f"\nScraping category: {base_url}") await scrape_category(page, cursor, base_url) conn.close() await browser.close() print("\nFinal scraping completed! All links saved to the database.") ``` The fetch\_product\_links() function serves as the overall entry point to scrape product links within various categories. The function controls the entire process of scraping by coordinating fetching category URLs, opening the web browser, and saving the scraped links in the database. First, the function retrieves the list of category URLs by calling the get\_category\_urls() function, which reads the URLs from a file. Second, in Playwright, a session of the browser Chromium is opened, and the script can interact with the web pages as if a human were doing so. Custom HTTP headers are set to simulate a real user, e.g., setting the User-Agent, Accept-Language, and Accept-Encoding headers. These headers help prevent the scraper from being caught by the website for automated traffic. For each category URL, it calls the scrape\_category() function, which performs the scraping for each and every category. While collecting the links, they are stored in an SQLite database so that the data is properly structured and easily accessed for future use. After parsing all the categories, the process properly closes the browser and database connection for resource release. The script concludes by outputting a success message, confirming all product links successfully scraped and saved in the database. Running the Scraping Process ``` # Initialize database init_db() # Run the async function asyncio.run(fetch_product_links()) ``` The last section of the script creates the database and subsequently calls the main scraping function in an asynchronous manner. First, the init\_db() function is executed to create the SQLite database and define the required table (product\_links) to store the scraped product links. This way, the database is prepared before scraping begins. Then, asyncio.run(fetch\_product\_links()) command is invoked to execute the fetch\_product\_links() function. Because fetch\_product\_links() is an async function, we utilize asyncio to handle the running of asynchronous operations such as browsing web pages and waiting for the content to load. Invoking asyncio.run() initiates the whole process of scraping, which imports the links of products from all categories, stores them in the database, and closes resources safely once the operation is done. This organization guarantees that the script runs smoothly, manages asynchronous operations correctly, and cleans up afterwards, offering a seamless and dependable scraping experience. ## STEP 2 : Product Data Scraping From Product Links ### Importing Libraries ``` import sqlite3 import random import time from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup ``` imported necessary libraries ### Database Setup and Operations ``` # Database setup and operations def connect_db(): """ Establish a connection to the SQLite database. This function: - Connects to the 'noon_eyeglasses.db' database. - Prints a message indicating the connection status. - Returns a SQLite connection object. Returns: - sqlite3.Connection: A connection object to interact with the database. """ print("Connecting to the database...") return sqlite3.connect("noon_eyeglasses.db") ``` The first step in this code is to establish a connection with the SQLite database where all the scraped product data will be stored. The function connect\_db() connects to a database named "noon\_eyeglasses.db". When the function is called, it prints a message to indicate that the connection process is starting. Once the connection is successfully made, it returns a connection object, which is used for further interactions with the database. This connection object allows the program to execute commands such as inserting new data or querying the existing data, helping us store and manage the product information efficiently. ### Modifying the Product Links Table ``` def alter_product_links_table(): """ Modify 'product_links' table to add the 'status' column if it does not exist. This function: - Connects to the SQLite database. - Retrieves the table schema using PRAGMA. - Checks if the 'status' column is already present. - If missing, adds 'status' (INTEGER, DEFAULT 0) to track processing status. - Commits the changes and closes the connection. The 'status' column helps in managing scraping workflows by indicating whether a product link has been processed. """ print( "Checking if 'status' column exists in product_links table..." ) conn = connect_db() cursor = conn.cursor() # Check if the 'status' column exists cursor.execute("PRAGMA table_info(product_links)") columns = cursor.fetchall() # Check if the 'status' column is present in the table if any(col[1] == 'status' for col in columns): print( "Column 'status' already exists. No changes made." ) else: # If the column doesn't exist, add it print( "Altering product_links table to add 'status' column..." ) cursor.execute( "ALTER TABLE product_links ADD COLUMN status INTEGER DEFAULT 0" ) conn.commit() print("Column 'status' added successfully.") conn.close() ``` In this section, the function alter\_product\_links\_table() checks whether the 'status' column exists in the product\_links table. This 'status' column is crucial for tracking the progress of product links during the scraping process. The function starts by connecting to the SQLite database and retrieving the table schema, which contains details about the columns of the table. It then checks if the 'status' column is already present. If the column is found, the function simply prints a message stating that no changes are needed. However, if the column is missing, it proceeds to add the 'status' column with a default value of 0\. This column will later be used to mark product links as processed (status 1) or unprocessed (status 0). Once the modification is done, the function commits the changes to the database and closes the connection, ensuring that the changes are saved and the system is ready for the next steps. ### Creating the Product Data Table ``` def create_noon_product_data_table(): """ Create the 'noon_product_data' table if it does not exist. This table stores extracted product details, including: - 'url' (TEXT): Product page URL. - 'page_title' (TEXT): Title of the product page. - 'model_number' (TEXT): Product's model number. - 'rating' (TEXT): Average customer rating. - 'review_count' (TEXT): Number of customer reviews. - 'sold_by' (TEXT): Name of the seller. - 'seller_rating' (TEXT): Seller's overall rating. - 'positive_review_percentage' (TEXT): Percentage of positive reviews. - 'specifications' (TEXT): Product specifications in text. - 'sale_price' (TEXT): Discounted price of the product. - 'discount_price' (TEXT): Original price before discount. - 'savings' (TEXT): Amount saved due to discount. The function: - Connects to the database. - Creates the table with necessary fields. - Commits changes and closes the connection. """ print("Creating 'noon_product_data' table...") conn = connect_db() cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS noon_product_data ( url TEXT, page_title TEXT, model_number TEXT, rating TEXT, review_count TEXT, sold_by TEXT, seller_rating TEXT, positive_review_percentage TEXT, specifications TEXT, sale_price TEXT, discount_price TEXT, savings TEXT )""") conn.commit() conn.close() ``` This section of the code defines the create\_noon\_product\_data\_table() function, which ensures that a table named noon\_product\_data exists in the database to store all the product details that will be scraped. The table is designed to hold essential information about each product, such as the product’s URL, title, model number, customer ratings, review count, seller details, specifications, prices, and the savings from discounts. If the table doesn’t already exist, the function creates it with the necessary fields to store this data in a structured way. After setting up the table, the function commits the changes to the database and then closes the connection, ensuring that everything is saved correctly for future use. This table will be crucial in organizing the scraped data for easy retrieval and analysis later on. ### Creating the Failed URLs Table ``` def create_failed_urls_table(): """ Create the 'failed_urls' table if it does not exist. This table stores URLs that failed during scraping, along with the failure reason for debugging. Table Structure: - 'failed_url' (TEXT): The URL that could not be scraped. - 'reason' (TEXT): Description of the failure cause. The function: - Connects to the database. - Creates the table if it is not already present. - Commits the changes and closes the connection. """ print("Creating 'failed_urls' table...") conn = connect_db() cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS failed_urls ( failed_url TEXT, reason TEXT )""") conn.commit() conn.close() ``` The create\_failed\_urls\_table() function is responsible for setting up a table called failed\_urls in the database to store information about URLs that could not be scraped. Sometimes, due to network issues, website changes, or other technical difficulties, certain URLs may fail during the scraping process. This table helps keep track of those failed attempts by storing the URL along with a description of the reason for the failure. The table includes two main columns: one for the URL (failed\_url) and one for the failure reason (reason). If this table does not already exist, the function will create it, ensuring that any failed URLs are logged for future analysis or retrying. After creating the table, the function commits the changes to the database and closes the connection, ensuring that the system is updated and ready for further operations. ### Retrieving URLs to Scrape ``` def get_urls_to_scrape(): """ Retrieve product URLs that need to be scraped. This function: - Connects to the database. - Queries the 'product_links' table for URLs where the 'status' column is set to 0 (pending scraping). - Fetches and returns a list of URLs. Returns: list: A list of product URLs that need scraping. """ print("Fetching URLs to scrape with status 0...") conn = connect_db() cursor = conn.cursor() cursor.execute( "SELECT product_link FROM product_links WHERE status = 0" ) urls = cursor.fetchall() conn.close() print(f"Found {len(urls)} URLs to scrape.") return [url[0] for url in urls] ``` The get\_urls\_to\_scrape() function is designed to fetch the product URLs that need to be scraped from the database. It connects to the database and queries the product\_links table for URLs where the 'status' column is set to 0, which indicates that these URLs are pending and have not been processed yet. The function retrieves these URLs and returns them in a list format. By doing so, it ensures that only the unprocessed URLs are selected for scraping. After fetching the URLs, the function closes the database connection and prints the number of URLs found. This helps in tracking how many URLs are waiting to be scraped in the current session. ### Updating URL Status ``` def update_url_status(url, status): """ Update the scraping status of a product URL. This function: - Connects to the database. - Updates the 'status' column in 'product_links' for the given product URL. - Commits the change and closes the connection. Args: url (str): The product URL to update. status (int): The new status value. Returns: None """ print(f"Updating status of URL {url} to {status}...") conn = connect_db() cursor = conn.cursor() cursor.execute( "UPDATE product_links SET status = ? WHERE product_link = ?", (status, url) ) conn.commit() conn.close() ``` The update\_url\_status() function is responsible for updating the scraping status of a specific product URL in the database. After scraping a product page, the status of that URL needs to be updated to reflect whether the scraping was successful or not. This function connects to the database, locates the product\_links table, and updates the status column for the given URL with the new status value. The status is typically used to track the progress of scraping, where a status of 0 might indicate pending, and a status of 1 might indicate completed. Once the update is made, the function commits the changes to the database and then closes the connection. This ensures that the database always reflects the most current status of each URL. ### Saving Failed URLs ``` def save_failed_url(url, reason): """ Save a failed URL along with the failure reason. This function: - Connects to the database. - Inserts the failed URL and its reason into the 'failed_urls' table. - Commits the change and closes the connection. Args: url (str): The product URL that failed to scrape. reason (str): The reason for the failure. Returns: None """ print(f"Saving failed URL {url} with reason: {reason}...") conn = connect_db() cursor = conn.cursor() cursor.execute( "INSERT INTO failed_urls (failed_url, reason) VALUES (?, ?)", (url, reason) ) conn.commit() conn.close() ``` The save\_failed\_url() function is used to store the URLs of products that could not be scraped, along with the reason for the failure. During the scraping process, sometimes a URL might fail due to various reasons such as network issues, changes in the website structure, or invalid links. This function ensures that these failed URLs are not ignored but instead saved for later review. It connects to the database and inserts the failed URL along with the reason into the failed\_urls table. This helps in tracking and troubleshooting issues during the scraping process. Once the data is inserted, the function commits the changes to the database and closes the connection. This function helps maintain an accurate record of which URLs need to be revisited or debugged. ### Browser Setup and Scraping Functions ``` # Browser setup and scraping functions def setup_browser(): """ Launch the browser using Playwright and create a new page. This function: - Starts a Playwright session. - Launches a Chromium browser instance. - Creates a new browser page. The browser runs in non-headless mode by default. Set `headless=True` to run in the background. Returns: tuple: (playwright, browser, page) - playwright: Playwright instance. - browser: Launched browser instance. - page: New browser page. """ print("Setting up the browser...") playwright = sync_playwright().start() browser = playwright.chromium.launch(headless=False) # Set to True for headless mode page = browser.new_page() return playwright, browser, page ``` The setup\_browser() function is responsible for setting up the environment necessary for the web scraping process by launching a web browser. It uses the Playwright library to manage the browser and navigate through web pages. This function initiates a Playwright session, launches a Chromium browser, and creates a new page for interacting with websites. By default, the browser runs in a non-headless mode, meaning you can see the browser window as it performs actions. If you want the browser to run in the background without a visible window, you can change the headless parameter to True. The function then returns the Playwright instance, the launched browser, and the newly created page, which are necessary for loading and interacting with the website during scraping. ### Setting Custom HTTP Headers for Web Scraping ``` def set_headers(page, url): """ Set HTTP headers for the request using a random user agent. This function: - Reads user agents from 'data/user_agents.txt'. - Selects a random user agent for each request. - Sets headers like 'user-agent', 'referer', and 'accept'. Args: page (Page): Playwright page instance. url (str): The target URL for setting the referer. Returns: None """ print("Setting headers for the page request...") with open("data/user_agents.txt", "r") as f: user_agents = f.readlines() user_agent = random.choice(user_agents).strip() print(f"Using user-agent: {user_agent}") headers = { "user-agent": user_agent, "referer": url, "accept": "application/json, text/plain, */*" } page.set_extra_http_headers(headers) ``` The set\_headers() function is designed to configure custom HTTP headers for web scraping requests made via Playwright. It improves the scraping process by simulating real browser requests, making it harder for websites to detect bots. The function reads a list of user agents from a file (data/user\_agents.txt) and randomly selects one for each request, mimicking different users. It also sets the referer header to the provided URL and the accept header to specify the types of content the browser can accept. This makes the request appear more legitimate, helping to bypass basic anti-scraping measures. By using Playwright’s page instance, the function ensures that all requests made from the page carry these headers, reducing the chances of getting blocked. ### Visiting a Web Page with Random Delay ``` def visit_page(page, url): """ Visit the given URL with a random delay to prevent detection and throttling. This function: - Introduces a delay between 3 to 5 seconds before making the request. - Navigates to the specified URL using Playwright. - Uses an increased timeout to handle slow-loading pages. Args: page (Page): Playwright page instance. url (str): The target URL to visit. Returns: None """ print(f"Visiting URL: {url}") time.sleep(random.uniform(3,5)) # Random delay between 5 and 8 seconds page.goto(url, timeout=200000) # Increase timeout if needed print(f"Page {url} loaded successfully.") ``` The visit\_page() function is crafted to navigate to a specified URL using Playwright, while introducing a random delay between 3 to 5 seconds before making the request. This randomness helps to simulate human-like behavior, making it more difficult for the website to detect and block the scraping activity. Additionally, the function sets an increased timeout (200,000 ms) to handle slow-loading pages, ensuring that even if the page takes longer to load, the function will still wait and complete the task. This approach minimizes the chances of encountering throttling or detection, promoting smoother scraping. ### Extracting Page Title ``` def extract_title(page): """ Extract and return the page title. This function: - Retrieves the title of the current page using Playwright's `title()` method. - Prints the extracted title for debugging. Args: page (Page): Playwright page instance. Returns: str: The extracted page title. """ title = page.title() print(f"Extracted page title: {title}") return title ``` The extract\_title() function is designed to retrieve the title of the current web page using Playwright's title() method. Once the title is extracted, it is printed for debugging purposes, helping to confirm that the correct page has been loaded. This function returns the page title as a string, which can be useful for various purposes, such as verifying that the right page has been scraped or for storing the title along with other product details in the database. ### Extracting Model Number ``` def get_model_number(soup): """ Extract the model number from the HTML content. This function: - Searches for a `
    ` element with class 'modelNumber' using BeautifulSoup. - Extracts and cleans the text content. - Splits the text by " : " to isolate the model number. - Returns the model number or a default message if not found. Args: soup (BeautifulSoup): Parsed HTML content. Returns: str: Extracted model number or a 'not found' message. """ model_number = soup.find( "div", class_="modelNumber" ) if model_number: model_text = model_number.text.strip() model_num = model_text.split(" : ")[-1] else: model_num = "Model number not found." print(f"Extracted model number: {model_num}") return model_num ``` The get\_model\_number() function is designed to extract the model number from the HTML content of a product page. It searches for a
    element with the class modelNumber using BeautifulSoup. If the element is found, it extracts the text, cleans it by removing any surrounding whitespace, and splits the text by " : " to isolate the model number. If the model number is not found, the function returns a default message, "Model number not found." The function prints the extracted model number and returns it as a string. This function is helpful for scraping product-specific information from e-commerce websites where model numbers are critical for identifying products. ### Extracting Product Rating and Review Count ``` def get_review_and_rating(soup): """ Extract the product rating and number of reviews from the HTML content. This function: - Searches for a `
    ` with class 'sc-9cb63f72-2 dGLdNc' to extract the rating. - Searches for a `` with class 'sc-9cb63f72-5 DkxLK' to extract the review count. - Cleans and returns the extracted values. - If elements are missing, returns a default message. Args: soup (BeautifulSoup): Parsed HTML content. Returns: tuple: (rating_value, review_count) - rating_value (str): Extracted rating or 'Rating not found.' - review_count (str): Extracted number of reviews or 'Review count not found.' """ rating = soup.find( "div", class_="sc-9cb63f72-2 dGLdNc" ) reviews = soup.find( "span", class_="sc-9cb63f72-5 DkxLK" ) if rating: rating_value = rating.text.strip() else: rating_value = "Rating not found." if reviews: review_count = reviews.text.strip() else: review_count = "Review count not found." print( f"Extracted rating: {rating_value}, " f"Review count: {review_count}" ) return rating_value, review_count ``` The get\_review\_and\_rating() function is designed to extract both the product's rating and the number of reviews from the HTML content. It searches for a
    element with the class sc-9cb63f72-2 dGLdNc to extract the product rating, and a element with the class sc-9cb63f72-5 DkxLK to extract the review count. After extracting these values, the function cleans the text to remove any unwanted spaces. If the elements are not found, the function returns default messages: "Rating not found" for the rating and "Review count not found" for the reviews. The function prints the extracted rating and review count and returns them as a tuple. This function is essential for scraping product feedback data from e-commerce websites. ### Extracting Seller Information ``` def get_sold_by(soup): """ Extract the seller information from the HTML content. This function: - Searches for a `` with class 'allOffers' to extract the seller name. - Cleans and returns the extracted text. - Returns a default message if the element is missing. Args: soup (BeautifulSoup): Parsed HTML content. Returns: str: The seller name or 'Seller information not found.' """ sold_by = soup.find( "span", class_="allOffers" ) if sold_by: sold_by_text = sold_by.text.strip() else: sold_by_text = "Seller information not found." print(f"Extracted sold by: {sold_by_text}") return sold_by_text ``` The get\_sold\_by() function is responsible for extracting the seller's information from the HTML content. It searches for a element with the class allOffers to retrieve the seller's name. If the element is found, the function cleans the extracted text by stripping any unwanted spaces. If the element is not present, the function returns a default message, "Seller information not found." The function prints the extracted seller name and returns it as a string, which helps in identifying the vendor or seller associated with the product. ### Extracting Seller Rating and Positive Review Percentage ``` def scrape_seller_details(soup): """ Extract the seller's rating and positive review percentage. This function: - Locates the seller rating using CSS selectors. - Extracts the percentage of positive reviews. - Returns 'N/A' if the data is not found. Args: soup (BeautifulSoup): Parsed HTML content. Returns: tuple: (seller_rating, positive_review_percentage) as strings. """ seller_rating_tag = soup.select_one( "div.sc-fb51bf29-0 span.sc-fb51bf29-1" ) if seller_rating_tag: seller_rating = seller_rating_tag.text.strip() else: seller_rating = "N/A" positive_rating_tag = soup.select_one( "div.sc-cf1d50e0-4 span" ) if positive_rating_tag: positive_rating = positive_rating_tag.text.strip() else: positive_rating = "N/A" print( f"Extracted seller rating: {seller_rating}, " f"Positive review percentage: {positive_rating}" ) return seller_rating, positive_rating ``` The scrape\_seller\_details() function is designed to extract the seller's rating and positive review percentage from the HTML content. It uses CSS selectors to locate the relevant elements. The function first searches for the seller rating in the div.sc-fb51bf29-0 span.sc-fb51bf29-1 element and extracts the text, returning "N/A" if not found. Similarly, it looks for the positive review percentage in div.sc-cf1d50e0-4 span, and if this is missing, it also returns "N/A". The function prints the extracted values and returns them as a tuple (seller\_rating, positive\_review\_percentage). This data is useful for evaluating the seller's reputation on the platform. Extracting Product Specifications ``` def scrape_specifications(soup): """ Extract product specifications from the table and return them as a dictionary. This function: - Searches for the first element in the HTML. - Iterates through table rows () and extracts key-value pairs from
    elements. - Stores specifications in a dictionary. - Returns an empty dictionary if no table is found. Args: soup (BeautifulSoup): Parsed HTML content. Returns: dict: A dictionary containing product specifications. """ specs = {} table = soup.find("table") if table: rows = table.find_all("tr") for row in rows: cols = row.find_all("td") if len(cols) == 2: key = cols[0].text.strip() value = cols[1].text.strip() specs[key] = value print( f"Extracted specifications: {specs}" ) return specs ``` The scrape\_specifications() function is designed to extract product specifications from a product page's HTML content. It first searches for the first element, which typically contains the specifications. Then, the function iterates through each row () of the table, extracting key-value pairs from the
    elements. These pairs are stored in a dictionary where the key is the specification name (e.g., "Model", "Color") and the value is the specification detail (e.g., "XYZ123", "Red"). If no table is found, the function returns an empty dictionary. The function prints the extracted specifications and returns them in the form of a dictionary, providing structured data for further processing. ### Scraping Price Details ``` def scrape_price_details(soup): """ Extract price details from the product page. This function: - Searches for elements containing sale price, discount price, and savings. - Extracts and cleans the text from the corresponding
    tags. - Returns default messages if any element is missing. Args: soup (BeautifulSoup): Parsed HTML content. Returns: tuple: (sale_price, discount_price, savings), where each is a string. """ sale_price_tag = soup.find( "div", class_="priceNow" ) discount_price_tag = soup.find( "div", class_="priceWas" ) saving_tag = soup.find( "div", class_="priceSaving" ) if sale_price_tag: sale_price = sale_price_tag.text.strip() else: sale_price = "Sale price not found." if discount_price_tag: discount_price = discount_price_tag.text.strip() else: discount_price = "Discount price not found." if saving_tag: saving = saving_tag.text.strip() else: saving = "Saving details not found." print( f"Extracted prices:\n" f" Sale price: {sale_price}\n" f" Discount price: {discount_price}\n" f" Savings: {saving}" ) return sale_price, discount_price, saving ``` The scrape\_price\_details() function extracts key pricing information from a product page. It looks for three primary elements on the page: the sale price (priceNow class), the discount price (priceWas class), and the savings (priceSaving class). The function attempts to locate these elements, cleans the extracted text, and returns it as a tuple. If any of the elements are missing, it returns a default message indicating that the price information was not found. This function prints the extracted details for debugging purposes and returns the data in a structured format (tuples of strings) for further processing or storage. ### Closing the Browser ``` def close_browser(playwright, browser): """ Main function to scrape product data and save it to the database. This function: - Launches the browser and sets request headers. - Navigates to the product page and extracts relevant details. - Parses the page content using BeautifulSoup. - Extracts key product information including: - Page title - Model number - Rating and reviews - Seller details - Specifications - Pricing information (sale price, discount price, savings) - Saves extracted data into the `noon_product_data` table in the database. - Updates the URL status in the `product_links` table upon successful scraping. - Handles errors by logging failed URLs in the `failed_urls` table. - Closes the browser and database connection after processing. Args: url (str): The product URL to be scraped. Returns: None """ print("Closing the browser...") browser.close() playwright.stop() ``` The close\_browser() function is responsible for safely closing the browser and stopping the Playwright session once the scraping process is complete. It ensures that the resources used by the browser and Playwright instance are properly released. This function accepts the playwright and browser instances as arguments, calls the close() method on the browser, and then stops the Playwright session with the stop() method. By calling this function at the end of the scraping task, it ensures a clean shutdown of the browser environment, preventing memory leaks and leaving the scraping environment ready for the next task. ### Scraping and Saving Product Data ``` # Scraping and data insertion def scrape_product_data(url): """ Scrapes product data from a given URL and inserts it into the database. This function: - Initializes a Playwright browser session and sets headers. - Extracts product details such as title, model number, ratings, reviews, seller details, specifications, and pricing information. - Saves the extracted data into the `noon_product_data` table. - Updates the `product_links` table to mark the URL as scraped. - Handles errors by saving failed URLs and closing the browser session properly. Args: url (str): The product page URL to scrape. Returns: None """ try: print( f"Starting to scrape product data for URL: {url}" ) playwright, browser, page = setup_browser() set_headers(page, url) visit_page(page, url) title = extract_title(page) # Get the page content and parse it using BeautifulSoup html_content = page.content() soup = BeautifulSoup( html_content, "html.parser" ) model_number = get_model_number(soup) rating, reviews = get_review_and_rating(soup) seller = get_sold_by(soup) seller_rating, positive_rating = scrape_seller_details(soup) specifications = scrape_specifications(soup) sale_price, discount_price, savings = scrape_price_details(soup) # Save the data to noon_product_data table conn = connect_db() cursor = conn.cursor() cursor.execute( """ INSERT INTO noon_product_data ( url, page_title, model_number, rating, review_count, sold_by, seller_rating, positive_review_percentage, specifications, sale_price, discount_price, savings ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) """, ( url, title, model_number, rating, reviews, seller, seller_rating, positive_rating, str(specifications), sale_price, discount_price, savings ) ) conn.commit() conn.close() # Update the status of the URL in product_links update_url_status(url, 1) close_browser(playwright, browser) except Exception as e: print( f"Error processing {url}: {e}" ) save_failed_url(url, str(e)) close_browser(playwright, browser) ``` The scrape\_product\_data() function is responsible for scraping product details from a given URL and saving the extracted data into the database. It starts by initializing a Playwright browser session and setting up the necessary headers before visiting the product URL. The function then extracts key product information, including the page title, which is retrieved using the extract\_title() function, as well as the model number, product rating, review count, seller information, seller rating, positive review percentage, product specifications, and pricing details such as the sale price, discount price, and savings. Once the data is extracted, it is inserted into the noon\_product\_data table in the database. After successful scraping, the status of the URL in the product\_links table is updated to reflect that the URL has been processed. If any error occurs during the scraping process, the failed URL and the error message are saved in the failed\_urls table, and the browser session is properly closed. This approach ensures a systematic handling of each product URL, with efficient error management and resource handling. ### Scraping All URLs ``` def scrape_all_urls(): """ Main function to scrape all product URLs. This function: - Retrieves all product URLs with a status of 0 from the `product_links` table. - Iterates through each URL and calls `scrape_product_data(url)`. - Ensures all URLs are processed sequentially. Args: None Returns: None """ print("Starting the scraping process for all URLs...") urls = get_urls_to_scrape() for url in urls: scrape_product_data(url) ``` The scrape\_all\_urls() function is designed to scrape all product URLs that are marked with a status of 0 (indicating they need to be processed) from the product\_links table in the database. The function retrieves these URLs and iterates through them, calling the scrape\_product\_data(url) function for each URL to extract and save the product data. This process ensures that all pending URLs are processed sequentially, one by one. The function manages the flow of the scraping process and ensures each URL is scraped and updated accordingly, providing a streamlined approach for handling multiple URLs. ### Main Entry Point for the Scraping Script ``` if __name__ == "__main__": """ Entry point for the scraping script. This script: - Ensures necessary database tables (`product_links`, `noon_product_data`, `failed_urls`) exist. - Calls `scrape_all_urls()` to start the scraping process. Execution: Run this script to initiate web scraping for all available product URLs. Args: None Returns: None """ alter_product_links_table() create_noon_product_data_table() create_failed_urls_table() scrape_all_urls() ``` The script's entry point is defined in the if name == "\_\_main\_\_": block, ensuring that the necessary database tables (product\_links, noon\_product\_data, and failed\_urls) are created or altered if they don't already exist. It then triggers the scrape\_all\_urls() function, which starts the scraping process for all URLs that need to be processed. By running this script, you initiate the web scraping process, which involves retrieving product data, inserting it into the database, and handling failed scraping attempts. The script ensures everything is set up and runs smoothly, starting from the table creation to the execution of scraping tasks. ## Conclusion The Noon Eyewear web scraping project is a highly advanced data gathering system coded in Python that applies a two-stage strategy for complete product data extraction. The initial stage utilizes an asynchronous framework with Playwright and BeautifulSoup to gather product URLs from category pages in an efficient manner, storing them in a SQLite database while managing pagination and applying dynamic content loading through smart scroll simulation. The second stage is dedicated to extracting detailed product information, running in parallel to provide accurate data scraping of product details, prices, seller names, and customer reviews. The system has built-in strong error handling features, such as random rotation of user agents, configurable timeouts, and extensive failure logging. Data integrity is ensured through a well-organized SQLite database with several tables monitoring scraping status and failures. The codebase exhibits careful consideration of performance optimization via asynchronous operations, database efficiency, and prudent request handling with random waits to evade detection. Some features that stand out include modular function design for ease of maintenance, concise logging throughout the process, and scalability through separate phases of scraping and status tracking for resume ability. Though the system entails cautious resource management by virtue of browser automation and perhaps the need to adjust for rate limiting in accordance with site policies. It nevertheless offers a sound basis for collecting detailed eyewear product information from Noon.com in an assured manner while ensuring data quality and process dependability. Connect with[ Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. ### How Argos Uses Pricing and Product Data to Sell URL: https://www.blog.datahut.co/post/how-argos-sells-watches-what-the-data-reveals/ Last updated: 2026-07-23T07:48:32.000Z Watches are more than just timekeepers, they're fashion statements, fitness companions, and even safety tools. But what drives pricing, features, and customer satisfaction in today’s competitive watch market? To explore this, we scraped and analyzed watch listings on Argos, one of the UK’s top multi-category retailers. Through exploratory data analysis (EDA), we uncovered trends across pricing, features, demographics, and customer reviews. Curious how this kind of analysis could help you understand your market better? Datahut specializes in custom web scraping and analytics solutions. ## Understanding the Argos Watch Landscape Argos offers a wide selection of watches under its Jewellery & Watches section, targeting men, women, and children alike. The range caters to diverse preferences and budgets from smartwatches to classic analog styles. ## Price Distribution by Category Using violin plots, we explored how prices vary across men's, women's, and children's watches. ![Pice distribution](https://www.blog.datahut.co/content/images/2026/07/img-364.png.webp) - Men’s Watches: Broad range, with several high-priced luxury models. - Women’s Watches: Clustered around mid-to-high price points—style meets affordability. - Children’s Watches: Primarily low-cost, simple designs focused on accessibility and basic functionality. Takeaway: Argos segments its inventory to cater to different customer personas, from style-conscious adults to tech-savvy athletes and budget-focused parents. ## What Features Matter to Whom? Smartwatches today do more than tell time, they track fitness, offer communication tools, and even support safety features. Here’s how Argos tailors functionality based on user groups: ![smart features distribution](https://www.blog.datahut.co/content/images/2026/07/img-365.png.webp) ### Men’s Smartwatches: Performance & Precision - Prioritized features: chronographs, GPS tracking, heart rate monitors, and stopwatches. - Designed for fitness enthusiasts, cyclists, and outdoor adventurers. - Communication features are available (e.g., making/answering calls), but not the main focus. - Emphasizes rugged design and technical accuracy over social connectivity. - Growth opportunity: Improve social and communication features to broaden appeal. ### Women’s Smartwatches: Style, Social, and Wellness - Blend connectivity (e.g., weather updates, notifications) with health monitoring. - Features include calorie tracking, heart rate monitoring, and aesthetic appeal. - Designed for everyday use, balancing fashion, fitness, and functionality. - Popular for users seeking seamless integration of wellness and social life. ### Children’s Smartwatches: Simplicity & Safety First - Focus on basic functions: pedometers, alarms, distance tracking. - Emphasize affordability, ease-of-use, and parental safety tools like GPS and calling. - Ideal for younger users new to smart devices—simple, fun, and secure. ### Key Feature Insights - Heart rate monitors, call features, and weather updates are heavily marketed toward women. - GPS, pedometers, and chronographs dominate men’s watch offerings. - Children’s watches focus on essential features, often lacking advanced capabilities like call or heart-rate functionality. - Adult smartwatches (men’s & women’s) are increasingly positioned as multi-purpose communication and health devices. - Each demographic reflects distinct priorities- performance, connectivity, or safety, guiding how brands design and price their offerings. ## Argos' Discount Strategy ![argos discount strategy](https://www.blog.datahut.co/content/images/2026/07/img-366.png.webp) - 93.6% of watches are sold at regular price. - Only 6.4% of products are discounted. Interpretation: Argos likely relies on value-based pricing rather than aggressive discounting, preserving perceived brand and product value. ## Brand Positioning: High-End vs. Budget ![brand positioning](https://www.blog.datahut.co/content/images/2026/07/img-367.png.webp) #### High-Priced Brands: - Samsung, Garmin, Fossil – average sale prices between £170 and £429. - Known for smart features and/or premium design. Low-Priced Brands: - Nickelodeon, Spy X, Citron – average sale prices between £7.5 and £11. - Target children's and low-budget markets. Insight: Argos captures both premium and budget segments, indicating a broad customer targeting strategy. ## What Do Customers Think? ![argos ratings](https://www.blog.datahut.co/content/images/2026/07/img-368.png.webp) Boxplot + Summary Statistics - Median rating: 4.5+ - Average rating: 4.47 - Majority of reviews fall between 4–5 stars - A few outliers show dissatisfaction ![average ratings](https://www.blog.datahut.co/content/images/2026/07/img-369.png.webp) Top-Rated Brands: Seksy, Guess, Jacques Du Manoir, Gabby’s Dollhouse (5.0 ratings) Conclusion: Overall high satisfaction, with a few exceptions, signaling strong product-market fit in most cases. ## Product Variety & Brand Dominance ![top brands by number of products ](https://www.blog.datahut.co/content/images/2026/07/img-370.png.webp) - Sekonda: 101 products (dominates Argos’ catalog) - Casio: 70 products - Citizen, Reflex Active, Lorus: offer mid-tier options Observation: Argos partners with brands across the spectrum—from entry-level to premium, ensuring diverse inventory coverage. ## Functional Attributes and Their Impact on Pricing One of the most revealing aspects of this analysis is how specific functional attributes like water resistance, bezel type, power source, and strap material, correlate with watch pricing. These features not only affect product value but also speak volumes about the intended use case and target customer. ### Water Resistance: The Deeper the Watch, the Higher the Price There’s a clear trend: as water resistance increases, so does the average price. Basic splash-proof watches are the most affordable, ideal for light use like rain or handwashing. As we move up the scale to 50m, 100m, and beyond, watches become more suitable for swimming, snorkeling, and professional diving—demanding stronger materials and advanced engineering. At the top end, diving-standard watches command premium prices and reflect serious performance. ![sales price by water resistance ](https://www.blog.datahut.co/content/images/2026/07/img-371.png.webp) ### Bezel Type: More Than Just a Design Element Bezel type also plays a major role in pricing. Watches with rotating bezels—commonly found in diving and tactical watches—have the highest average prices. These bezels serve functional purposes like tracking elapsed time. In contrast, fixed bezels are generally aesthetic and fall into the mid-range, while watches with no bezel tend to be minimalist, entry-level products with lower price points. ![avg sale price vs bezel type ](https://www.blog.datahut.co/content/images/2026/07/img-372.png.webp) ### Power Type: What’s Inside Matters The mechanism powering a watch significantly affects its cost. Mechanical watches are the most expensive, often handcrafted and appreciated for their complexity and luxury appeal. Kinetic and solar-powered watches follow, combining advanced features with self-charging capability, perfect for tech-forward or eco-conscious users. Battery-powered watches are the most affordable, dominating the mass-market and children’s segments where affordability and simplicity matter most. ![power type](https://www.blog.datahut.co/content/images/2026/07/img-373.png.webp) ### Strap Material: Comfort, Style, and Price Strap materials aren’t just about aesthetics—they’re closely tied to pricing and brand positioning. Titanium and gold-plated straps top the pricing chart, associated with durability, prestige, and luxury. Stainless steel offers a premium feel without going ultra-luxury. Meanwhile, leather and polyurethane provide a mid-range option balancing comfort and cost. Plastic, rubber, and faux leather straps are the most affordable, often used in kids’ watches or sporty, casual designs. ![strap type](https://www.blog.datahut.co/content/images/2026/07/img-374.png.webp) ### The Bigger Picture for Retailers and Brands These patterns show that price is rarely arbitrary, it’s shaped by the functional value, design complexity, and perceived lifestyle use of each watch. For product strategists and category managers, these insights can guide: - Smarter pricing models based on feature sets - Clearer product tiering for budget, mid-range, and premium shoppers - Better alignment between product design and customer needs In short, understanding the “why” behind pricing can help retailers like Argos build more competitive and consumer-friendly assortments. ## Conclusion: What This Means for Retailers & Brands EDA reveals how data can tell a compelling story. From feature preferences to brand segmentation and customer sentiment, the insights in Argos’ watch inventory can inform smarter decisions in: - Pricing strategy - Product development - Inventory planning - Market targeting ### Want Insights Like These? [Datahut ](https://www.datahut.co/?ref=blog.datahut.co)can help you unlock hidden trends in your online catalog or even your competitor’s. Contact us for custom web scraping and data analysis services tailored to your industry. Related posts FAQ SECTION ### FAQs 1\. What insights does the data reveal about Argos’s watch sales? The data reveals that Argos’s watch sales are driven by mid-range pricing, popular fashion brands, and seasonal promotions that align with gift-buying trends like holidays and Valentine’s Day. 2\. Which brands dominate Argos’s watch category? Brands such as Casio, Sekonda, and Fossil are among the top sellers on Argos, reflecting strong consumer preference for affordable and stylish options. 3\. How does Argos’s pricing strategy affect its sales performance? Argos adopts a competitive pricing strategy by balancing affordability with brand variety. Frequent discounts and bundle offers help increase sales volume while appealing to price-sensitive buyers. 4\. What types of watches perform best on Argos? Analog and digital watches with mid-tier price points perform the best, with men’s and unisex models attracting higher purchase rates than luxury or niche categories. 5\. How can businesses use Argos watch data for market insights? By analyzing Argos’s product, pricing, and discount patterns, businesses can benchmark their pricing strategies, identify top-selling models, and forecast consumer trends in the watch segment. ### Automate Ray-Ban Product Discovery Using Web Scraping URL: https://www.blog.datahut.co/post/how-to-automate-ray-ban-product-discovery-a-web-scraping-approach/ Last updated: 2026-07-23T07:48:32.000Z ## Introduction Ever been curious about how programmers extract large blocks of information from websites without manually copying and pasting? The answer lies in web scraping, a technique of web data extraction automation. Instead of manually scraping product descriptions or prices, programmers program to issue HTTP requests, extract HTML data, and parse it to get particular data.With the aid of special libraries and software, web scraping can navigate web pages, handle HTTP responses, and parse HTML or XML, making it a viable method of harvesting structured data. Be it for tracking price movements, extracting product information, or scraping contact details, the process cuts down the time by hundreds of hours if done manually.However, web scraping is not all about programming—it is also about knowing a website's structure, its robots.txt file, and its terms of service in order to scrape both legally and ethically. Ray-Ban is an iconic eye-wear brand that sells timeless styles of high-quality sunglasses and optical frames. Founded in 1936, the company first designed anti-glare glasses for U.S. Air Force pilots. Over the years, Ray-Ban has gained the status of a cultural phenomenon: wayfarers and aviators have become legendary, while classic models have gained legendary status through their sheer ubiquity. Presently, Ray-Ban is one of the international brands that offer a variety of sunglasses, eyeglasses, and prescription lenses for men and women, as well as children. Styles range from classic to contemporary, thereby satisfying consumers worldwide. Scrape all data from the websites of Ray-Ban for information on what they are offering, pricing strategies, and how they position themselves in the market. The scraping process is divided into two main phases, each implemented in a separate Python script: - URL Collection: Ray-Ban's categories pages, for example sunglasses or eyeglasses are accessed by the first script and the product URLs are collected. It makes use of Playwright to automate the browser interaction. It handles pop-ups and scrolls through a page to make sure that it has fetched all the products. Collected URLs with their category and target gender are persisted in a SQLite database. - Product Data Extraction The second script retrieves all the stored URLs from the database and then visits each product page to extract more elaborate information. This script makes use of browser automation by Playwright and Beautiful Soup for parsing HTML content. The script extracts data such as product name, collection, color options, model code, frame description, pricing, and discounts. All the detailed product information is saved back into the SQLite database. ## Tools and Technologies This Ray-Ban web scraping project relies on three key technologies: Playwright, Beautiful Soup, and SQLite. Each one plays an integral role in the data collection and storage process. Playwright is a powerful web automation library that allows the scraper to control web browsers programmatically. It's like an invisible hand opening up a web browser, typing in URLs, clicking on buttons, and scrolling through pages-all much faster than any human mind can do. In this project, Playwright is essential in crossing the Ray-Ban website. Handling dynamic content and interacting with elements such as popups or "Load More" buttons is also important for the project. It could even simulate different devices or user agents to avoid detection as a bot. Where the ability to wait for certain elements to load is really handy is in dealing with dynamic product listings like Ray-Ban's. Once the Playwright has loaded the web page, Beautiful Soup goes into work. Beautiful Soup is a Python library specifically for parsing HTML content. If you think of a webpage as a tree-like structure, Beautiful Soup helps you climb this tree to find exactly the information you need. In the context of this Ray-Ban scraper, Beautiful Soup is used to extract specific data elements like product names, prices, or color options from the HTML of each product page. It's like having a smart assistant that can quickly read through a complex document and highlight all the important information for you. The third key technology in this project is SQLite. SQLite is a lightweight, file-based database system. Think of it as a high-class filing cabinet built into your programme. In this Ray-Ban scraper, SQLite has two very important purposes. Firstly, SQLite saves the URLs of all products for which initial scanning information was found. That allows the scraper to remember which pages it needs to visit again to gather detailed information, even if the program is stopped and restarted. SQLite captures all the data extracted from each page in terms of the detailed information of the product. This builds up a structured collection of queryable data for any of the Ray-Ban products under consideration. What's beautiful about SQLite is that it's easy to implement and use, but quite powerful to manage the large amount of data created with web scraping. These three tools together comprise the backbone of the Ray-Ban scraping project: Playwright takes care of interacting with the web; Beautiful Soup extracts the relevant data; and SQLite is robust storage for storing the data. These three combined will allow for efficient, structured, and reliable gathering of data from the Ray-Ban website, which can then be stored as a valuable dataset for further analysis or usage. ## Data Cleaning and Refinement Sometimes, after extracting raw data by web scraping, it must be cleaned and purified. It is necessary to ensure the reliability and consistency of the collected data to analyze it or use it in another form. Usually, a powerful tool like OpenRefine may be used for more manual working with messy data or pandas can be utilized in Python when working in more of a programmatic way. The tasks at this stage encompass standardizing formats, missing values, removal of duplicates, and keeping track of the data obtained for consistency. These data can be used to do meaningful market analysis, compare products, or integrate them into other systems. By combining effective web scraping strategies with careful data cleaning, this would then result in an all-inclusive and accurate dataset of Ray-Ban products, meaning insight for the company into its eyewear offerings and pricing. ## Urls Collection This script in Python, then, would collect product information systematically from the website Ray-Ban, much like a digital shopping assistant that would visit various sections of Ray-Ban's online store. The script will use Playwright-the tool that supports and automates browser-to navigate through various product categories of sunglasses, eyeglasses, and smart glasses designed for the different demographics like men and women, and children, etc. It will behave like a normal visitor to the website by choosing random user agents and browsing naturally. For each category it navigates through, it will handle popup windows, scroll down the pages, click "Load More" buttons for revealing more products, and extract URLs of products. This information is then collated together and compiled in a SQLite database. Upon completion, it offers a comprehensive catalogue of Ray-Ban that can be easily accessed and analyzed later. The script is developed with sensitivity to website behavior, including error handling for common issues such as timeout, and even adding in delay between actions to avoid saturating the server. ### Import Section ``` import sqlite3 import asyncio from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError from bs4 import BeautifulSoup import random import sys ``` The import section will bring in all the tools necessary for this web scraping project. SQLite3 is used for database operations to store the scraped data. Asyncio enables asynchronous programming that makes the scraping process more efficient. Playwright is the main web automation tool that controls the browser, while BeautifulSoup helps parse the HTML content. Random is used to add delays and select user agents, and sys takes care of system level operations such as checking for files and programmable exits. ### load\_user\_agents(file\_path='user\_agents.txt') ``` def load_user_agents(file_path='user_agents.txt'): """ Loads user agents from a text file. Args: file_path (str): Path to the text file containing user agents. Defaults to 'user_agents.txt'. Returns: list: A list of user agent strings. Raises: SystemExit: If the file is not found or is empty. """ try: with open(file_path, 'r') as file: user_agents = [line.strip() for line in file if line.strip()] if not user_agents: print(f"Error: The file {file_path} is empty.") sys.exit(1) return user_agents except FileNotFoundError: print(f"Error: User agent file not found: {file_path}") sys.exit(1) ``` This function opens and reads a text file full of different browser user agents, which are text strings telling websites what kind of browser and system is trying to access it. Imagine it like having a collection of different disguises for your scraper. The function uses a simple file reading operation to get these user agents and puts them into a list that the scraper can use later. If something is going wrong - for instance, if your file is nowhere to be found or is empty - the function will notify you and halt the script. If you weren't using user agents and didn't have this early failure point, websites could simply identify that you were running a scraper rather than a normal browser. The whole function is actually built with the intention of failing early, which is a better thing than to start running into issues during actual scraping work. ### initialize\_database(db\_name='rayban\_products.db') ``` def initialize_database(db_name='rayban_products.db'): """ Initializes the SQLite database by creating a connection and setting up the table structure. The function connects to the SQLite database with the name provided in the `db_name` argument. If the database does not exist, it will be created. The function then creates a table named 'products' with columns for storing product URLs, categories, and gender information. If the table already exists, it will not be recreated, ensuring that existing data is preserved. Args: db_name (str): The name of the SQLite database file. Defaults to 'rayban_products.db'. """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS urls ( url TEXT NOT NULL, category TEXT NOT NULL, gender TEXT NOT NULL ) ''') conn.commit() conn.close() ``` Consider this function as an attempt to set up a filing cabinet for all that information you're going to collect regarding products. This function creates a SQLite database file (or simply opens it if it already exists) and ensures there's a table within it with the proper structure to store product URLs, categories, and gender information. The table is simply called 'urls' and functions very similarly to a spreadsheet with three columns: the product webpage address, what kind of product it is (sunglasses or eyeglasses), and who it's made for - men, women, kids, etc. The nice thing about this function is that it uses "IF NOT EXISTS" within its SQL command, so if you run the script multiple times, it won't accidentally erase all your existing data. Kind of like making sure you don't have a filing cabinet already before building another one. Once everything is in place, it closes the database connection, kind of like locking up the filing cabinet when you are done. ### save\_to\_database(url,category,gender,db\_name='rayban\_products.db') ``` def save_to_database(url, category, gender, db_name='rayban_products.db'): """ Saves the product URL, category, and gender information to the SQLite database. The function establishes a connection to the SQLite database specified by the `db_name` argument It then inserts a new record into the 'products' table, storing the provided `url`, `category`,and `gender`. The connection to the database is closed after the operation to ensure that resources are properly released. Args: url (str): The URL of the product to be saved. category (str): The category of the product (e.g., 'sunglasses', 'eyeglasses'). gender (str): The target gender for the product (e.g., 'Men', 'Women'). db_name (str): The name of the SQLite database file. Defaults to 'rayban_products.db'. """ conn = sqlite3.connect(db_name) cursor = conn.cursor() cursor.execute(''' INSERT INTO urls (url, category, gender) VALUES (?, ?, ?) ''', (url, category, gender)) conn.commit() conn.close() ``` This function actually goes ahead and saves this information in your database. Every time the scraper encounters a product page, this function opens up the database connection to the filing cabinet, creates a new record with the product's URL, what category it belongs to, and who it's designed for, and then safely closes the connection again. Sort of like having a secretary who knows exactly how to file each piece of information into its right place. The function is minimalist and task-oriented-it does one thing: save data, and it does that well. It uses the SQL INSERT command, appropriate changes being committed before closing it in itself, then closes the database to be sure everything has been cleaned up and in order. ### close\_popup(page) ``` async def close_popup(page): """ Attempts to close any popup that appears when first accessing the site. Args: page: Playwright page object """ try: # Wait a moment for any popups to appear await page.wait_for_timeout(2000) # Try multiple possible selectors for the close button close_button_selectors = [ 'button.close-button', # Common class name for close buttons 'button[aria-label="Close"]', # Accessibility label '.modal-close', # Common modal close class '.popup-close', # Common popup close class 'button.dismiss-button', # Common dismiss button class '[data-testid="close-button"]', # Test ID '//button[contains(@class, "close")]' # XPath for buttons containing 'close' in class ] for selector in close_button_selectors: try: # Try to find and click the close button if selector.startswith('//'): # Handle XPath selector await page.wait_for_selector(selector, timeout=2000, state='visible') await page.click(selector) else: # Handle CSS selector await page.wait_for_selector(selector, timeout=2000, state='visible') await page.click(selector) print("Successfully closed popup") return except: continue # If no close button found, try pressing Escape key await page.keyboard.press('Escape') print("Attempted to close popup with Escape key") except Exception as e: print(f"Error handling popup: {e}") ``` This one is your popup-fighting hero. Websites, mainly e-commerce sites, will often show a popup asking you to subscribe to their newsletters or provide some discount for you. Scraper time can be seriously messed with, so this function is used to close these, trying the various ways to dismiss unwanted interruptions. It is like a security guard with all the variations of how politely an unwanted interruption can be dismissed. The function waits a short amount of time to allow the pop-ups to occur, then attempts several different methods to locate and click on close buttons in order to find close, dismiss, or X buttons. If the function can't find any buttons to click, it then defaults to pressing the Escape key. It is a particularly clever function because it doesn't give up if one method fails, but rather keeps trying different approaches until it either succeeds or has tried all of the possible methods. ### scrape\_rayban\_product\_urls(category\_url, category, gender) ``` async def scrape_rayban_product_urls(category_url, category, gender): """ Asynchronously scrapes product URLs from the Ray-Ban website for a specific category and gender. This function performs several key operations: 1. Launches a Chromium browser instance with a random user agent 2. Navigates to the specified category URL 3. Handles any popup dialogs that appear 4. Scrolls through the page to load all products (handles lazy loading) 5. Clicks "Load More" button if present to reveal additional products 6. Extracts product URLs from the loaded page 7. Saves the collected URLs to a SQLite database Args: category_url (str): The full URL of the Ray-Ban category page to scrape (e.g., 'https://www.ray-ban.com/usa/sunglasses/men-s') category (str): The product category identifier (e.g., 'sunglasses', 'eyeglasses', 'smart-glasses') gender (str): The target gender/age group for the products (e.g., 'Men', 'Women', 'Kids', 'Unisex') Raises: PlaywrightTimeoutError: If page loading or element selection timeouts occur Exception: For any other errors during the scraping process Notes: - Uses Playwright for browser automation - Implements scrolling to handle lazy-loaded content - Includes random delays between actions to mimic human behavior - Saves results directly to a SQLite database - Closes browser resources properly even if errors occur """ base_url = "https://www.ray-ban.com" user_agents = load_user_agents() async with async_playwright() as p: browser = await p.chromium.launch(headless=False) user_agent = random.choice(user_agents) context = await browser.new_context(user_agent=user_agent) page = await context.new_page() try: # Navigate to the page await page.goto(category_url, wait_until="domcontentloaded", timeout=120000) # Try to close any popup that appears await close_popup(page) # Continue with the rest of the scraping process last_height = await page.evaluate('document.body.scrollHeight') while True: await page.evaluate('window.scrollTo(0, document.body.scrollHeight)') await page.wait_for_timeout(2000) new_height = await page.evaluate('document.body.scrollHeight') if new_height == last_height: break last_height = new_height await page.wait_for_selector('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div > div.rb-load-more', timeout=60000) while True: try: load_more_button = await page.query_selector('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div > div.rb-load-more > button') if load_more_button: await load_more_button.click() await page.wait_for_timeout(random.randint(3000, 5000)) else: break except PlaywrightTimeoutError: print("Timeout error occurred while waiting for 'Load More Products' button.") break html_content = await page.content() soup = BeautifulSoup(html_content, 'html.parser') product_elements = soup.select('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div > div.rb-products.grid > a') product_urls = [base_url + element['href'] for element in product_elements] for url in product_urls: save_to_database(url, category, gender) print(f"Scraped {len(product_urls)} products for {category} - {gender}") except PlaywrightTimeoutError: print(f"Page.goto timeout exceeded while trying to load {category_url}") except Exception as e: print(f"An error occurred: {e}") finally: await context.close() await browser.close() ``` This is the 'workhorse' part of the scraping operation; basically, a professional shopper who knows exactly how to navigate through Ray-Ban's website and gather product information. The function begins by launching a browser, using Playwright, with a randomly selected user agent, one of those disguises we discussed earlier. It navigates to the specified category page on Ray-Ban's website and then begins gathering product URLs. At the same time, the function performs several complex tasks, among which attempting to close all pop-ups that appear, scrolling the full page so that all products are loaded in case some web pages use lazy loading that displays only the first products as one scrolls, clicks "Load More" buttons in case they exist, and extracts any product URLs found. It is also well designed to wait for the actual content loading and includes random delays between actions to make the browsing pattern look more human-like. It includes error handling for common problems such as timeout errors and network problems that may go wrong. If something does go wrong, it ensures to properly close the browser and clean up after itself. All the URLs it collects get saved to the database using the save\_to\_database function we discussed earlier. ### main() ``` async def main(): """ Main function to initialise the database and run the scraping tasks asynchronously. This function initialises the SQLite database by calling `initialize_database()`, ensuring that the necessary table structure is in place. It then sequentially runs the `scrape_rayban_product_urls()` function for various product categories and genders, scraping and storing product URLs from each category on the Ray-Ban website. The asynchronous nature of the function allows for efficient handling of multiple web scraping tasks. """ initialize_database() # usage with different categories and genders await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/men-s', 'sunglasses', 'Men') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/women-s', 'sunglasses', 'Women') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/toddlers', 'sunglasses', 'Toddlers') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/little-kids', 'sunglasses', 'Little-Kids') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/kids', 'sunglasses', 'Kids') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/sunglasses/teenager', 'sunglasses', 'Teenagers') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/men-s', 'eyeglasses', 'Men') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/women-s', 'eyeglasses', 'Women') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/toddlers', 'eyeglasses', 'Toddlers') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/little-kids', 'eyeglasses', 'Little-Kids') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/kids', 'eyeglasses', 'Kids') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/eyeglasses/teenager', 'eyeglasses', 'Teenagers') await scrape_rayban_product_urls('https://www.ray-ban.com/usa/ray-ban-meta-smart-glasses', 'smart-glasses', 'Unisex') print('Product URLs have been scraped and saved to the database.') if __name__ == '__main__': asyncio.run(main()) ``` The major function will be compared to an orchestra conductor; it harmonizes all the functions together and makes them work. It begins with ensuring that the database is ready. Then, it orchestrates a series of scraping operations on the different categories and genders of Ray-Ban products by calling the initialize\_database function. It repeats the scraping function to feed each combination of product category (sunglasses, eyeglasses, smart-glasses) and target demographic (men, women, kids, etc). As the function makes use of async/await patterns to carry out this operation efficiently, it's actually like having multiple workers collecting information simultaneously but in an organised way. When everything is done, it prints a message to let you know the scraping is complete. ## Product Data Extraction This is a Python code implementing an advanced web scraping system to provide detailed product information from the Ray-Ban website. Using asynchronous programming with Playwright for browser automation, BeautifulSoup for HTML parsing, and SQLite for data storage, this script takes off by first initializing a database to store URLs and scraped data and systematically processes each unscraped URL. It travels through each product page, noticing possible roadblocks such as country selection pop-up, scrolling to lazy-loaded content, and extracting name, collection, price, color and technical specifications of the products. Scrape data is saved to the database with the URL marked as processed. The system has error handling, logging, and randomised user agents to make it reliable and detectable. This comprehensive approach will result in efficient scalabe scraping of large quantities of product pages while maintaining structured database information. ### Import Section ``` import sqlite3 import asyncio import random import os from playwright.async_api import async_playwright from bs4 import BeautifulSoup import logging import json ``` This section imports all the required libraries for scraping the data of the product. It includes pandas for data manipulation, sqlite3 for database operations, asyncio for asynchronous programming, os for file operations, playwright for web automation, BeautifulSoup for HTML parsing, logging, and JSON handling. Such imports lay out the basis of a solid web scraping script to operate on so many aspects of collecting, processing, and storing data. ### load\_user\_agents Function ``` def load_user_agents(file_path): """ Loads a list of user agents from a specified file, filtering out empty lines. This function is used to provide rotating user agents for web scraping to avoid detection and potential IP blocks. Args: file_path (str): Path to the file containing user agents, one per line Returns: list: A list of cleaned user agent strings Raises: FileNotFoundError: If the specified file_path does not exist """ if os.path.exists(file_path): with open(file_path, 'r') as f: # Create list comprehension to strip whitespace and filter empty lines return [line.strip() for line in f if line.strip()] else: raise FileNotFoundError(f"The file {file_path} does not exist.") ``` The \`load\_user\_agents\` function is another important part of web scraping; it's a function used to increase the likeliness of mimicking different web browsers and avoid detection. This function reads a list of user agent strings from a file, the contents of which can be rotated during the course of your scrape so that all requests appear as if they are coming from different browsers or devices. It begins by checking if a specified file exists using \`os.path.exists(file\_path)\`. If a file is found, it opens the file and reads its contents. A list comprehension is used to process each line of the file, stripping whitespace and filtering out any empty lines. This efficient approach ensures that only valid user agent strings are included in the final list. The function raises a \`FileNotFoundError\` with a custom error message if the specified file is not found. This type of error handling is important because it alerts the user right away if there's an issue with the user agent file, preventing the scraper from running without this critical component. A separate file for user agents makes updating and maintaining the list of user agents easy without having to modify the main script. ### parse\_product\_name Function ``` def parse_product_name(soup): """ Extracts the product name from the BeautifulSoup object of a product page. Targets the main product title element using a specific CSS selector path. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: str: The product name if found, 'N/A' if the element is not present """ # CSS selector for the product name heading product_name_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.sticky-sidebar > div > div > div.rb-product-information > div.rb-product-information__product-name-and-wishlist > h1.rb-product-name') return product_name_elem.get_text(strip=True) if product_name_elem else 'N/A' ``` \`parse\_product\_name\` will extract the name of the product from a product page's HTML content. It takes as an argument a BeautifulSoup object, an object that already represents the parsed HTML content of a product page. This is an important function in the scraping process for accurately identifying and then categorising the products. The following function uses a particular CSS selector to find the element containing the name of the product. The selector in this case: \`body > div.rb-app\_\_main.static-header.loaded.rb-app\_\_header--static > div.rb-pdp-page > div > div.rb-pdp\_\_scrollable-area > div.rb-pdp\_\_sidebar > div.sticky-sidebar > div > div > div.rb-product-information > div.rb-product-information\_\_product-name-and-wishlist > h1.rb-product-name\` is pretty specific, and this means that the HTML structure of the target website must be complex and the product name must be deeply nested within the DOM. If the element is found, it uses the \`get\_text(strip=True)\` function to extract the text content of the element, removing any leading or trailing whitespace. In the case that the element is not found-this could be if the structure of the page has changed, or if there is a mistake on the way the page loads-this function returns 'N/A.' This fallback ensures that the scraping process can continue even if some product names can't be extracted, hence maintaining data consistency. ### parse\_collection Function ``` def parse_collection(soup): """ Extracts the collection name from the BeautifulSoup object of a product page. Collection name typically indicates the product line or series. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: str: The collection name if found, 'N/A' if the element is not present """ # CSS selector for the collection name element collection_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.sticky-sidebar > div > div > div.rb-product-information > div.rb-product-information__status-message-label > div > div > p') return collection_elem.get_text(strip=True) if collection_elem else 'N/A' ``` Thus, the \`parse\_collection\` function is designed using a similar pattern as that of the \`parse\_product\_name\` function but for collecting the name of a collection of a product. In most e-commerce scenarios, the name of a collection is a series or product line and would be of use in categorizing and analyzing the product. It also employs a specific CSS selector to locate the element that contains the name of the collection. The selector \`body > div.rb-app\_\_main.static-header.loaded.rb-app\_\_header--static > div.rb-pdp-page > div > div.rb-pdp\_\_scrollable-area > div.rb-pdp\_\_sidebar > div.sticky-sidebar > div > div > div.rb-product-information > div.rb-product-information\_\_status-message-label > div > div > p\` might indicate a common find inside a paragraph on a product page in a section that bears the name of the collection. It returns the text content of the element in case it is found or 'N/A' otherwise in a similar way like the \`parse\_product\_name\` function does. In this way, all the parsing functions have a uniform approach that helps maintain uniform data structure even in cases where some information is missing from some product pages. ### parse\_colors\_number Function ``` def parse_colors_number(soup): """ Extracts the number of available colors from the BeautifulSoup object. This information is typically displayed in the color selection section. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: str: The number of colors available if found, 'N/A' if the element is not present """ # CSS selector for the colors count element colors_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.sticky-sidebar > div > div > div.rb-right-shoulder__info > div.rb-colours > div.rb-colours__title') return colors_elem.get_text(strip=True) if colors_elem else 'N/A' ``` The \`parse\_colors\_number\` function aims to extract the number of different colours available for a given product. This might prove to be very useful for inventory analysis or in studies on product variety, or as a feature in any tool that helps to compare products. The function addresses an element of the page, typically holding colour count information. The CSS selector applied(\`'body > div.rb-app\_\_main.static-header.loaded.rb-app\_\_header--static > div.rb-pdp-page > div > div.rb-pdp\_\_scrollable-area > div.rb-pdp\_\_sidebar > div.sticky-sidebar > div > div > div.rb-right-shoulder\_\_info > div.rb-colours > div.rb-colours\_\_title'\`) suggests that normally, such information is held in a title element of the color selection part of the product page. Like all other parsing functions, it will return a 'N/A' if the element was not found; otherwise, it returns the text content of the element. This is something whose output might then need further processing to provide only the numeric value if that's how the website presents the color count (e.g., "5 colours available" might need to be parsed to just "5"). ### parse\_model\_code Function ``` def parse_model_code(soup): """ Extracts the model code from the BeautifulSoup object of a product page. Model code is a unique identifier for the product variant. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: str: The model code if found, 'N/A' if the element is not present """ # CSS selector for the model code element code_elem = soup.select_one('#-answer >div > div.rb-product-details__model-code') return code_elem.get_text(strip=True) if code_elem else 'N/A' ``` The function \`parse\_model\_code\` will retrieve the unique identifier for a specific model code of a product variant. Model codes are of great importance in the e-commerce business and inventory management since this helps in the unique identification of different product variants. This particularly applies to situations where multiple colors, sizes, or configurations of the same product are available. This function uses a different method of selecting the target element than previous ones used. Instead of the long nested CSS selector, it uses an ID selector (\`'#-answer >div > div.rb-product-details\_\_model-code'\`). That would imply that the model code is probably found in a more regularly structured part of the page, possibly in a product details or specifications section. The function adheres to the standard pattern returning the text content if the element exists, and 'N/A' otherwise. The model code data extracted by this function can be valuable for cross-referencing products between different systems or databases and tracking specific product variants in inventory or sales analyses. ### parse\_frame\_description Function ``` def parse_frame_description(soup): """ Extracts the detailed frame description including various features and specifications. Creates a dictionary mapping feature labels to their corresponding values. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: dict: Dictionary containing feature labels as keys and their values as values """ frame_description = {} # Select all feature elements features = soup.select('.rb-product-detail__features > div') for feature in features: # Extract label and value for each feature label = feature.select_one('.rb-product-detail__label').get_text(strip=True) value = feature.select_one('.rb-product-detail__feature > span').get_text(strip=True) frame_description[label] = value return frame_description ``` The \`parse\_frame\_description\` function is somewhat more complex than the previous parsing functions. It is supposed to collect a very detailed description of the product's frame with other features and specifications. This creates a well-formed representation of the particular features of the product, which can be extremely valuable for really detailed product comparisons, or for filling up really comprehensive databases with the information about products. Unlike all other functions that return a single string value, this function returns a dictionary. This function starts by initializing an empty dictionary called \`frame\_description\`. It then chooses all elements with the class \`.rb-product-detail\_\_features > div\`, representing probably a list of feature elements on the product page. For each feature element, the function extracts the two pieces of information- the label (which is the dictionary key) and the value (which is the dictionary value). It uses particular selectors to locate these in each element: the label with \`.rb-product-detail\_\_label\` and the value with \`.rb-product-detail\_\_feature > span\`. This works well for fetching multiple features along with their details in an organized way. The resulting dictionary gives a full impression of the product's specifications, which are easily processed or stored in a database to then be analyzed further. This level of detail is indispensable for products for which technical specifications play a crucial role for consumers, as with eyewear or other special products. ### parse\_price\_details Function ``` def parse_price_details(soup): """ Extracts pricing information including MRP, sale price, and discount details. Handles both regular and sale pricing scenarios. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: tuple: Contains (mrp, sale_price, discount) - mrp (str): Original/Maximum Retail Price - sale_price (str): Discounted price if available - discount (str): Discount percentage if available """ # First try to find regular price price_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.rb-sticky-bar > div.rb-sticky-bar-left > div.rb-sticky-bar-left__title > div > span.rb-prices__normal') if price_elem: # If regular price found, no discount scenario mrp = price_elem.get_text(strip=True) return mrp, mrp, 0 else: # Handle sale price scenario price_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.rb-sticky-bar > div.rb-sticky-bar-left > div.rb-sticky-bar-left__title > div > span.rb-prices__list') sale_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.rb-sticky-bar > div.rb-sticky-bar-left > div.rb-sticky-bar-left__title > div > span.rb-prices__discounted') discount_elem = soup.select_one('body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-pdp__sidebar > div.rb-sticky-bar > div.rb-sticky-bar-left > div.rb-sticky-bar-left__title > div > span.rb-promo-badge.rb-promo-badge--small') # Extract values with fallbacks mrp = price_elem.get_text(strip=True) if price_elem else '0' sale_price = sale_elem.get_text(strip=True) if sale_elem else '0' discount = discount_elem.get_text(strip=True) if discount_elem else '0' return mrp, sale_price, discount ``` The \`parse\_price\_details\` function should handle the sometimes tricky business of gathering pricing information from product pages. It can be applied equally to simple cases where prices are straightforward and extracted, as well as to more complex cases where products are on sale, making it relevant in a wide range of pricing schemes found in e-commerce. It first tries to find a regular price element. When it does find one, it assumes that no discount has been applied and returns the same value for both MRP and sale price, with a discount of 0\. This covers products that are not on sale. If no regular price is found, it then looks for sale pricing elements. It searches for three separate pieces of information: original price (MRP), discounted price, and discount percentage. Each of these are extracted by using the proper CSS selectors targeting different elements on the page. The function uses fallback mechanism, just in case some of the following elements are not found; consequently it returns '0' by default. Thus, the function will always return a consistent tuple of (mrp, sale\_price, discount), possibly with some information missing from the page. It makes the function robust against variations of page structure or missing price information. ### parse\_color\_options Function ``` def parse_color_options(soup): """ Extracts all available color options for the product. Processes the color variant buttons to get their alt text descriptions. Args: soup (BeautifulSoup): Parsed HTML content of the product page Returns: str: Comma-separated string of available color options, or 'N/A' if none found """ # Find the container div for color options target_div = soup.find('div', class_='rb-colours-list') if target_div: # Find all color variant buttons and extract their image alt texts buttons = target_div.find_all('button', class_='rb-colour-variant') alt_texts = [img['alt'] for button in buttons for img in button.find_all('img') if 'alt' in img.attrs] return ', '.join(alt_texts) return 'N/A' ``` The \`parse\_color\_options\` function parses out any information on all the different color options available for a product. It comes in especially handy for products that are available in more than one color, as it gathers all the options on the color. The function starts with the search for a div element by class \`rb-colours-list\` - probably containing all colour option elements. If this container exists, it proceeds to look inside it for all button elements with the class \`rb-colour-variant\`. These buttons typically represent individual colour options. The function retrieves the \`alt\` text of any img elements contained within the button for each button it finds. This typically contains a name or description of the color. It is a nice trick, using accessibility features (alt text) to get meaningful color information that might be more likely to be in use and more descriptive than some effort to parse colour codes or names from other attributes. Lastly, the function concatenates all the colour names it has extracted into a single string separated by commas. When no colour options are found, in the sense that either the container div was missing or there were no buttons of color variant type, the function will return 'N/A'. This way, even if colour information is unavailable, the function will always return a string and the data structure will always be consistent. ### scrape\_product\_details Function ``` async def scrape_product_details(page, url, category, gender): """ Asynchronously scrapes detailed product information from a Ray-Ban product page using Playwright. Handles various page interactions including country selection popups, scrolling, and dynamic content loading. Args: page (Page): Playwright page object for browser interaction url (str): The product URL to scrape category (str): Product category (e.g., 'sunglasses', 'eyeglasses') gender (str): Target gender for the product Returns: dict: A dictionary containing scraped product details including: - url: Product URL - category: Product category - gender: Target gender - name: Product name - collection: Collection name - number_of_colors: Available color count - model_code: Product model code - frame_description: Dictionary of frame details - mrp: Maximum Retail Price - sale_price: Current sale price - discount: Discount percentage - colors: Available color options """ try: # URL encode spaces to ensure proper formatting encoded_url = url.replace(" ", "%20") # Navigate to the page with extended timeout for slow connections response = await page.goto(encoded_url, timeout=60000) logging.info(f"Scraping {url} - Response status: {response.status}") # Check for successful page load if response.status != 200: logging.error(f"Failed to load page, status code: {response.status}") # Handle country selection popup that appears on first visit try: # Complex selector for the country selection button (first country option) await page.locator('#rb-header-app > div.modal-wrapper.modal-wrapper--header.modal-wrapper--display > div.modal-content-wrapper > div > div.rb-modal-content > div > div > span > div.rb-country-overlay-modal__flag-container > a:nth-child(1)').click() logging.info("Country selection popup handled successfully.") # Wait for popup animation to complete await asyncio.sleep(2) except Exception as e: logging.warning(f"Country selection popup could not be handled: {e}") # Scroll the page to trigger lazy loading of content for _ in range(10): await page.mouse.wheel(0, 1000) # Scroll down 1000 pixels await asyncio.sleep(1) # Wait for content to load # Get initial page content and parse with BeautifulSoup content = await page.content() soup = BeautifulSoup(content, 'html.parser') # Selectors for product details accordion button button_selector = "body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-left-shoulder > div.rb-accordions > div.rb-accordion.rb-accordion--with-custom-icon.rb-accordions__product-details > button" open_button_selector = "body > div.rb-app__main.static-header.loaded.rb-app__header--static > div.rb-pdp-page > div > div.rb-pdp__scrollable-area > div.rb-left-shoulder > div.rb-accordions > div.rb-accordion.rb-accordion--with-custom-icon.rb-accordion--is-open.rb-accordions__product-details > button" # Check if product details section is already open is_open = await page.query_selector(open_button_selector) if not is_open: # Click to open product details section await page.click(button_selector) logging.info("Product details section opened.") # Wait for content to load after clicking await page.wait_for_timeout(2000) # Get updated page content after opening details section content2 = await page.content() soup2 = BeautifulSoup(content2, 'html.parser') except Exception as e: logging.error(f"Could not extract product details from {url}: {e}") # Return all scraped data in a structured dictionary return { 'url': url, 'category': category, 'gender': gender, 'name': parse_product_name(soup), 'collection': parse_collection(soup), 'number_of_colors': parse_colors_number(soup), 'model_code': parse_model_code(soup2), 'frame_description': parse_frame_description(soup2), 'mrp': parse_price_details(soup)[0], 'sale_price': parse_price_details(soup)[1], 'discount': parse_price_details(soup)[2], 'colors':parse_color_options(soup) } ``` The \`scrape\_product\_details\` function is a wrapper for the web scraping operation. It is an asynchronous function that utilizes Playwright to step through an individual product page to extricate detailed information. It is a "one-stop shop" that encapsulates a whole interaction with a web page and retrieval of data points on a product. The function first encodes the URL to handle spaces, navigates to the page with an extended timeout for slow connections, logs the response status. This is important for monitoring the flow of scraping as the status code needs to not be 200. A good characteristic of this function is that it deals with dynamic page elements. It tries to handle a country selection popup that may appear on the first visit to the site. This really shows prudence in dealing with the typical, real-world challenges of web scraping. The function implements scrolling to also have an effect of triggering lazy loading of content, which would ensure all the necessary information is loaded before the extraction process begins. The function then applies the parsing functions defined above (\`parse\_product\_name\`, \`parse\_collection\`, etc.), extracting different pieces of information from the page. In so doing, it also handles the opening of a product details section if it's not already open-an example of how to interact with the page in order to obtain hidden information. Finally, it compiles all the extracted data into a structured dictionary, providing one with a comprehensive set of information about the product. ### init\_database Function ``` def init_database(): """ Initializes SQLite database and creates necessary tables for storing product data. Creates two main tables: 'urls' for tracking URLs to scrape and 'data' for storing product information. Also handles database schema updates by adding new columns if needed. Returns: sqlite3.Connection: Database connection object Tables Created: urls: - id: Primary key - url: Product URL - category: Product category - gender: Target gender - scraped: Flag indicating if URL has been processed data: - All product details fields corresponding to scrape_product_details output """ # Create connection to SQLite database conn = sqlite3.connect('rayban_products.db') cursor = conn.cursor() # Create URLs table if it doesn't exist cursor.execute(''' CREATE TABLE IF NOT EXISTS urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, category TEXT, gender TEXT, scraped INTEGER DEFAULT 0 ) ''') # Check if scraped column exists, add if missing cursor.execute("PRAGMA table_info(urls)") columns = [column[1] for column in cursor.fetchall()] if 'scraped' not in columns: cursor.execute("ALTER TABLE urls ADD COLUMN scraped INTEGER DEFAULT 0") # Create data table for storing product information cursor.execute(''' CREATE TABLE IF NOT EXISTS data ( url TEXT, category TEXT, gender TEXT, name TEXT, collection TEXT, number_of_colors TEXT, model_code TEXT, frame_description TEXT, mrp INTEGER, sale_price INTEGER, discount INTEGER, colors TEXT ) ''') conn.commit() return conn ``` The \`init\_database\` function is responsible for initializing the SQLite database that will hold all the scraped data. This function also shows good practices in the management of databases and design of schemas for web scraping projects. The function creates two main tables: 'urls' and 'data'. The 'urls' table is designed to track URLs that need to be scraped, including metadata like category and gender, and a flag to indicate whether the URL has been processed. The 'data' table is structured to store all the detailed product information that will be scraped. An interesting feature of this function is that it allows the database schema to be updated. It checks whether the 'scraped' column exists in the 'urls' table and adds it if it is missing. This forward-thinking approach allows for easy updates to the database structure without needing to recreate the entire database. It returns a connection object, which is quite a good practice since it allows the calling code to also manage the database connection lifecycle. ### load\_urls\_from\_db Function ``` def load_urls_from_db(conn): """ Retrieves all unscraped URLs from the database along with their associated metadata. Args: conn (sqlite3.Connection): Database connection object Returns: list: List of dictionaries containing unscraped URLs with their metadata: - url: Product URL - category: Product category - gender: Target gender Note: Only returns URLs where scraped=0 in the database """ cursor = conn.cursor() # Select only unscraped URLs cursor.execute("SELECT url, category, gender FROM urls WHERE scraped = 0") # Convert results to list of dictionaries for easier handling return [{'url': row[0], 'category': row[1], 'gender': row[2]} for row in cursor.fetchall()] ``` The \`load\_urls\_from\_db\` function is responsible for fetching unscraped URLs from the database. This function is fundamental to the scraping workflow by providing the list of the URLs that need to be processed. The function performs a SQL query to fetch the records of URLs for which 'scraped' status is 0, or in other words, URL records that have not yet been processed. Then, through a list of dictionaries, convert the query results. Hence, it becomes easier to iterate over the list created by the main scraping function and access the metadata associated with the scraped URLs. This function allows the process to be paused and resumed effectively as it will always start with the URLs that haven't been processed yet by retrieving only the unscraped URLs. ### update\_url\_status Function ``` def update_url_status(conn, url): """ Updates the scraped status of a URL in the database to mark it as processed. Args: conn (sqlite3.Connection): Database connection object url (str): The URL to mark as scraped Note: Sets scraped=1 for the specified URL in the urls table """ cursor = conn.cursor() # Mark URL as scraped cursor.execute("UPDATE urls SET scraped = 1 WHERE url = ?", (url,)) conn.commit() ``` The \`update\_url\_status\` is one of the simplest workflows yet most important in scraping that marks a URL in the database as scraped after the processing. It takes two parameters: database connection and URL. What it does is perform the SQL UPDATE statement that simply sets the 'scraped' flag of given URL to 1\. The function immediately commits its change, ensuring that it's always keeping the database consistent. This function can update the status of each URL after the scraping so as not to scrape over and over again and easily track progress in long-running scraping operations. ### save\_to\_db Function ``` def save_to_db(conn, product_data): """ Saves scraped product information to the database. Handles conversion of complex data types (like dictionaries) to JSON for storage. Args: conn (sqlite3.Connection): Database connection object product_data (dict): Dictionary containing all scraped product information Must contain all fields corresponding to the data table schema Note: Converts frame_description dictionary to JSON string for storage Commits the transaction immediately after insertion """ cursor = conn.cursor() # Convert frame description dictionary to JSON string frame_description_json = json.dumps(product_data['frame_description']) # Insert product data into database cursor.execute(''' INSERT INTO data (url, category, gender, name, collection, number_of_colors, model_code, frame_description, mrp, sale_price, discount, colors) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) ''', (product_data['url'], product_data['category'], product_data['gender'],product_data['name'], product_data['collection'], product_data['number_of_colors'],product_data['model_code'], frame_description_json, product_data['mrp'],product_data['sale_price'], product_data['discount'], product_data['colors'])) conn.commit() ``` The \`save\_to\_db\` function saves product data scrapped to the database. This function is used for saving complex data types in a SQLite database. One interesting feature of this function is how it handles the 'frame\_description' field. Since this field is a dictionary, which SQLite cannot store directly, this dictionary is converted to a JSON string before being inserted. This makes it easy to store structured data in an easily retrievable and parsable format. It makes use of parameterised SQL queries for the inserting process and hence avoids the use of SQL injection attacks. Then it commits the transaction, just after the insertion happens. So, in result, every product data has saved instantly after processing. ### main Function ``` async def main(): """ Main execution function that orchestrates the entire scraping process. Manages database connections, browser instances, and the scraping workflow. Process Flow: 1. Initializes database connection 2. Loads unscraped URLs 3. For each URL: - Creates new browser instance with random user agent - Attempts to scrape product details - Saves data to database - Updates URL status - Closes browser instance 4. Closes database connection Note: Uses Playwright in non-headless mode for browser automation Implements error handling and logging for each step Runs asynchronously for better performance """ # Initialize database and load unscraped URLs conn = init_database() urls_to_scrape = load_urls_from_db(conn) user_agents = load_user_agents('user_agents.txt') # Start Playwright context async with async_playwright() as playwright: for url_data in urls_to_scrape: url = url_data['url'] category = url_data['category'] gender = url_data['gender'] print(f"Scraping: {url}") # Launch new browser instance for each URL browser = await playwright.chromium.launch(headless=False) # Create new context with random user agent context = await browser.new_context(user_agent=random.choice(user_agents)) page = await context.new_page() try: # Scrape product details and save to database product_data = await scrape_product_details(page, url, category, gender) save_to_db(conn, product_data) update_url_status(conn, url) logging.info(f"Successfully saved data for {url}") except Exception as e: logging.error(f"Error scraping {url}: {e}") finally: # Clean up browser resources await page.close() await context.close() await browser.close() # Close database connection conn.close() # Entry point of the script if __name__ == "__main__": asyncio.run(main()) ``` The \`main\` function is the orchestrator of the entire scraping process. It is an asynchronous function that ties all other components of the scraper together. This function begins initializing the database and loads the list of URLs to scrape while loading the list of user agents that are used for randomization of browser's identity for any request. The main characteristic of this function is that it utilizes Playwright for automating the interactions in the browser. For every URL, it starts a new browser instance with a random user agent. This results in a new context for each scraping operation, thus helping avoid detection and potential IP blocks. The function employs a try-except block for each URL; this way, it will allow scraping to continue even when a failure is caused by one URL. It records successes and errors, which is useful information for monitoring the scraping operation. After processing each URL, it cleans up resources by the browser; this is crucial for memory management with long-running scraping. Lastly, the function closes the database connection to manage resources properly. At the entry point of the script, calling \`asyncio.run(main())\` allows asynchronous execution of the whole process of scraping, which could improve performance, particularly in network operations and concurrent scraping. ## Conclusion Web scraping has revolutionized how we gather and analyze data, making it possible to automate product discovery and gain insights efficiently. This project demonstrated how web scraping can be used to extract detailed product information from Ray-Ban’s eyewear collections, streamlining the research process that would otherwise be tedious and time-consuming. By leveraging Playwright and Beautiful Soup, we navigated the complexities of dynamic web content, ensuring accurate and structured data collection. While web scraping is a powerful tool, it is essential to adhere to ethical and legal considerations, respecting website policies and terms of service. Moving forward, the extracted data can be used for price tracking, trend analysis, or even building a recommendation system. This project not only highlights the potential of automation in e-commerce research but also opens doors to further innovations in data-driven decision-making. Connect with[ Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. AUTHOR I’m Shahana, Data Analyst at Datahut. I specialize in building data pipelines that turn complex web data into actionable insights for brands in fashion, eyewear, and retail. At Datahut, we’ve spent over 10 years helping companies automate product discovery, monitor competitors, and make data-driven decisions. In this blog, I walk you through how web scraping can be used to track Ray-Ban’s product range across multiple platforms—saving time, improving visibility, and supporting strategic planning. If you're looking to automate product discovery for your brand, connect with us through the chat widget on the right. We’re happy to help. ### Fashion Pricing Insights Retailers Can Learn from H&M URL: https://www.blog.datahut.co/post/what-fashion-retailers-can-learn-from-h-m-s-data-driven-pricing/ Last updated: 2026-09-07T09:44:32.000Z In the ever-changing fashion industry , pricing isn't just about numbers—it's about perception, positioning, and precision. One brand that has mastered this art is H&M. With its ability to offer trend-led fashion at accessible prices while dabbling in premium offerings, H&M presents a goldmine of pricing lessons for fashion retailers . [Recent data-driven analysis , powered by web scraping techniques , reveals just how strategic H&M’s pricing really is](https://www.blog.datahut.co/post/h-m-s-pricing-strategy-detailed-analysis-data/). From offering multi-tiered product lines to making data-backed discount decisions, the brand’s playbook is a blueprint for retailers aiming to optimize margins without compromising customer satisfaction or customer loyalty . Here’s what fashion companies can learn from H&M’s pricing strategy and how to apply those lessons. ## Balance Affordability with Premium Appeal At first glance, H&M is a budget-friendly brand. Dig deeper, and you'll find it’s also an aspirational one. - Price Range : ₹299 to ₹29,999 - Median Price : ₹1,799 - 97.6% of products priced under ₹5,000 This pricing architecture allows H&M to attract two kinds of shoppers: price-conscious buyers looking for everyday wear and fashion-forward customers seeking exclusivity through limited-edition collections like H&M Rokh (₹12,107 avg.) and Studio Collection (₹7,229 avg.). Takeaway for Retailers : Structure your catalog to include both affordable essentials and premium products. Think core staples for everyday shoppers and curated, design-forward collections for your fashion consumers . This dual-pricing strategy helps capture a wider customer base without diluting your brand. ## Use Discounts Strategically- Not Habitually Discounts can be a powerful tool but only when used wisely. H&M offers a masterclass in discount restraint: - Only 1.65% of H&M’s products were found to be discounted. - Most common discounts: 30%–40%. - Deep cuts (50%–60%) are rare and reserved for specific clearance periods. This approach maintains brand value and avoids customer behavior habits of "waiting for sales." Takeaway for Retailers :Stop the endless cycle of discounting. Instead, build a brand worth paying full price for. Offer strategic markdowns during peak seasons or end-of-season clearances, not as your default pricing model. By doing so, you can improve inventory accuracy and reduce excess inventory risks. ## Let Data Drive Your Pricing Decisions H&M doesn’t guess. It gathers. Every pricing decision is informed by real-time data , often gathered using web scraping and analytics. For example, scraped data helped analyze: - Pricing distributions by category - Discount frequency and depth - Impact of materials, fits, and sleeve lengths on pricing These insights help the brand respond quickly to market shifts , competitor activity, and customer preferences . Takeaway for Retailers : Invest in tools like web scraping and data analytics . Services like Datahut help fashion firms pull product and pricing data from competitors or marketplaces in real time. With this, you can optimize prices, monitor trends, and respond faster than traditional retail cycles allow. Incorporating AI-driven demand forecasting and machine learning models can further enhance your supply chain operations and reduce inaccurate demand forecasting risks. ## Differentiate Collections Based on Consumer Segments One size doesn’t fit all—and neither does one collection. H&M targets different segments through its collection-specific pricing: - New Arrivals : Trendy and affordable (\~₹1,737 avg.) - Studio Collection : Mid-luxury (\~₹7,229 avg.) - Rokh Collection : Designer and premium (\~₹12,107 avg.) Each collection speaks to a unique audience, from college students to professionals to trendsetters. Takeaway for Retailers : Don’t just create products- create personas. Structure collections that serve specific demographics, lifestyles, and budgets. This allows for sharper targeting and stronger fashion branding . Additionally, leveraging social media platforms and digital platforms can amplify your reach and engagement with these diverse consumer groups. ## Material Quality & Design Affect Willingness to Pay Another layer of H&M’s pricing is driven by product attributes—what it’s made of and how it fits. High-priced materials: - Polyester Lining : ₹3,486 avg. - Cotton Denim : ₹2,481 avg. - Linen : ₹2,292 avg. Fit-Based Pricing: - Oversized Fit : ₹2,581 avg. - Skinny Fit : ₹1,552 avg. - Slim Fit : ₹1,861 avg. This shows how fashion meets functionality—and how design trends influence price perception. Takeaway for Retailers :Track which materials and designs resonate most with your audience and price accordingly. Premium fabrics or trendy fits can command higher prices—as long as you tell the right story through your product pages and marketing. Consider incorporating recycled materials or sustainable fashion supply chain practices to appeal to eco-conscious consumers and reduce your carbon footprint . ## Focus on High-Value Categories Some categories naturally command higher prices. At H&M, these include: - Coats : ₹7,047 avg. - Blazers : ₹4,354 avg. - Outdoor Trousers : ₹3,900+ avg. These items are perceived as functional, durable, and aspirational—justifying the higher price tags. Takeaway for Retailers : Know where your margin-rich opportunities lie. Whether it’s outerwear, formalwear, or occasion-specific clothing, identify high-value categories and invest in those for maximum ROI. Align these efforts with environmental sustainability goals to enhance your brand’s competitive edge . ## Conclusion: Analyze, Adapt, and Act H&M’s pricing strategy isn’t just about affordability—it’s about intentionality. By blending mass-market accessibility with premium touches, using data to guide every move, and avoiding the trap of over-discounting, H&M shows how fashion pricing can be both profitable and perceptive. Fashion retailers looking to stay ahead need to: - Analyze real-time pricing trends - Build diversified collections - Understand the impact of materials, fits, and functionality - Make pricing a function of both art and data In the words of the data: smart pricing isn’t static. It evolves with your consumer—and the market. ## Ready to Level Up Your Pricing Strategy? Harness the power of data with tools like web scraping and expert analysis. Services like [Datahut](https://www.datahut.co/?ref=blog.datahut.co) can equip you with the competitive insights you need to win in today’s fast-paced retail landscape . Are you a fashion retailer looking to benchmark your pricing strategy against industry leaders like H&M? Let us help you identify what works, adapt successful strategies, and stay competitive. Reach out today to get started. H&M's pricing strategy becomes even more valuable when supported by large-scale product data collection and analysis. To learn how such datasets are built, explore our tutorial on [how to scrape H&M product data using Python](https://www.blog.datahut.co/post/how-to-scrape-h-m-product-data-using-python/), which walks through the complete process of gathering H&M product information for competitive intelligence and retail analytics. Inspired by H&M’s strategic brilliance? Drop us a line or share your thoughts- we’d love to hear how you’re using data to price smarter! ### How AI Is Reshaping Startups Backed by Y Combinator URL: https://www.blog.datahut.co/post/y-combinator-2025-how-ai-is-reshaping-startups-and-markets/ Last updated: 2026-07-23T07:48:32.000Z In 2025, over 72% of new startups in Y Combinator are powered by artificial intelligence , signaling a seismic shift in how technology is driving innovation across industries. From automating mundane tasks to revolutionizing entire sectors, AI has moved beyond being a buzzword—it’s now the backbone of modern entrepreneurship. [Y Combinator (YC)](https://www.ycombinator.com/companies?ref=blog.datahut.co) has long been a beacon for early-stage startup founders , offering mentorship, funding, and access to a vast network of investors. Its portfolio reflects not only the evolution of technology but also the shifting priorities of global markets. From its humble beginnings in 2005 to becoming a launchpad for some of the world’s most successful companies, YC continues to redefine what it means to build and scale a startup. With advice from leaders like Garry Tan , Paul Graham , and Michael Seibel , YC has helped shape countless industry-defining companies . This analysis explores three major shifts in YC’s ecosystem: the explosive growth of AI-driven startups, the maturation of fintech, and the expanding global footprint of entrepreneurship. Together, these trends paint a picture of a dynamic and rapidly evolving startup landscape. ## Key Trends in YC Startups (2005–2025) ![key trends in YC](https://www.blog.datahut.co/content/images/2026/07/img-310.png.webp) ## A. Explosive Growth Followed by Market Corrections YC saw steady growth from 2005 to 2015, with a surge in startups peaking in 2020–2021\. However, post-2021, there’s been a decline due to economic slowdowns and a more cautious investment climate. Key Insight : The sharp drop in 2025 suggests fewer startups entering YC, reflecting either stricter selection criteria or broader macroeconomic challenges. This decline underscores the importance of resilience and adaptability in today’s competitive market. For many early-stage companies , achieving product-market fit remains critical to survival. ## B. Active Startups: Rising but Volatile The number of active startups closely follows total startup trends, with a steep rise from 2016 to 2021\. A decline in 2022, followed by a modest recovery in 2023–2024, suggests that while many startups thrive, others struggle to survive long-term. Key Insight : Stronger startups are persisting, while weaker ones have exited the market. This trend highlights the growing emphasis on sustainable business models and long-term value creation. Many successful startups attribute their success to customer obsession and a relentless focus on solving real-world problems. ## C. Acquisitions vs. Public Listings Acquisitions peaked around 2019–2020 but have since dropped significantly, likely due to economic uncertainty or a slowdown in big tech buyouts. Public listings remain rare, with IPO activity stagnating. Key Insight : Acquisitions remain the dominant exit strategy, reinforcing the idea that major tech companies prefer acquiring innovation rather than building it from scratch. This trend also reflects investor confidence in startups that demonstrate clear market traction. For example, Dylan Field and Jared Friedman have both emphasized the importance of network effects in scaling software companies. ## AI Boom: The Backbone of Modern Startups ![analysis of top industry tags](https://www.blog.datahut.co/content/images/2026/07/img-311.png.webp) ![ai companies accepted to ycombinator](https://www.blog.datahut.co/content/images/2026/07/img-312.png.webp) ## A. AI’s Unstoppable Rise AI-focused startups grew from 871 in 2024 to 1,140 in 2025 , accounting for 53% of all newly created YC startups . AI is now embedded across diverse domains: - Business & Enterprise Solutions : Automation in SaaS, B2B operations, and financial decision-making. - Developer Tools & Automation : AI-driven DevOps, debugging, and open-source contributions. - Fintech & Financial Services : Fraud detection, personalized banking, and risk assessment. - Healthcare & Biotech : Diagnostics, drug discovery, and patient management. - Retail & E-commerce : Recommendation engines, chatbots, and recruitment automation. Key Insight : AI is no longer confined to tech-heavy industries; it’s transforming every sector, from healthcare to retail, by enabling smarter, faster, and more efficient solutions. Seminal figures like Harj Taggar and Aravind Srinivas have highlighted the transformative potential of AI in reshaping traditional industries. ## B. Generative AI: Beyond the Hype Generative AI startups increased from 214 to 262 (+22.4%) , maintaining \~23% of all AI startups. Applications include: - AI-powered content creation, data visualization, and customer support. - Automated productivity tools in media, business, and consumer products. Key Insight : Generative AI is proving to be more than just a passing trend. Its ability to create new content, optimize workflows, and enhance creativity makes it a cornerstone of modern innovation. Raphael Schaad and David Lieb have both spoken about the leap from designer to founder and the role of generative AI in this transition. ## C. AI Startup Resilience AI startups show high survival rates, with relatively low inactivity. Acquisitions remain a key exit strategy, with notable peaks in 2017 and 2020. Key Insight : AI startups are resilient, highly investable, and attractive acquisition targets, ensuring continued venture capital interest. Their adaptability and scalability make them stand out in a crowded market. Co-founder & CEO Josh Reeves has often emphasized the importance of an equity stake in aligning incentives for long-term success. ## Fintech Evolution: From Disruption to Sustainability ![fintech companies accepted](https://www.blog.datahut.co/content/images/2026/07/img-313.png.webp) ## A. Peak Growth (2016–2021) Fintech startups skyrocketed, fueled by digital banking, crypto, lending, and payment solutions gaining mainstream adoption. ## B. Decline Post-2022 New fintech startups dropped significantly in 2024–2025, indicating possible market saturation or a shift in investor focus toward AI-integrated finance solutions. ## C. Long-Term Sustainability While new entrants slowed, active fintech startups remain strong, suggesting long-term sustainability. Many have secured funding, partnerships, or a strong customer base. Key Insight : Fintech is evolving into AI-powered financial solutions that prioritize automation, compliance, and efficiency. This transition marks a shift from rapid disruption to building lasting, impactful technologies. Tony Xu , co-founder & CEO of DoorDash, has highlighted the importance of potential customers in shaping product design. ## Global Expansion: Beyond Silicon Valley ![geographical distribution](https://www.blog.datahut.co/content/images/2026/07/img-314.png.webp) ## A. U.S. Dominance The U.S. remains the epicenter of YC startups, with 3,538 startups (up from 3,183) . San Francisco leads with 1,330 startups , though its numbers are declining due to high living costs and remote work trends. Key Cities : - New York : 403 startups, the strongest alternative to Silicon Valley. - Los Angeles : 92 startups, showing a slowdown. - Palo Alto & Mountain View : Declining numbers signal decentralization. ![analysis of top tech hubs](https://www.blog.datahut.co/content/images/2026/07/img-315.png.webp) ## B. Emerging Global Hubs International hubs solidify their positions: - London : 110 startups, leading Europe. - Bengaluru : 95 startups, India’s top startup city. - Toronto, Paris, Mexico City : Indicating North America, Europe, and Latin America’s growing significance. Key Insight : While the U.S. dominates, international hubs like India, the UK, and Canada are rising, showcasing the global nature of innovation and entrepreneurship. Boom Supersonic , founded by individuals with a background in aerospace engineering , exemplifies how diverse expertise drives innovation. ## Comparative Analysis (2005–2024 vs. 2005–2025) ![analysis of y combinator startups](https://www.blog.datahut.co/content/images/2026/07/img-316.png.webp) ## A. Growth Metrics - Total startups increased from 4,666 (2005–2024) to 5,173 (2005–2025) , adding 507 new companies (+10.9%) .Active startups grew from 3,307 to 3,600 , though their share dipped slightly from 70.9% to 69.6% . ## B. AI Dominance - AI is the standout sector, with 33 out of 46 new startups in 2025 tagging “artificial-intelligence.” ## C. Market Maturity - The ecosystem shows signs of maturing, with stabilization in active companies despite increased competition. Key Insight : Success now depends on deeper, more impactful innovation—not just being part of the YC network. Office Hours and mentorship programs continue to play a crucial role in guiding fast-growing YC startups . ## Conclusion YC continues to be a driving force in the startup world, but the rapid expansion of past years has given way to a more measured, competitive landscape. ## Key Takeaways : - AI is no longer a trend; it’s the foundation upon which startups are built. - Fintech is maturing, focusing on AI-powered automation and compliance. - Geographically, the U.S. dominates, but international hubs like India, the UK, and Canada are rising. - Success now depends on deeper, more impactful innovation, not just being part of the YC network. As we move forward, the startups that thrive will be those that leverage AI in meaningful, transformative ways—building solutions that endure in an increasingly competitive and technologically advanced world. Connect with [Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. Whether you’re tracking industry trends or analyzing market data, Datahut empowers your business with actionable intelligence. AUTHOR I’m Aarathi J, Marketing Manager at Datahut. With over 5 years of experience in data-driven marketing, I’ve collaborated with fashion, retail, and tech brands to turn raw data into strategic growth opportunities. At Datahut, we’ve spent more than a decade helping startups and enterprises alike unlock market intelligence using web scraping, AI, and custom data solutions. We’ve worked with fast-scaling tech companies—including those backed by accelerators like Y Combinator—to give them a competitive edge through timely, actionable data. If you're a startup founder or investor exploring how AI and data automation can fuel smarter decisions, start a conversation with us using the chat widget on the right. Let’s turn your data challenges into growth opportunities. ### How to Automate Lowe’s Product Data Extraction Using Web Scraping URL: https://www.blog.datahut.co/post/how-to-automate-lowe-s-product-data-extraction-using-web-scraping/ Last updated: 2026-07-23T07:48:32.000Z ## Introduction Did you know that most of the e-commerce businesses utilize web scratching to track competitor assessing, screen thing availability, and analyze client estimation? In today's data-driven economy, businesses depend on mechanized data collection to choose up a competitive edge. Instead of physically gathering points of interest, companies can use web scraping to extract vast amounts of data quickly and efficiently. This project focuses on automating the extraction of product data from Lowe's, one of the largest home improvement retailers. By utilizing Python , the scraper navigates through Lowe's location, collects principal thing details—including titles, costs, evaluations, and reviews—and stores the data in an SQLite database. The structured data can be used for market analysis, price comparisons, and inventory tracking, helping businesses make data-driven decisions. This documentation is designed for e-commerce managers, data analysts, and developers who want to leverage web scraping for business intelligence. Whether you're monitoring pricing trends, analyzing customer feedback, or researching market demand, this project provides a reliable and scalable approach to collecting valuable product insights. ## What is Web Scraping? Web scraping is the process of automatically extracting data from websites using scripts or programs. Instead of manually copying and pasting information, a web scraper accesses web pages, retrieves relevant data, and organizes it for further analysis. This technique is widely used in various industries, particularly in e-commerce, finance, and research, where large-scale data collection is necessary. Businesses use web scraping for multiple purposes, such as tracking competitor prices, analyzing customer reviews, monitoring stock availability, and gathering product specifications for database management. By automating data extraction, companies can save time, reduce errors, and gain real-time insights into market trends. In this project, web scraping enables the automated collection of product data from Lowe’s website. Using Playwright, the scraper efficiently navigates through the website, interacts with dynamic elements, and extracts structured data while handling challenges like changing page structures, missing data, and connection failures. ## overview of the project The initial stage consists of collecting the URLs of the product pages from the Lowe’s category pages. The scraper works by stepping into different segments of the website, detects HTML anchor () tags with product links, and saves them to an SQLite database. The scraper tries to maintain data integrity by cleaning duplicate URLs and retrieving information that will be used later. After acquiring the product URLs, the scraper extracts the data from every single product page and gets information such as product name, cost, ratings, total reviews, stock status and other details. The scraper collects information only through static HTML selectors, so only the necessary information is targeted. The information that was captured is then saved in an SQLite database for structured and easy access to data in the future. To increase reliability and performance, this project encompasses a number of vital functionalities. It captures dynamic content with the help of Playwright with the guarantee that elements present in the JavaScript will be rendered correctly. It also features error-handling for common problems, such as pages failing to load, losing connection, and data being stored in an unsatisfactory state. Moreover, the scraper logs all steps taken to help analyze the set of completed work and any mistakes that occurred within it. This project aids companies and researchers by automating the data extraction process from Lowe’s. ## An Overview of Libraries for Seamless Data Extraction ### Playwright Playwright is a powerful automation library, and users can control web browsers entirely via the power of code. It supports a rich browser API and is extremely well-suited for web scrapers that rely on content-heavy JavaScript rendering. For example, it makes easy interactive work with web pages--scrolling down through pages of content, clicking buttons or filling out forms, or taking screenshots of web pages just in time. ### Asyncio Asyncio is a standard library in Python which supports the concept of asynchronous programming. The asyncio library provides support to write concurrent code by using async and await syntax. In web scraping, Asyncio is important because it allows for running numerous requests in parallel to ensure increased efficiency of the process. Asyncio allows the scraper to do other work while waiting for responses from web servers. This makes the entire time taken to scrape a set of URLs reduced, which consequently makes the scraping process more responsive and faster. ### SQLite3 SQLite3 is a lightweight disk-based database that can be easily set up and used within Python applications. It is very useful for storing structured data, like the URLs and product details fetched during the scraping process. By using SQLite3, scrapers can create a persistent database that avoids duplicate entries and organizes data efficiently. In this way, it becomes easy to query and retrieve scraped information for further analysis or reporting. ### Random It is a native library of Python and can generate random numbers with the help of it so that users can choose their random elements. It can include random delay between successive requests for the website while performing web scraping. Thus, including random sleep time, web scrapers will look like human browsing instead of bots. So, they are less likely to be noticed by anti-bot mechanisms in the websites. This makes it possible to connect consistently with the target website, and hence one is sure to follow all its policies for usage. ### Beautiful Soup Beautiful Soup library uses Python for parsing HTML and XML documents. It provides a good and well-structured parse tree of the page source; hence searching document structures is easy. It scrapes through the web, using BeautifulSoup to fetch special elements from HTML pages, probably product titles, prices, descriptions, or specifications. Its syntax is very friendly for scrapers as it lets them query and manipulate the HTML content based on tags, attributes, and text in an efficient way that enables the extraction of accurate and reliable data. ## Why SQLite Outperforms CSV for Web Scraping Projects Given its simplicity, reliability and efficiency, SQLite is rather a good option to consider when it comes to the order of storing scraped data. Being self-sufficient, serverless and requiring no administrative tasks. It writes data on disks which is also useful in these types of tasks since a lot of information in this case URLs have to be saved and fast as well. This lightweight construction enables the user to fetch back the data at immense speeds in the case of complex queries rather than the use of CSV files that becomes difficult and slow with too much data set. ## STEP 1 :Product Link Scraping ### Importing Libraries ``` import asyncio from playwright.async_api import async_playwright import random import sqlite3 ``` This code imports essential libraries for web scraping and data handling. It includes requests for making HTTP requests, random and time for adding delays, sqlite3 for database interaction. ### Defining Base URLs for Categories ``` # Base URLs for each category BASE_URLS = { "security_cameras": "https://www.lowes.com/pl/home-security/security-surveillance-cameras/security-cameras/4294546211?offset={}", "smart_doorbells_locks": "https://www.lowes.com/pl/smart-home/smart-home-security/smart-doorbells-locks/37721669146465?offset={}", "smart_home_bundles": "https://www.lowes.com/pl/smart-home/smart-devices/smart-home-bundles/2311714614847?offset={}", "smart_light_bulbs": "https://www.lowes.com/pl/smart-home/smart-lighting/smart-light-bulbs/37721669146457?offset={}", "smart_speakers_displays": "https://www.lowes.com/pl/smart-home/smart-devices/smart-speakers-displays/2011455432077?offset={}", "home_alarms_sensors": "https://www.lowes.com/pl/home-security/home-alarms-sensors/1217527669?offset={}" } ``` This section of the code configures a dictionary named BASE\_URLS to store the base URLs for various categories of products at Lowe's. The URL for each category such as "security cameras," "smart doorbells and locks," etc is associated with it. Also, the URLs include an offset parameter (offset={}), to deal with pagination while web scraping multiple pages of products belonging to each category. In the code, later down the page, the base URL will be formatted with all offset values to navigate throughout the list of products for that category. This configuration allows dynamic page construction for different pages; this will make the scraping more scalable over different product categories. ### Setting Offset Values for Pagination ``` # Offset values for pagination in each category OFFSET_VALUES = { "security_cameras": [0, 24, 48, 72, 96], "smart_doorbells_locks": [0, 24, 48, 72, 96, 120, 144], "smart_home_bundles": [0, 24, 48, 72, 96], "smart_light_bulbs": [0, 24, 48], "smart_speakers_displays": [0], # Only one page for this category "home_alarms_sensors": [0, 24, 48] } ``` In the following section, a dictionary OFFSET\_VALUES is defined, in which offsets for pagination of all the categories of products are maintained. The Lowe's website displays a limited number of products per page, and these offset values are used to navigate through multiple pages. For example, the "security cameras" category has products spread across five pages, with each page displaying 24 products, hence the offset values of \[0, 24, 48, 72, 96\]. Other categories, like "smart speakers and displays," have only one page, so the offset value is \[0\]. These values will be appended to the URLs in the BASE\_URLS dictionary so that all product data are scraped from multiple pages for each category, so no data is missed. ### Base URL for Constructing Complete Product Links ``` # Base URL to prepend to relative URLs BASE_URL_PREFIX = "https://www.lowes.com" ``` Here is the definition of the constant BASE\_URL\_PREFIX that is the base URL for the Lowe's website. Sometimes, relative URLs (i.e., without the full domain) are scraped for product links. These relative URLs are prepended using BASE\_URL\_PREFIX so that all extracted links are complete and valid. By appending the relative paths to this base URL, the scraper can generate full product URLs that can be stored in the database or used for further data extraction. ### Setting Up the SQLite Database ``` # Database setup def setup_database(db_path): """ Set up the SQLite database for storing scraped product URLs. This function creates a connection to an SQLite database using the provided file path and initializes a table called `product_links` if it doesn't already exist. The table contains two columns: - `product_url`: Stores unique product URLs. - `status`: An integer that represents the processing status of each URL, which is initialized to 0 by default. Args: db_path (str): The file path of the SQLite database. Returns: sqlite3.Connection: A connection object to interact with the SQLite database. """ conn = sqlite3.connect(db_path) cursor = conn.cursor() # Create the table with product_url and status (default 0) cursor.execute(''' CREATE TABLE IF NOT EXISTS product_links ( product_url TEXT UNIQUE, status INTEGER DEFAULT 0 ) ''') conn.commit() return conn ``` The setup\_database function initializes the SQLite database used to store the scraped product URLs. It connects to the database from the provided file path (db\_path) and creates an empty table named product\_links that doesn't exist in that database. The table now contains two columns: namely, product\_url for unrepeatable product links as well as status that serves to monitor the processing status for each URL which is set 0 by default. The function returns a connection object that allows interaction with the SQLite database, which guarantees that scraped data will be efficiently stored and accessed for further use. ### Inserting URLs into the Database ``` # Insert URLs into the SQLite table, ignoring duplicates def insert_urls_to_db(conn, urls): """ Insert a list of product URLs into the SQLite database, ignoring duplicates. This function inserts each URL from the provided list into the `product_links` table. If a URL already exists in the table, it will be ignored to avoid duplicate entries. Args: conn (sqlite3.Connection): The connection object for the SQLite database. urls (list of str): A list of product URLs to be inserted into the database. Returns: None """ cursor = conn.cursor() cursor.executemany(''' INSERT OR IGNORE INTO product_links (product_url) VALUES (?) ''', [(url,) for url in urls]) conn.commit() ``` The insert\_urls\_to\_db function was used to insert a set of product URLs into the SQLite database, while ignoring duplication. This function takes two arguments-conn, which is the SQLite connection object used to interact with the database, and urls, the list of product URLs to insert. It uses a cursor to perform an INSERT OR IGNORE SQL statement, which tries to insert every URL into the product\_links table. The OR IGNORE clause would ensure that if the URL exists in the table, then it is skipped, thus preventing duplications. Finally, it commits the change using conn.commit(). This function does not return anything but handles the insertion of unique URLs into the database in an efficient manner. ### Scraping Product URLs from a Webpage ``` # Scrape URLs from a given page async def scrape_page(page, url, all_hrefs): """ Scrape product URLs from a given page and append them to the provided list. This function navigates to the specified URL using a Playwright `page` object, scrolls to the bottom of the page to load all dynamic content, and extracts URLs from `` tags within `

    ` elements. The URLs are then formatted with the base URL if they are relative links and appended to the provided list. Args: page (playwright.async_api.Page): The Playwright page object used for navigation and scraping. url (str): The URL of the page to be scraped. all_hrefs (list of str): A list to store the scraped product URLs. Returns: None Raises: Exception: Logs any error encountered during scraping and continues the process. """ try: await page.goto(url, timeout=60000) await page.wait_for_load_state('load') # Scroll to the bottom to load all content previous_height = await page.evaluate("document.body.scrollHeight") while True: await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await asyncio.sleep(2) current_height = await page.evaluate("document.body.scrollHeight") if current_height == previous_height: break previous_height = current_height await asyncio.sleep(5) # Ensuring all content is fully loaded # Locate tags within

    and extract href h3_links = await page.locator("h3 a").all() hrefs = [await link.get_attribute('href') for link in h3_links if await link.get_attribute('href') is not None] # Prepend the base URL to relative URLs complete_urls = [BASE_URL_PREFIX + href if not href.startswith('http') else href for href in hrefs] all_hrefs.extend(complete_urls) print(f"Scraped {len(complete_urls)} URLs from {url}") except Exception as e: print(f"Error scraping {url}: {e}") ``` The scrape\_page function is a method that scrapes product URLs from a given webpage by using Playwright's asynchronous API. It takes a Playwright page object, a target url, and a list all\_hrefs in which the scraped URLs will be stored. The function navigates to the page, waits for that page to load, and then scrolls down to the end of the page using execution of JavaScript to ensure that dynamic content is fully loaded, which it does by continuous scrolling and checking the page's height until the same height is shown twice; that is, no other content is loading. After this, it catches all tags inside

    tags and it returns the href attribute of all the of them, which are product URLs. If any of the foregoing are not absolute (i.e. do not have a complete path), it prepends BASE\_URL\_PREFIX to the beginning in order to make them absolute. All these URLs are put into the all\_hrefs list so they are all prepared for further processing. The function catches errors by logging them but continues to run, thus allowing robust scraping operations. Finally, it prints the number of URLs scraped from the given page, which helps track the scraping progress. ### Scraping URLs for Specific Categories and Inserting Them into the Database ``` # Process each category and insert URLs to the database async def process_category(browser, conn, category, base_url, offsets): """ Scrape product URLs for a specific category and insert them into the database. This function handles scraping product URLs for a specific category by iterating through pages using the provided offset values. For each page, it formats the base URL with the current offset, opens the page in a new Playwright browser tab, scrapes the URLs, and then closes the tab. After scraping all pages, the URLs are inserted into the SQLite database. A random delay is introduced between each page request to avoid overloading the server. Args: browser (playwright.async_api.Browser): The Playwright browser instance used for opening new pages. conn (sqlite3.Connection): The SQLite database connection object. category (str): The name of the category being scraped (for logging purposes). base_url (str): The base URL of the category with a placeholder for pagination offset. offsets (list of int): A list of offset values for pagination to iterate over pages. Returns: None """ all_hrefs = [] # To store scraped URLs for this category for offset in offsets: url = base_url.format(offset) page = await browser.new_page() await scrape_page(page, url, all_hrefs) await page.close() # Random delay before next request delay = random.uniform(2, 5) print(f"Waiting for {delay:.2f} seconds before the next request...") await asyncio.sleep(delay) # Insert URLs to the database after the category scraping is done insert_urls_to_db(conn, all_hrefs) ``` The process\_category function is responsible for scraping product URLs for a given category and inserting them into an SQLite database. It needs a Playwright browser instance, an SQLite connection conn, category name, a base\_url which contains pagination placeholder and offsets - the list of integers for page numbers to be scraped. In the list, it then proceeds to loop for each offset; format the base\_url for current pagination; and open new tab in the Playwright browser. It now calls the scrape\_page function to fetch URLs from the page and append them to all\_hrefs list. After scraping every page, it closes the tab and introduces a random delay between 2 and 5 seconds to prevent flood crashing. After parsing through the entire pages that concern category, all URLs fetched at all\_hrefs is added to the database for making further usage of this fetched URLs later using a function named as insert\_urls\_to\_db method which inserts the fetched URL into database in a safe way. It allows to loop through many pages for multi-structured categories . ### Running the Main Scraping Process for Multiple Categories ``` # Main function to run the scraping process async def run(): """ Main function to execute the web scraping process for multiple categories. This function orchestrates the entire scraping process by: 1. Setting up the SQLite database to store product URLs. 2. Launching a Playwright browser instance. 3. Iterating through predefined categories and scraping URLs from multiple pages using offset values for pagination. 4. Inserting the scraped URLs into the database. 5. Closing the browser and database connection after the scraping process is complete. Args: None Returns: None """ db_path = 'lowes_webscraping.db' # Set up database connection conn = setup_database(db_path) async with async_playwright() as p: browser = await p.chromium.launch(headless=False) # Scrape each category for category, base_url in BASE_URLS.items(): print(f"Scraping category: {category}") await process_category(browser, conn, category, base_url, OFFSET_VALUES[category]) await browser.close() # Close the database connection conn.close() # Run the asynchronous function asyncio.run(run()) ``` The run function is the orchestrator of the whole web scraping process. It first initializes the connection to an SQLite database through the setup\_database function with a database file called lowes\_webscraping.db. The database will store all product URLs scraped. It launches a Playwright browser instance, which can be used for automated interaction with web pages. The function prints the category name for every defined category in the BASE\_URLS dictionary and then calls process\_category to scrape URLs from multiple pages. The process\_category function handles navigation through every category's pages by predefined pagination offsets, scraping URLs, and inserting them into the database. After all categories have been processed, the browser is closed, and the database connection is terminated to ensure proper resource management. This design allows the scraping of many categories in sequence, all of whose URLs will be stored efficiently within one database, and manages the resources such as browser and database connection correctly throughout. ## STEP 2 :Detailed Product Data Scraping From Product Links ### Importing Libraries ``` import asyncio from playwright.async_api import async_playwright from bs4 import BeautifulSoup import random import sqlite3 ``` This code imports essential libraries for web scraping and data handling. It includes requests for making HTTP requests, random and time for adding delays, sqlite3 for database interaction, and BeautifulSoup from the bs4 library for parsing HTML content. ### Initializing the SQLite Database and Tables ``` # Initialize SQLite database and tables def init_db(): """ Initializes the SQLite database and creates necessary tables. This function establishes a connection to the SQLite database and creates two tables: `final_product_data` for storing successfully scraped product details, and `failed_urls` for logging URLs that could not be scraped, along with associated error messages. The `final_product_data` table includes the following columns: - id: Primary key, auto-incremented. - product_url: The URL of the product. - title: The title of the product. - price: The price of the product. - review_count: The number of reviews for the product. - rating: The rating of the product. - description: A brief description of the product. - specification: Additional specifications of the product. - features: Key features of the product. The `failed_urls` table includes the following columns: - id: Primary key, auto-incremented. - product_url: The URL of the product that failed to scrape. - error_message: The error message describing the reason for the failure. Returns: None """ conn = sqlite3.connect('lowes_webscraping.db') c = conn.cursor() c.execute(''' CREATE TABLE IF NOT EXISTS final_product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_url TEXT, title TEXT, price TEXT, review_count TEXT, rating TEXT, description TEXT, specification TEXT, features TEXT ) ''') c.execute(''' CREATE TABLE IF NOT EXISTS failed_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_url TEXT, error_message TEXT ) ''') conn.commit() conn.close() ``` This is the initialization function, init\_db intended to initialise an SQLite database exclusively used for storing all scraped data of products scraped from any website. Each time a connection is initialized with lowes\_webscraping.db, the required two tables are created namely final\_product\_data and failed\_urls in this case: the latter is used as a holding place for products which do not have web pages:. These include properties such as the URL of the product, title, price, review count, rating, description, specifications, and key features. The primary key for each record in this table is automatically incremented. On the other hand, the failed\_urls table contains any URL which could not be scraped and also an error message for explaining the cause of failure. That way, any problem which may occur in the process is recorded for further analysis. Both the tables are created only if they do not exist; this ensures the function will not raise errors if run multiple times because of the creation of a table that already exists. Lastly, the method commits all made changes to the database and closes the connection, thus completing the configuration of the database structure ### Inserting Scraped Data into the Database ``` # Function to insert scraped data into the database def insert_data(data): """ Inserts scraped product data into the SQLite database. This function connects to the SQLite database and inserts a new record into the `final_product_data` table. The record includes details about a product obtained from web scraping. Parameters: data (dict): A dictionary containing the product details to be inserted, which must include: - product_url (str): The URL of the product. - title (str): The title of the product. - price (str): The price of the product. - review_count (str): The number of reviews for the product. - rating (str): The rating of the product. - description (str): A brief description of the product. - specification (str): Additional specifications of the product. - features (str): Key features of the product. Returns: None: This function does not return a value. It commits the data to the database and closes the connection. Raises: sqlite3.Error: If an error occurs while interacting with the database. """ conn = sqlite3.connect('lowes_webscraping.db') c = conn.cursor() c.execute(''' INSERT INTO final_product_data (product_url, title, price, review_count, rating, description, specification, features) VALUES (?, ?, ?, ?, ?, ?, ?, ?) ''', (data['product_url'], data['title'], data['price'], data['review_count'], data['rating'], data['description'], data['specification'], data['features'])) conn.commit() conn.close() ``` The insert\_data function is used for inputting the scraped product data into the SQLite database. On execution, this function opens a connection to the lowes\_webscraping.db database and is ready to add a new record to the final\_product\_data table. The data that are to be inserted are passed into the function as a dictionary, and this dictionary is to have all the parameters inside it. These parameters include product's URL, title, price, number of reviews, ratings, description, specifications, and some key features. These data are then added by applying a parameterized SQL statement for the INSERT instruction provided. After executing the insertion command, the function commits the changes to the database in order to save the new record and then closes the connection. Any sqlite3.Error that might occur in the course of interacting with the database will raise an error, and therefore, error handling will occur appropriately in the insertion process. This is so because the product information scrapped will be saved in an appropriate manner for retrieval later for analysis. ### Logging Failed URLs into the Database ``` # Function to insert failed URLs into the database def insert_failed_url(url, error_message): """ Inserts failed URL data into the SQLite database. This function connects to the SQLite database and logs a failed URL along with an associated error message into the `failed_urls` table. This can be useful for debugging and tracking issues encountered during the web scraping process. Parameters: url (str): The URL of the product that failed to be scraped. error_message (str): A message describing the error that occurred while attempting to scrape the product. Returns: None: This function does not return a value. It commits the data to the database and closes the connection. Raises: sqlite3.Error: If an error occurs while interacting with the database. """ conn = sqlite3.connect('lowes_webscraping.db') c = conn.cursor() c.execute(''' INSERT INTO failed_urls (product_url, error_message) VALUES (?, ?) ''', (url, error_message)) conn.commit() conn.close() ``` The insert\_failed\_url function is designed to write the failure of URLs with their relevant error messages to the SQLite database. This logging method is important for debugging purposes and tracking problems that will occur during web scraping. When called, the function opens a connection to the lowes\_webscraping.db database and prepares it to insert a new row in the failed\_urls table. It takes two parameters: url, which is the URL of the product that was not able to be scraped, and an error\_message, which gives a description of the problem that occurred when trying to scrape. The function makes use of a parameterized SQL INSERT statement to add this data safely to the database, thus preventing SQL injection. After running the insertion command, it commits the changes to the database so that the logged failure is persisted, and then closes the connection. Like any other database interaction, if there is an error in this step, an sqlite3.Error is raised so that proper error handling can be done. This function plays a critical role in keeping a record of failures during scraping, which will help in debugging and also increase the reliability of the web scraping workflow. ### Updating URL Status After Successful Scraping ``` # Update URL status after successful scraping def update_url_status(url): """ Updates the scraping status of a product URL in the SQLite database. This function connects to the SQLite database and updates the `status` of a specific product URL in the `product_links` table to indicate that the scraping for this URL has been completed successfully. This is useful for tracking the progress of scraped URLs. Parameters: url (str): The URL of the product whose status is to be updated. Returns: None: This function does not return a value. It commits the changes to the database and closes the connection. Raises: sqlite3.Error: If an error occurs while interacting with the database. """ conn = sqlite3.connect('lowes_webscraping.db') c = conn.cursor() c.execute(''' UPDATE product_links SET status = 1 WHERE product_url = ? ''', (url,)) conn.commit() conn.close() ``` The update\_url\_status function is designed to update the status of a product URL in the SQLite database, more specifically in the product\_links table, to indicate that scraping for that URL has been successfully done. This functionality is important in tracking the progress of scraped URLs and maintaining an accurate record of which products have been processed. Whenever it is called, this function opens a connection to the lowes\_webscraping.db database and awaits sending a SQL UPDATE statement. It needs one parameter, and in this case, that will be 'url' of the particular product page which status is to be updated. It uses a parameterized query so as to not to allow any kind of SQL injections. On executing the UPDATE statement which sets the status field to 1 (meaning successful) it commits the transaction, saving the changes made on the database and closing its connection. In any of these database interactions having error, an sqlite3.Error raises, which can handle proper error management. The main function in this entire piece of code is about having visibility in the process of scraping and ensuring data integrity in the workflow of collection. ### Fetching Unscraped URLs from the Database ``` # Function to fetch unscraped URLs from the database (status=0) def get_unscraped_urls(): """ Fetches URLs from the database that have not yet been scraped. This function connects to the SQLite database and retrieves all product URLs from the `product_links` table where the `status` is set to 0, indicating that these URLs have not been scraped yet. Returns: list: A list of product URLs that have not been scraped. Raises: sqlite3.Error: If an error occurs while interacting with the database. """ conn = sqlite3.connect('lowes_webscraping.db') c = conn.cursor() c.execute(''' SELECT product_url FROM product_links WHERE status = 0 ''') urls = [row[0] for row in c.fetchall()] conn.close() return urls ``` The get\_unscraped\_urls function will obtain all the URLs of products in the SQLite database, which haven't been passed through the web scraping workflow yet. A connection to the lowes\_webscraping.db database will conduct a SQL SELECT statement against the product\_links table, selecting all those URLs where the status equals 0\. That means those haven't been scraped yet. The function uses a cursor to execute the query, then it fetches all rows returned. Then, it compiles the fetched URLs into a list comprehension to take the first element from every row, which coincidentally is the product URL. After retrieving the list of unscrewed URLs, the function closes the database connection to free resources and returns the list back to the caller. If there are any problems in dealing with the database, then an sqlite3.Error is raised so potential problems are handled correctly. This function is critical to managing the scraping process because it will allow the scraper to very efficiently identify which URLs still need processing. ### Extracting Features from HTML Content ``` # Process HTML content for extracting features def process_html_content(html_content): """ Extracts features from the provided HTML content using BeautifulSoup. This function parses the given HTML content to find all tables with the class 'TableWrapper-sc-ys35zb-0'. It extracts the key-value pairs from the table's rows and constructs a string representation of the features. The dictionary name is taken from the table's header (thead), and the key-value pairs are formatted as 'key:value'. Args: html_content (str): The HTML content to be processed. Returns: str: A string representation of the extracted features from the tables. If no tables are found, an empty string is returned. Raises: AttributeError: If the structure of the HTML content does not contain the expected table elements. """ soup = BeautifulSoup(html_content, 'html.parser') tables = soup.find_all('table', class_='TableWrapper-sc-ys35zb-0') features = [] for table in tables: thead = table.find('thead') dictionary_name = thead.get_text(strip=True) if thead else 'No dictionary name found' table_features = f"{dictionary_name}:{{" tbody = table.find('tbody') if tbody: key_value_pairs = [] for row in tbody.find_all('tr'): cells = row.find_all('td') if len(cells) > 1: key_value_pairs.append(f"{cells[0].get_text(strip=True)}:{cells[1].get_text(strip=True)}") table_features += ','.join(key_value_pairs) + "}" features.append(table_features) return ' '.join(features) ``` Process\_HTML\_Content is designed to present the features of content using BeautifulSoup given the HTML. For a content object, it should create a BeautifulSoup object to sufficiently parse the content before finding relevant features from the content. The function searches within each table element with a class TableWrapper-sc-ys35zb-0 since features will appear inside it in some sort of structured way. It then tries to locate the header of that table, or 'thead', to create a name of the dictionary which can define the contents of the particular table. In case there is no header found, it defaults to 'No dictionary name found'. This function checks the table for a body. That's where the key-value pairs would be located. The function loops through rows in the tbody to collect data from pairs by first and second cells for every row if more than two cells exist in a row. The obtained pairs are set in the key:value form and added to the list. After collecting the pairs for a table, those are concatenated into a string representing the features in the format dictionary\_name: {key:value,key:value}. Lastly, the function joins the strings of feature information from the tables into a single space-delimited string. If the processing does not contain any tables, the function returns an empty string. While parsing, it also raises an AttributeError when the structure of the HTML deviates from that expected by parsing; this then allows error handling in case parsing does not continue as one would hope. This function is very critical as it converts unstructured data from HTML into structured form for easier analysis or storage. ### Scrolling Through a Web Page for Dynamic Content Loading ``` # Scroll through the product page to load all content async def scroll_page(page): """ Scrolls through a webpage to load all dynamic content. This asynchronous function simulates scrolling down the page until no new content is loaded. It keeps scrolling until the height of the document remains constant, indicating that all content has been loaded. Args: page : The page object representing the webpage to scroll. Raises: Exception: Raises an exception if there is an issue with page evaluation or scrolling. Usage: await scroll_page(page) """ previous_height = await page.evaluate("document.body.scrollHeight") while True: await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") await asyncio.sleep(2) current_height = await page.evaluate("document.body.scrollHeight") if current_height == previous_height: break previous_height = current_height ``` The scroll\_page function is an asynchronous method that is designed to simulate scrolling across a webpage to load all content dynamically. When invoked, it first retrieves the initial height of the document's body using page.evaluate("document.body.scrollHeight"). This will be used as a comparison point to see if more contents are being loaded as scrolling through the page. The function then enters an infinite scroll down to the bottom of the page using the command window.scrollTo(0, document.body.scrollHeight). The function pauses for 2 seconds after scrolling using await asyncio.sleep(2) to allow loading time for the new content. This is how it looks after each scroll: function checks the current height of the document's body, if this height hasn't changed since the last one then the content hasn't been loaded in; in which case it quits the loop. If this is different then it updates the variable previous\_height to the new one, allowing it to keep on scrolling. This continues on till the end of all dynamic content loading. This function ensures that the operation of scrolling does not get in the way of the other tasks because asynchronous programming is used, making this function more efficient for dealing with web scraping processes, mainly on pages that rely on JavaScript to load more content. It is a part of web scraping workflows where complete content extraction is an important requirement for getting proper data collection. ### Extracting the Product Title from a Web Page ``` # Extract product title async def extract_title(page): """ Extracts the product title from a webpage. This asynchronous function retrieves the product title from the specified HTML element on the page. It searches for the element using a predefined CSS selector. If the title element is found, its text content is returned; otherwise, an empty string is returned. Args: page : The page object representing the webpage to extract the title from. Returns: str: The extracted product title if found, or an empty string if the title element does not exist. Usage: title = await extract_title(page) """ title_element = await page.query_selector('h1.styles__H1-sc-11vpuyu-0.krJSUv') if title_element: return (await title_element.text_content()).strip() return '' ``` The extract\_title function is an asynchronous function that fetches the title of a product from a given webpage library. On invocation, it accepts one argument: page, which is the webpage object whose title will be extracted. The first thing that is done in the function is to locate the product title element, which is identified by a defined CSS selector h1.styles\_\_H1-sc-11vpuyu-0.krJSUv. This is a defined CSS selector targeting the

    HTML element that has the product title. After that, it uses page object's query\_selector() method to locate this one. Once found, the function extracts text content from that one using text\_content() applied on the title element. The function will remove any leading or trailing whitespace from the extracted title with the strip() method before returning it. If the title element is not found on the page, the function returns an empty string. This is an extremely important function to web scraping workflows because this usually holds all the title information that may be needed for identification and categorization of a specific product. The encapsulation of this extraction logic in the context of an asynchronous function makes it efficient despite other tasks then running in parallel, waiting for the title extraction operation. ### Extracting the Review Count from a Product Web Page ``` # Extract review count async def extract_review_count(page): """ Extracts the review count from a product webpage. This asynchronous function retrieves the review count from the specified HTML element on the page. It searches for the element using a predefined CSS selector. If the review count element is found, its text content is returned; otherwise, an empty string is returned. Args: page: The page object representing the webpage to extract the review count from. Returns: str: The extracted review count if found, or an empty string if the review count element does not exist. Usage: review_count = await extract_review_count(page) """ reviews_element = await page.query_selector('div[data-testid="trigger"] span.sc-dhKdcB.kPpRJe') if reviews_element: return (await reviews_element.text_content()).strip() return '' ``` The function extract\_review\_count is an asynchronous function that is supposed to fetch the review count on a given product webpage. It only has one parameter: page; that is, it takes in a page object where extractivity of the review count will be performed. Inside the function, the action starts with finding a review count element by CSS selector: div\[data-testid="trigger"\] span.sc-dhKdcB.kPpRJe This selector finds a inside a
    that has a data-testid attribute. The function uses page object's query\_selector method to find this element. If the review count element exists, it will fetch text content using the text\_content() method then remove leading or trailing white spaces with the strip() method. This makes sure the review count returned is clean and well formatted. In a case where the review count element does not exist, the function returns an empty string. This function will go a long way in web scraping since it will enable collection of really important product feedback metrics for the sake of analysis, comparison, and understanding customer sentiment. Because it's an asynchronous function, other operations will continue to go on in parallel while the waiting for review count extraction runs its course. ### Extracting Product Rating from a Web Page ``` # Extract rating async def extract_rating(page): """ Extracts the rating from a product webpage. This asynchronous function retrieves the rating of a product from the specified HTML element on the page. It searches for the element using a predefined CSS selector. If the rating element is found, its 'aria-label' attribute is returned; otherwise, an empty string is returned. Args: page : The page object representing the webpage to extract the rating from. Returns: str: The extracted rating if found, or an empty string if the rating element does not exist. Usage: rating = await extract_rating(page) """ rating_element = await page.query_selector('div.sc-kAyceB.ijRZHV span[role="img"]') if rating_element: return (await rating_element.get_attribute('aria-label')).strip() return '' ``` This is an asynchronous function named extract\_rating, which pulls the rating of a product from a particular webpage. It requires one parameter: page. This is the webpage object from which it extracts the rating. The first thing this function does is find the rating element; a predefined CSS selector targets the rating information that is usually held in a , typically within a
    that carries certain class names. If it finds a rating element, it pulls out the value of its aria-label attribute. Typically this is text describing the rating as, for example, "4.5 out of 5 stars". The function removes any leading and trailing white space from that value before it returns. Otherwise, it returns an empty string-that is, it could not find a rating element on the page. This is one of the most important functionalities required for analysis on product performance based on user ratings. It will aid in better data collection, hence giving a deeper insight into the satisfaction of customers and quality of the product. ### Extracting Product Price from a Web Page ``` # Extract price async def extract_price(page): """ Extracts the price from a product webpage. This asynchronous function retrieves the price of a product from the specified HTML element on the page. It looks for the main price container and extracts both the dollar and cent parts of the price. If both parts are found, they are combined into a single string representing the final price; otherwise, an empty string is returned. Args: page : The page object representing the webpage to extract the price from. Returns: str: The extracted price if found, or an empty string if the price element does not exist. Usage: price = await extract_price(page) """ # Select the correct div that contains the price price_element = await page.query_selector('div.main-price.medium.split.split-left') if price_element: # Extract the dollar part and cent part of the price dollar_part = await price_element.query_selector('span.item-price-dollar') cent_part = await price_element.query_selector('span.PriceUIstyles__Cent-sc-14j12uk-0.bktBXX.item-price-cent') if dollar_part and cent_part: # Combine dollar and cent parts to form the final price return f"{(await dollar_part.text_content()).strip()}{(await cent_part.text_content()).strip()}" return '' ``` The extract\_price function is asynchronous, and it would take in the price of a given product from a given page. It accepts an argument; this argument will be a page object. This price will be fetched from that page object. Initially, this function would go ahead and select the div element which contains the basic information of price using a pre-defined CSS selector. Inside this block, it is looking for two specific elements of the price: the dollar and cent. It does this by requesting the elements containing these values - identified by their respective CSS classes. If it finds both parts of the price, the function returns their text content, after stripping any whitespace from around those values, then combines the two to make a single string that represents the full price. If either the dollar or the cent portion does not exist, it leaves the function returning a NULL string, meaning the price could not be pulled. This is extremely important for the purpose of accumulating accurate pricing information that would otherwise be used in further analyzing products and the competitive market for a product. ### Extracting Product Description from a Web Page ``` # Extract description async def extract_description(page): """ Extracts the product description from a product webpage. This asynchronous function retrieves the product description from the specified HTML element on the page. It locates the description within the appropriate div structure and returns the text content, stripped of any leading or trailing whitespace. If the description element is not found, an empty string is returned. Args: page : The page object representing the webpage from which to extract the description. Returns: str: The extracted product description if found, or an empty string if the description element does not exist. Usage: description = await extract_description(page) """ # Select the correct div that contains the description description_element = await page.query_selector('span.accordion-content.opened div.sc-esYiGF.bIlmEA.overviewWrapper div.romance') if description_element: # Return the text content of the description, stripped of any leading/trailing whitespace return (await description_element.text_content()).strip() return '' ``` The extract\_description function is an asynchronous function intended to fetch the product description from a given webpage. It takes one argument, page, which is the page object that will be used to fetch the description. The function starts by finding the correct element holding the product description by using a pre-defined CSS selector that correctly points to the required div structure. If the description element exists, the function outputs its text content after eradicating any leading and trailing whitespace to ensure a clean output. If the description element does not exist, the function would return an empty string, signaling that the description could not be extracted. This is essential for getting exhaustive product information and improving the quality of data gathered for further analysis or display. ### Extracting Product Specifications from a Web Page ``` # Extract specifications async def extract_specifications(page): """ Extracts product specifications from a product webpage. This asynchronous function retrieves the product specifications from the specified HTML element on the page. It locates the specifications within a list structure and returns the text content of each bullet point, joined into a single string with line breaks. If the specifications element is not found, an empty string is returned. Args: page : The page object representing the webpage from which to extract the specifications. Returns: str: The extracted product specifications if found, or an empty string if the specifications element does not exist. Usage: specifications = await extract_specifications(page) """ # Select the correct div that contains the specifications specs_element = await page.query_selector('span.accordion-content.opened div.specs ul.bullets') if specs_element: # Get all the bullet points (li) elements inside the specs list bullet_points = await specs_element.query_selector_all('li p') if bullet_points: # Extract the text content of each bullet point and return them as a joined string specs = [await li.text_content() for li in bullet_points] return '\n'.join(specs) return '' ``` The extract specifications function is asynchronous. It will collect the specifications of a product from the appropriate webpage. It accepts a page object containing data about the product as a single argument. The function looks first for the correct structure, which is an unordered list (
      ) wrapped around a containing a list of specifications that have a class of bullets. If the specifications element is successfully located, then the function collates all of the separate bullet points within
    • tags and captures the text content of each one. The separate text strings are concatenated together as one string with line breaks in between so that the resulting string can be easily readable. In case the element cannot be located, then the function will return an empty string; that means that no specification data exists. This method is able to retrieve detailed product specifications in an efficient manner, thus improving the overall data collection process for further analysis or presentation. ### Extracting All Product Details from a Web Page ``` # Extract all product details async def extract_product_details(page, data): """ Extracts all relevant product details from a product webpage. This asynchronous function aggregates multiple data extraction functions to retrieve essential product information from the specified HTML page. The extracted details are stored in the provided `data` dictionary. Args: page: The page object representing the webpage from which to extract the product details. data (dict): A dictionary that will hold the extracted product details. The function updates this dictionary with the following keys: - 'title' - 'review_count' - 'rating' - 'price' - 'description' - 'specification' Returns: None: The function modifies the `data` dictionary in place and does not return any value. """ data['title'] = await extract_title(page) data['review_count'] = await extract_review_count(page) data['rating'] = await extract_rating(page) data['price'] = await extract_price(page) data['description'] = await extract_description(page) data['specification'] = await extract_specifications(page) ``` It is there for the sake of getting complete product details from a webpage through various extraction functions. This is an asynchronous function that takes two parameters, namely page, holding the product's HTML content through the object, and data, a dictionary where all the details will be stored. The following function includes several asynchronous calls to particular extraction functions such as extract\_title, extract\_review\_count, extract\_rating, extract\_price, extract\_description, and extract\_specifications. All of these are meant to retrieve a specific piece of information, say, product title, review count, rating, price, description, or specifications. Then the results of those calls are stored within the data dictionary under its corresponding keys. This function does this pretty effectively by updating the dictionary in place, hence gathering all product details into one structure for easy access and further processing. It simplifies data collection workflows as it makes sure that the information one needs is fetched quite efficiently from the webpage at hand. It is pretty easy to use in larger data collection workflows because of the nature in which the function operates: instead of returning any value, it updates the dictionary directly. ### Scraping Individual Product Pages ``` # Scrape individual product pages async def scrape_page(page, url): """ Scrapes an individual product page for details and stores the data in a database. This asynchronous function navigates to the specified product URL, extracts relevant product details such as title, price, review count, rating, description, specifications, and features, and then saves the information to a database. If any errors occur during the scraping process, the function logs the URL and error message to a separate database table for failed URLs. Args: page : The page object representing the browser context to navigate and extract data from. url (str): The URL of the product page to scrape. Returns: None: The function modifies the data dictionary in place, inserts the scraped data into the database, and updates the URL status in the database. It does not return any value. Raises: Exception: If an error occurs during the scraping process, it is caught and logged. """ data = { 'product_url': url, 'title': '', 'price': '', 'review_count': '', 'rating': '', 'description': '', 'specification': '', 'features': '' } try: await page.goto(url, timeout=60000) await page.wait_for_load_state('load') await asyncio.sleep(2) # Scroll to the bottom await scroll_page(page) # Extract product details await extract_product_details(page, data) # Extract HTML content for additional features html_content = await page.content() data['features'] = process_html_content(html_content) # Insert data into the database insert_data(data) update_url_status(url) except Exception as e: print(f"Error scraping {url}: {e}") insert_failed_url(url, str(e)) ``` The scrape\_page function is defined with the aim of scraping the details on an e-commerce website regarding the product. Details scraped from the web page are written to a database. The function is asynchronous and depends on two parameters: it takes in a page, meaning the browser context that has enabled navigation and data extractions, and url- the specific URL of a product page that needs scraping. When this function is called it creates a dictionary called data and holds all the information there is about the product, it includes URL, title, price, review count, rating, description, specifications, and features. Afterwards, it tries to load the given URL by accessing the page object's method goto with a time-out of 60 seconds so it waits enough for the page to load. After ensuring the page is loaded, it introduces a short delay to let the loading process settle. Then it calls the function scroll\_page to scroll to the bottom of the page, which normally needs to be scrolled down to initialize the lazy loading of product details. The function calls extract\_product\_details passing both page and data dictionary for retrieving and storing the product information, retrieve the complete HTML content of the page which further is processed by the hypothetical process\_html\_content function for any other feature related to a product. All data having been collected, the algorithm begins by inserting the data in the database via insert\_data. The status of the URL is updated inside the database with the aid of update\_url\_status. If it catches an exception during the above step for reasons of network problems, or modification in the layout of the page, an error would occur that would then be captured and logged in, plus recording the failure of a certain URL within a distinct table using insert\_failed\_url function. This strong error handling enables proceeding with the scraping procedure even when there is an issue with certain pages, which makes it much more resilient and data-rich. ### Main Function to Control the Scraping Process ``` # Main function to control the scraping process async def run(): """ Controls the overall web scraping process, initializing the database, fetching unscraped URLs, and launching the browser to scrape each page. This asynchronous function is responsible for coordinating the scraping workflow by performing the following steps: 1. Initializes the SQLite database and creates the necessary tables. 2. Retrieves a list of product URLs that have not been scraped yet. 3. Launches a Chromium browser instance using Playwright. 4. Iterates over each URL, creating a new page for scraping. 5. Calls the `scrape_page` function to extract data from each product page. 6. Closes the page after scraping and waits for a random delay before proceeding to the next URL, simulating human-like browsing behavior. 7. Closes the browser once all URLs have been processed. Returns: None: This function does not return any value, as its primary purpose is to manage the scraping workflow. Raises: Exception: Any exceptions occurring during the scraping process will be raised and should be handled appropriately in the calling context. """ init_db() urls = get_unscraped_urls() async with async_playwright() as p: browser = await p.chromium.launch(headless=False) for url in urls: page = await browser.new_page() await scrape_page(page, url) await page.close() delay = random.uniform(2, 4) print(f"Waiting for {delay:.2f} seconds before navigating to the next URL...") await asyncio.sleep(delay) await browser.close() # Run the asynchronous function asyncio.run(run()) ``` The main run function orchestrates everything involving the web scraping process it performs necessary initialization, retrieves some URLs, and takes responsibility in scrapping each of those product pages. The execution is asynchronous, whereupon init\_db() will activate SQLite's in-memory database along with required tables to save off this scraping information. Then it accesses the list of remaining available but yet to be scraped of a list of product URLs made of calls upon the get\_unscraped\_urls() function. Lastly, it uses the async\_playwright() context manager from Playwright to create a new Chromium browser. This makes the headless parameter as False so the window is visible during scraping if any bug needs to be solved while creating a process. Then enters the function with an infinite loop iterating each of the URls fetched previously:. For each URL, a new page is created using browser.new\_page() and the scrape\_page function is called to extract relevant data from the product page. Once all the scrapes for each URL are done, the page is closed with await page.close(). The function also introduces a random delay between 2 and 4 seconds using random.uniform(2, 4) to mimic human-like browsing behavior, allowing the program to wait before moving on to the next URL. This pause is essential to avoid detection by the website's anti-scraping measures. Once all URLs are processed, the browser is closed by await browser.close(). There is no return of any value because the core objective is to manage overall scraping very efficiently. In case an exception occurs during the running of this function, then it must be raised and should, therefore be caught in an appropriate way in the calling context, hence making the scraping robust and reliable. Lastly, the function applied is called using asyncio.run(run()) to start an asynchronous event loop to run the scraping task:. ## Libraries and Versions This code utilizes several key libraries to perform web scraping and data processing. The versions of the libraries used in this project are as follows: BeautifulSoup4 (v4.12.3) for parsing HTML content, Requests (v2.32.3) for making HTTP requests, and Playwright (v1.46.0) for automating browser interactions. These versions ensure smooth integration and functionality throughout the scraping workflow. ## Conclusion The goal of the project has been accomplished as product data extraction from Lowe’s is fully automated via Playwright in Python, thus adhearing to the original scope of data automation. Dynamic content, pagination and errors are well managed, providing an overall robust solution for business and research purposes. The collected information lends itself perfectly for price monitoring, market research, and inventory management updating web scrapping to decision making's value chain. Connect with[ Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. ### How to Stay Anonymous While Web Scraping at Scale URL: https://www.blog.datahut.co/post/how-to-maintain-anonymity-when-web-scraping-at-scale-expert-tips/ Last updated: 2026-07-23T07:48:32.000Z Maintaining anonymity while web scraping at scale requires a combination of technical measures and strategic planning. Below is a structured guide covering proxy usage, request obfuscation, headless browsing with anti-detection, fingerprint avoidance, scalable infrastructure, defeating anti-scraping measures, and adaptive strategies. Each section includes best practices, examples, and tool recommendations to help keep your scraping activities anonymous and efficient. ## Managing Residential, Datacenter, & Rotating Proxies - to enable anonymity. Anonymizing web scraping activities demands the usage of proxies. Residential proxies route requests through real user devices (ISP-assigned IPs), making them appear as ordinary user traffic and offering high anonymity​. These are ideal for stealth since they originate from real networks and are more complicated to block than cloud server IPs. On the other hand, data center proxies come from cloud servers and are not affiliated with ISPs. They are cheaper and faster but easier for websites to flag as bots due to their recognizable IP ranges.​ For large-scale scraping, a rotating proxy setup is recommended – this means using a pool of proxy IPs that automatically change with each request or at set intervals. Rotating through many IP addresses helps distribute your requests and avoid rate-limit bans on any single IP​. Many proxy providers offer auto-rotating networks where each request is assigned a new IP or rotated after a certain time window. Proxy strategies: Use a mix of proxy types suited to your target site. For example, residential proxies excel at bypassing IP-based blocks​, while data center proxies can be helpful for high-volume scraping if the target isn’t strict about IP reputation. Ensure your proxy solution supports geolocation targeting if you need to appear from specific countries (many residential proxy services let you choose regions). If possible, prefer sticky sessions (keeping the same IP for a short duration when needed) for tasks like multi-page navigation or logins, but still rotate IPs periodically to avoid long-term profiling. Constantly monitor proxy health and avoid free or public proxies – those are often shared, slow, and quickly banned. Premium proxies with large pools of IPs are worth the investment for serious projects.​ Your script should randomly rotate through a list of proxies and user agents. Each request is sent from a different IP and with a different client identity, reducing the chance of detection. Adjust the timing and rotation strategy as needed (e.g., rotate on each request or after a fixed number of requests per IP). Proper error handling (e.g., retrying with a new proxy on timeouts) should be added for robustness. ## Request Obfuscation (Headers, User Agents, Cookies, Timing) Web servers often detect scrapers by their network request patterns and headers. To blend in with regular traffic, your scraper’s HTTP requests must look as realistic as possible. This involves randomizing and faking certain parts of the request: User-Agent strings: Always supply a User-Agent header imitating a common browser (Chrome, Firefox, mobile Safari, etc.). Don’t use the default ones from HTTP libraries, as those are obvious (e.g., Python’s requests default is python-requests/2.x, a dead giveaway​). Maintain a pool of modern User-Agent strings and rotate them so each request isn’t identical​. Ensure the User-Agent and other headers match (for instance, a Chrome User-Agent should be accompanied by typical Chrome headers). HTTP Headers: Real browsers send a variety of headers like Accept, Accept-Language, Accept-Encoding, Connection, and sometimes Referer. Simulate these as needed. For example, include an Accept-Language header (e.g., “en-US,en;q=0.9”) and an Accept header matching browsers (text/html,application/xhtml+xml,...)​ . Many scrapers get caught by sending too few or inconsistent headers. Compare a browser’s headers to your scraper’s using a service like httpbin.org to ensure you have all the standard fields.​ Cookies and Sessions: Use cookies like a regular user. For example, handle Set-Cookie headers and resend them on subsequent requests using a session object. This can make your scraper seem like a repeat visitor rather than a fresh client on every request. Randomize or clear cookies when starting a completely new session or when switching identities. Some anti-bot systems track cookie consistency; completely blocking cookies can raise suspicion, so it’s often better to accept and use them as a browser would. Request timing & ordering: Avoid making requests in a perfectly periodic or fast manner. Humans have irregular browsing patterns. Introduce random delays between requests (as shown in the code above) to mimic human pacing​. Also, avoid constantly hitting pages in the exact same sequence or frequency. If possible, shuffle the order in which you scrape pages or inject occasional pauses. For example, instead of scraping 1000 pages in one burst, you might scrape in smaller batches with breaks. Vary the time of day your scraper runs if it’s a continuous process (e.g., not every day at precisely 00:00)​. These tactics help defeat simple rate-based IP bans and more advanced behavioral analysis. Referer and Navigation simulation: If feasible, sometimes set the Referer header to a logical previous page (e.g., if scraping a product page, set the referer to the category page). This isn’t always necessary, but on some sites it can make your traffic pattern resemble a user clicking through links rather than a bot directly fetching every page. Similarly, performing a search on the site and then navigating to items (when possible) can emulate user behavior. Always verify the request your code produces – it should closely resemble a real browser’s request. By setting realistic headers and varying them, you make it much harder for a site to filter out your scraper based on “odd” HTTP signatures. ## Headless Browsing (Selenium, Puppeteer, Playwright & Bot Evasion) When target websites employ heavy JavaScript or advanced anti-scraping measures, using a headless browser can help your scraper behave more like a real user. Headless automation tools like Selenium (Python, Java), Puppeteer (Node.js), and Playwright allow your scraper to load pages just as a browser would, running all scripts and rendering content. They enable you to simulate human actions such as clicking buttons, scrolling, filling out forms, and navigating complex sites​. This is crucial for sites that require interaction (e.g., clicking “Load more” or logging in) or that deliberately delay content rendering to foil simple scrapers. However, using a headless browser by itself doesn’t guarantee anonymity. Many websites use scripts to detect automation. Common giveaways include the navigator.webdriver flag (which is true in Selenium by default), the absence of typical plugins, or the HeadlessChrome substring in the User-Agent. To avoid these, developers use stealth techniques: Stealth plugins and patches: For Puppeteer, the popular puppeteer-extra-plugin-stealth plugin automatically fixes many headless tells (it masks webdriver, modifies APIs to expected values, etc.). Similarly, Selenium users can employ libraries like selenium-stealth or undetected-chrome driver, which launch Chrome in a stealthy mode (patching WebDriver flag, disabling automation extensions, and more)​ Playwright has a stealth library as well. By integrating these, a headless browser can appear nearly identical to a real browser. For example, adding stealth(driver, ...) in Selenium will adjust attributes (languages, vendor, platform, etc.) to mimic a regular Chrome on Windows. Customizing headless behavior: You can manually tweak the browser context. This includes setting a realistic User-Agent (headless tools allow this), enabling graphics (some detection scripts check for WebGL fingerprints), and even injecting scripts to spoof functions. For instance, you might preload a script to override navigator.permissions.query to not reveal automation. There are open-source scripts on GitHub (like Browserleaks or stealth.min.js from puppeteer-extra) that address these fingerprint points. Headless vs Headful: If stealth mode isn’t enough, an alternative is running in non-headless mode (a full browser) controlled by automation. This way, the native browser is precisely as a user would have it. You can run a full Chrome/Firefox in a virtual display or sandbox, so it’s not visible but still not in headless mode. This eliminates headless-only flags. It’s more resource-intensive but sometimes necessary for sites with aggressive bot detection. Using headless browsers effectively allows you to scrape content that isn’t reachable with simple HTTP requests, and combined with stealth tactics; you can evade many bot detection systems. Just be mindful that headless browsers consume more CPU and memory, so you’ll need to scale your infrastructure accordingly (discussed in a later section). ## Anti-Fingerprinting Methods (Preventing Browser Fingerprinting) Websites increasingly rely on browser fingerprinting to identify bots. Fingerprinting involves collecting dozens of environment attributes – like your screen resolution, OS, time zone, installed fonts, canvas/WebGL rendering data, and even subtle TLS handshake traits – to form a unique “fingerprint” of your device.​ If your scraper’s fingerprint remains constant or has values that no real user’s browser would have, anti-bot systems can latch onto that and block you. To combat fingerprinting, you have two main approaches: using specialized anti-detect browsers or environments, and spoofing or randomizing key fingerprint data. Anti-detect Browsers & Multi-Profile Tools: There are tools to create isolated browser profiles that each have a distinct fingerprint. These tools basically run actual browser instances (Chrome or Firefox forks) but feed them fake profile data – e.g., one profile might simulate Windows 10 with Chrome at 1920x1080, and another might be an Android phone with Chrome Mobile​. They also often integrate proxy management for each profile. Using such a platform, you can manage many virtual “identities” for your scraper, each appearing on websites as a different user. Multilogin, for instance, masks your digital fingerprint and ensures each browser profile looks like an actual, unique device.​ These tools cost money but are very powerful – they handle the low-level spoofing of canvas, audio context, fonts, and so on, which is difficult to do manually. If you’re conducting large-scale scraping or managing multiple accounts, an anti-detect browser can be a one-stop solution to prevent cross-site and cross-session fingerprint linking. Manual fingerprint spoofing: If using standard tools like Selenium or Puppeteer, you can still spoof many fingerprint components manually. Some examples: override the Canvas API to return a constant (or randomized) image so that canvas fingerprinting can’t track you; similarly, override the WebGL renderer info to match your claimed device. Adjust your browser’s reported timezone and locale to match your proxy’s region. Randomize plugin lists or use a plugin to generate believable values. For TLS fingerprinting (server-side TLS client hello profiling), tools like Curl-Impersonate can mimic the TLS signature of real browsers.​ In fact, headless tools often have slightly different TLS handshakes; using an open-source library to impersonate Chrome/Firefox at the TLS level can close that gap. These manual methods require significant effort and testing but can be script-automated if you have to roll your own solution. Match proxies to fingerprints: A clever detail is to match your proxy IP’s properties to your claimed client. For instance, if your browser profile claims to be an Android phone, use a mobile proxy (an IP from a cellular network) so that everything aligns.​ If you pretend to be a user in Germany, use a German IP. Mismatched signals (e.g., a “Chrome Windows” fingerprint coming from a data center in another country) could raise suspicion. Many anti-detect platforms let you bind a proxy to a profile and even adjust the fingerprint to fit the IP’s geolocation. In summary, anti-fingerprinting is about making your automated browser look like a unique human user and doing so consistently. High-end anti-bot systems might track dozens of parameters – you need to either suppress those (block them from being read, which is hard without breaking site functionality) or fake them in a believable way. Using established anti-detect tools is often the straightforward path. As a simpler stop-gap, ensure each scraping bot at least has a different User-Agent, screen size, and IP, and consider resetting or randomizing the fingerprint occasionally so it doesn’t build a repetitive history. Remember that even with perfect technical spoofing, unusual behavior (too-fast navigation, no real mouse movement, etc.) can still give away a bot – so combine fingerprint avoidance with the behavioral tactics discussed below. ## Anonymity Infrastructure for Large-Scale Scraping When scraping at scale, infrastructure plays a big role in both efficiency and anonymity. A single machine or IP is not enough; you’ll want to distribute the load and design a robust system. Key considerations include the distribution of tasks, concurrency control, and fault tolerance, all while keeping your identity hidden. Distributed scraping: Instead of one process doing all the work, use multiple worker processes or servers to run scrapers in parallel. A distributed architecture can handle more volume and also allows you to originate traffic from many places (a plus for anonymity). For example, you might deploy scraper instances on cloud servers in multiple regions, each using its own set of proxies. This horizontal scaling improves throughput and avoids bottlenecks. Anonymity infrastructure: From a privacy standpoint, having distributed infrastructure means you should also distribute your anonymity tools. For example, use different proxy pools on different nodes to avoid correlation. You might even use multiple proxy providers (one node using Provider A’s IPs, another using Provider B) so no single provider sees all your traffic. Containerization can help here by bundling distinct proxy credentials per container. Also, monitor each node’s IP reputation – sometimes, entire cloud regions get temp-banned by a site, in which case shifting your workload to other regions (or using residential proxies on those nodes) is a solution. In short, design your scraper like a scalable service. A possible architecture: a central scheduler service assigns URLs to a fleet of workers; each worker runs in a container/VM with its own proxy configuration and scraping logic; they report back data to a central database. This setup can handle failure gracefully (if one worker IP is banned, tasks can be retried on another), and it’s horizontally scalable. Just remember that scaling up also means scaling your anti-detection measures across the board. ## Handling CAPTCHAs and Anti-scraping systems while maintaining anonymity Websites employ various anti-scraping defenses. To maintain anonymity and scraping efficiency, you need tactics to detect and bypass these mechanisms: CAPTCHAs: CAPTCHAs (“Completely Automated Public Turing test to tell Computers and Humans Apart”) are challenges like image selections or puzzles designed to stump bots. If your scraper hits a CAPTCHA, one approach is to solve it using external services. There are Captcha resolvers offering API-based solving: they farm out the challenge to human solvers and return the answer to you.​ This can be integrated into your scraper (e.g., send the image to Captcha Resolver API and get back the text). However, solving CAPTCHAs costs money and adds delay (15-30+ seconds), and at large scale it may be too slow and pricey. The preferred strategy is to avoid triggering CAPTCHAs in the first place.​ Many CAPTCHAs are deployed selectively – e.g., after a certain number of rapid requests or on suspicious patterns. By using the techniques discussed (rotating IPs, realistic headers, slow pacing), you can often scrape under the radar such that CAPTCHAs don’t appear. If a site uses something like Cloudflare, which throws CAPTCHA/JS challenges by default, you might employ a headless browser to navigate it (Cloudflare’s bot screen can be passed by a real browser solving the JS challenge). There are also specialized anti-CAPTCHA tools (e.g., Cloudflare solver in Python like cloudscraper. JavaScript challenges and bot detection scripts: Anti-bot providers (Cloudflare, Akamai, DataDome, PerimeterX, etc.) use scripts that run in the browser to analyze behavior and environment. They might collect a fingerprint, observe how quickly the page is rendered, and even simulate user interactions to see if they occur. Bypassing these often requires executing the JS (hence using headless browsers) and possibly adding delays or interactions. For example, some challenges check that certain events (like mouse movements) happen. Tools like Puppeteer/Playwright allow you to generate fake mouse movements or scroll events to satisfy these checks. If you encounter an anti-bot wall, examine what it’s looking for – browser dev tools or network logs can show if there’s a specific script handing out tokens that you need to replicate. ## Adaptive Strategies (Monitoring Detection & Dynamic Evasion) Even with all precautions, you must be prepared to adapt on the fly. A hallmark of successful large-scale scraping is continuous monitoring of your scraper’s performance and the target’s reactions and adjusting accordingly: - Detecting detection (!): Build in checks for signs that your scraper has been spotted. This could be an increase in HTTP error codes (403 Forbidden, 429 Too Many Requests), receiving CAPTCHA pages instead of content, or getting served misleading content (some sites present fake data to suspected scrapers). Implement logic to recognize these. For example, if the page content contains phrases like “verify you are human” or has Blocked, flag that. Keep an eye on unexpected redirects or login prompts for pages that shouldn’t require login. If you use headless browsers, watch for alert popups or specific DOM changes that indicate a bot challenge. Logging is crucial here: keep logs of requests and responses (or at least response codes). - Auto-adjust rate and patterns: If your monitoring shows a lot of failures or blocks, have your system back off automatically. You can incorporate an exponential backoff algorithm: when encountering a block, wait a bit and retry; if blocked again, wait longer, etc.​ This helps in two ways – it reduces pressure on the site (lessening the suspicious activity) and gives time for any temporary IP bans to possibly lift. Likewise, if a particular proxy IP gets banned, stop using it immediately and switch to a fresh one (and maybe don’t return to the bad one for many hours). Ideally, your scraper fleet can dynamically drop and replace proxies that go bad. - Dynamic proxy/user-agent switching: In an adaptive system, you might maintain multiple identities. If identity A starts getting blocked frequently, switch to identity B (different user agent, different proxy pool). This is akin to a getaway car – don’t keep banging on the front door when you’ve been seen; try a different approach. Some scrapers even cycle through a set of user profiles per session. - Monitor site changes: Websites can change their HTML structure or anti-bot measures at any time. If your extraction logic suddenly fails (e.g., CSS selectors no longer find the data), the site might have redesigned or introduced new obfuscation (like randomly generated element IDs). Use automated tests or minor pilot scrapes to detect layout changes. For example, track the number of data items found or use assertions; if they drop to zero, raise an alert. ZenRows recommends monitoring for changes in the site’s structure and adjusting your parser accordingly.​ Being adaptable means your scraper can be quickly updated to handle such changes — sometimes even automatically if you can make your parsing logic flexible (e.g., using text cues in addition to fixed XPaths). - Notifications and fail-safes: Set up alerts when certain thresholds are met, such as: X% of recent requests were blocked, or the scrape rate dropped dramatically, or unusual content detected. This allows you (or your system) to intervene promptly. A fail-safe could be to temporarily pause scraping that site when too many blocks occur, to avoid burning all your proxies unnecessarily and to give the site time to cool down suspicion. - Continual improvement: Treat the scraping operation as iterative. Each time you get blocked, analyze how and why. Maybe you notice the site started checking for a particular header – you can then add that to your requests (cat-and-mouse). Or you realize all your proxies from a specific subnet got banned – perhaps avoid clustering requests via similar IP ranges. Use insights from failures to update your strategy. It’s helpful to keep a knowledge base of what anti-bot techniques each target site employs so you can anticipate problems. For instance, if you know Site A uses Akamai (which might fingerprint aggressively), you’ll prioritize headless+stealth for it, whereas Site B has basic IP rate limiting, so you rely more on proxy rotation there. In summary, an adaptive scraper is one that monitors itself and the target and can modify its behavior without a complete manual overhaul. It’s like having a stealth vehicle that changes its route if it senses roadblocks ahead. This can mean the difference between a scraper that works for one week and then gets shut out and one that runs for months continuously. By combining adaptive techniques with all the prior measures (proxies, obfuscation, headless stealth, etc.), you create a resilient scraping system that maintains anonymity and effectiveness even as targets evolve their defenses. ## Conclusion: Anonymity in large-scale web scraping is achievable by layering multiple strategies. Use proxies to hide your IP, disguise your requests to look human, leverage headless browsers with stealth to bypass interactive checks, and vary your fingerprint. All the while, stay within legal and ethical bounds and design a scalable, monitorable system that can adjust as needed. With careful implementation, your scrapers can collect vast amounts of data while flying under the radar of anti-bot detection.​ Always remember that this is an arms race – keep learning and refining your techniques as websites deploy new defenses. Happy (responsible) scraping! Related posts ### 5 FAQs 1\. Why is anonymity important when web scraping at scale? Maintaining anonymity helps prevent IP bans, ensures scraping continuity, and reduces the risk of detection. It protects your scraping infrastructure from being flagged or blocked by websites using anti-bot systems. 2\. What tools can help maintain anonymity in web scraping? Tools like proxy rotators, VPNs, and headless browsers (Playwright, Puppeteer) help disguise your scraping activities. Services that offer residential or rotating proxies also add an extra layer of anonymity. 3\. How do rotating proxies help in anonymous web scraping? Rotating proxies assign a different IP for each request or session, mimicking human browsing behavior. This prevents websites from identifying and blocking requests from a single IP source. 4\. What are the common mistakes that can expose your scraper’s identity? Using the same IP for multiple requests, sending requests too fast, failing to randomize user agents, or ignoring robots.txt directives can all reveal your scraper’s identity and lead to IP bans. 5\. How can you balance anonymity with ethical web scraping? Always respect website terms of service, scrape publicly available data, avoid personal or sensitive information, and ensure compliance with privacy laws like GDPR. Ethical scraping practices build credibility and reduce legal risks. ### Why Amazon Sellers Must Analyze Competitor Reviews URL: https://www.blog.datahut.co/post/why-scrape-competitor-amazon-reviews/ Last updated: 2026-07-23T07:48:32.000Z If you're an Amazon seller, you spend hours trying to make your product stand out—tweaking descriptions, adjusting pricing, and running ads. But what if there were a smarter way to gain a competitive edge? The answer? Your competitors' reviews. Think about it—your potential customers are already telling you exactly what they love and hate about products in your niche. The problem is going through thousands of reviews manually is impossible. That’s where Amazon review scraping comes in. In this post, we'll discuss how analyzing your competitors' reviews can boost your sales, how to do it legally, and how Datahut can help you automate the process. ## The Power of Amazon Reviews in Influencing Consumer Decisions Before moving into the technical aspects of scraping, it’s essential to understand why Amazon reviews are so important. According to a recent study by PowerReviews, 99.9% of consumers read reviews when shopping online, and 96% of consumers specifically look for negative reviews. These statistics highlight the importance of reviews in influencing purchasing decisions. Positive reviews can boost sales, while negative reviews can provide valuable feedback for product improvement. For Amazon sellers, reviews are not just a reflection of customer satisfaction—they are a direct line to understanding what customers love or hate about a product. By scraping competitor reviews, you can gain insights into what customers say about similar products, identify gaps in the market, and refine your offerings to meet customer needs better. ## Industry-Specific Use Cases for Scraping Amazon Reviews Scraping competitor reviews isn’t a one-size-fits-all strategy. Different industries can benefit from this practice in unique ways. 1\. Consumer Electronics Customers often leave detailed reviews about specific product features in the consumer electronics industry. By scraping these reviews, you can identify which features are most appreciated and which ones are causing frustration. For example, if customers consistently complain about the battery life of a competitor’s smartwatch, you can highlight your product’s superior battery performance in your marketing campaigns. Scraping reviews can also reveal insights into how customers interact with software interfaces, connectivity features, or durability aspects. If customers frequently mention difficulties setting up a competitor’s smart home device, you can simplify your product’s setup process and emphasize this in your product descriptions. Additionally, scraping reviews can help you identify recurring issues like overheating, slow performance, or compatibility problems, allowing you to address these concerns in your product design and marketing. This ensures your product meets and exceeds customer expectations in a highly competitive market. 2\. Fashion In the fashion industry, trends change rapidly, and customer preferences can be fickle. Scraping reviews can help you identify emerging trends and understand what customers want regarding style, fit, and fabric. For instance, if customers are raving about a particular type of fabric in a competitor’s product, you can consider incorporating similar materials into your designs. Reviews can also provide insights into sizing issues, such as whether a product runs large or small and how customers feel about the fit of certain styles. If customers frequently mention that a competitor’s jeans are too tight around the waist, you can adjust your sizing charts or design to offer a more comfortable fit. Additionally, scraping reviews can reveal color preferences, styling tips, or packaging feedback. For example, if customers appreciate eco-friendly packaging in a competitor’s product, you can adopt similar practices to appeal to environmentally conscious buyers. By leveraging these insights, you can tailor your product lines to meet customer expectations better and stay ahead of fashion trends. 3\. FMCG For fast-moving consumer goods (FMCG), customer preferences vary widely based on taste, packaging, and price. Scraping reviews can help you understand what customers value most in these products. For example, if customers frequently mention that they prefer eco-friendly packaging, you can make sustainability a key selling point for your products. Reviews can also highlight specific flavors, textures, or ingredients that customers love or dislike, allowing you to refine your product formulations. If customers consistently praise a competitor’s coffee for its rich flavor but complain about its high price, you can offer a similarly high-quality product at a more competitive price. Additionally, scraping reviews can reveal insights into packaging convenience, such as resealable bags or portion-controlled servings, influencing purchasing decisions. If customers frequently mention that a competitor’s snack packaging is difficult to open, you can design your packaging to be more user-friendly. By analyzing competitor reviews, you can also identify pricing sensitivities or promotional strategies that resonate with customers, helping you optimize your pricing and marketing efforts to better compete in the FMCG space. ## Gaining a Competitive Edge Through Scraped Reviews Scraping competitor reviews isn’t just about gathering data—it’s about staying ahead of the competition. Here’s how you can use review insights to your advantage: ### Spot Competitor Weaknesses and Capitalize on Them If customers repeatedly complain about something—like poor battery life, flimsy materials, or bad customer service—you can do better. Highlight those improvements in your product descriptions and marketing to attract frustrated customers looking for a better option. ### Understand What Customers Want Sometimes, what you think customers want and what they want are two different things. By analyzing reviews, you can see recurring requests—like better packaging, more size options, or improved durability—and use that information to upgrade your products before your competitors do. ### Fine-Tune Your Marketing Strategy Amazon is flooded with similar products, so positioning matters. If reviews show that people love a certain feature in a competitor’s product, highlight that feature in your marketing. Conversely, if they hate something, turn that into your unique selling point. Example: If reviews show customers complaining about confusing assembly instructions for a competitor’s standing desk, you can advertise yours as “Easy Assembly—No Tools Required” and include a short demo video. ### Predict Market Trends Before They Peak Review scraping helps you see patterns before they explode. If increasing customers mention a specific material, design, or feature, that’s a sign the market is shifting. You can use this information to adjust your inventory and be one of the first sellers offering exactly what people want. Example: If reviews for running shoes mention a preference for lightweight, breathable materials, you can prepare your next stock accordingly, giving you a head start on the trend. ### Improve Customer Satisfaction & Retention Understanding competitor reviews doesn’t just help you win new customers—it helps you keep them. If you proactively address common pain points (like slow shipping or poor durability), customers will be less likely to leave negative reviews on your products. Happy customers lead to higher ratings, better rankings, and more repeat buyers. ### Case Study A pet products company specializing in pet carriers faced challenges on Amazon in 2018\. Despite offering a high-quality product, their carrier was ranked below 10,000 in the "Pet Supplies" category, with a modest 3.8-star rating. Sales were stagnant, advertising efforts were ineffective, and product descriptions failed to resonate with customers. The company needed a strategy to improve visibility, customer satisfaction, and sales. #### Challenge The company struggled to understand why competitors were outperforming them. Their product was durable and well-made, but customers weren’t connecting with it. They needed to identify what customers truly valued in pet carriers and adjust their product and marketing strategies accordingly. #### Solution The company decided to analyze over 25,000 reviews from its top five competitors. This analysis revealed three key customer priorities: 1. Airline Approval: Customers valued carriers that were certified for airline travel, as it made traveling with pets easier. 2. Air Circulation: Poor ventilation was a common complaint in negative reviews, indicating a significant pain point. 3. Dual Access: Many customers preferred carriers that could be opened from both the top and side, a feature rarely highlighted by competitors. Based on these insights, the company made several strategic changes: - They obtained airline certification for their carriers and prominently featured this in their product descriptions. - They redesigned the ventilation system to address air circulation issues. - They added a top door to their carriers for easier access. - They rewrote their product descriptions to focus on the features customers cared about most. #### Results The changes had a dramatic impact. Within three months, the company’s pet carrier jumped into the top 1,000 products in the "Pet Supplies" category. Their average rating improved from 3.8 to 4.6 stars, and sales more than tripled. Returns dropped significantly, and their advertising campaigns became more effective as they aligned with customer priorities. Customer reviews reflected the improvements. One customer wrote, "Finally, an airline-approved carrier without any hassle! The ventilation is amazing compared to other carriers I've tried." This feedback validated the company’s efforts and reinforced its commitment to meeting customer needs. #### Continuous Improvement The company didn’t stop at its initial success. They continued to monitor customer feedback and market trends. When eco-friendly materials gained traction, they incorporated recycled materials into their carriers, further solidifying their position as a forward-thinking brand. #### Key Takeaways This case study highlights leveraging customer reviews to drive business decisions. Key lessons include: - Customer-Centric Approach: Understanding and addressing customer pain points can significantly improve product design and marketing. - Competitor Analysis: Analyzing competitor reviews can reveal unmet customer needs and opportunities for differentiation. - Adaptability: Success on Amazon requires continuous monitoring of customer feedback and market trends and the willingness to adapt quickly. #### Conclusion The pet products company transformed its Amazon business by analyzing competitor reviews and focusing on customer priorities. Their story demonstrates the power of data-driven decision-making and the importance of listening to customers. For Amazon sellers, this case study is a valuable example of how competitor review analysis can lead to improved product offerings, increased sales, and long-term success. ## Sentiment Analysis Scraped reviews don’t just tell you what customers say; they also reveal how they feel. Sentiment analysis helps you measure whether customer feedback is positive, negative, or neutral—giving you a clear picture of how your brand (or your competitors) is perceived. ### How Sentiment Analysis Works Sentiment analysis uses Natural Language Processing (NLP) to determine the emotional tone behind words. By analyzing patterns in language, we can categorize reviews into positive, neutral, or negative sentiments. For example: - Positive review: “This smartwatch has an amazing battery life and a sleek design!” - Neutral review: “The smartwatch is okay, but nothing special.” - Negative review: “Battery drains too fast, and the strap is uncomfortable.” ### Tools & Libraries for Sentiment Analysis If you want to perform sentiment analysis on scraped Amazon reviews, here are some tools to help: - VADER (Valence Aware Dictionary and sentiment Reasoner): Great for analyzing short texts like social media posts and reviews. It assigns sentiment scores to words and phrases, making it ideal for Amazon reviews. - TextBlob: A beginner-friendly Python library that simplifies sentiment analysis. It provides a polarity score that tells whether a review is positive, negative, or neutral. - NLTK (Natural Language Toolkit): A comprehensive NLP library that allows deeper text processing, including tokenization, sentiment analysis, and more. ### How Sentiment Analysis Helps Amazon Sellers Track Customer Satisfaction Trends: By analyzing reviews over time, you can spot patterns in customer sentiment. If sentiment suddenly shifts from positive to negative, it might indicate a drop in quality, a new competitor gaining traction, or an external issue like shipping delays. Keeping an eye on these trends helps you react swiftly and maintain customer trust. Identify Negative Patterns in Competitor Products: Sentiment analysis allows you to pinpoint consistent issues in competitor reviews. If multiple customers mention the same complaint—such as poor durability or a misleading product description—you can use that knowledge to improve your product and highlight those enhancements in your marketing. Improve Product Development: If many customers complain about a feature, you can refine your next product version accordingly. For instance, if customers frequently mention that a laptop bag is too small for larger laptops, you can introduce a version with more spacious compartments. Enhance Customer Support: Sentiment analysis can help your customer service team identify and address recurring issues before they escalate. If a specific issue is repeatedly mentioned in reviews, proactively offering solutions—like better instructions or an extended warranty—can reduce negative feedback and improve customer satisfaction. Optimize Marketing Strategies: Understanding customers' feelings about a competitor’s product can shape your marketing messages. If buyers consistently praise a competitor’s customer service but criticize product durability, you can position your brand as offering superior quality and excellent service to attract dissatisfied customers. ## Legal and Ethical Considerations for Scraping Amazon Reviews While scraping Amazon reviews can provide valuable insights, it’s important to do so in a way that complies with Amazon’s Terms of Service (TOS). Amazon prohibits the use of automated tools to scrape data from its site, so it’s crucial to ensure that your scraping practices are ethical and legal. To build trust with your readers, it’s important to emphasize the importance of ethical scraping. This includes: ### Follow Amazon’s TOS and Guidelines Amazon prohibits unauthorized data scraping using automated bots. To stay compliant, scrape only publicly available review data and avoid bypassing Amazon’s security measures. ### Respect Robots.txt Amazon’s robots.txt file provides guidelines on which parts of the site can be crawled. Following these rules helps ensure that your scraping activities remain ethical and within accepted limits. ### Limit Request Rates Sending too many requests to Amazon’s servers in a short period can trigger rate limits or even get your IP blocked. To prevent this, limit the frequency of scraping requests and use delays between them. ### Use Proxy Servers and Ethical Scraping Methods To avoid overloading Amazon’s servers, consider using rotating proxies and ethical scraping practices that minimize disruption to the website. ### Opt for Amazon’s API When Possible If your business relies on Amazon data, consider using Amazon’s Product Advertising API instead of scraping. This ensures compliance while still providing valuable product data. By following these best practices, you can extract valuable competitor insights while maintaining a trustworthy and compliant approach. ## Integrating Review Insights into Your Business Strategy Scraped reviews are only valuable if they lead to actionable business improvements. Here’s how Amazon sellers can integrate these insights into their strategies: ### Product Development By analyzing reviews, you can identify common complaints or desired features. If a competitor’s customers frequently complain about a fragile product, you can focus on durability in your design. Likewise, if people praise a certain feature, you might consider incorporating it into your next version. ### Marketing Strategies Understanding customer sentiment helps craft more effective marketing messages. If people rave about how lightweight a competitor’s product is, you can emphasize your product’s portability. If buyers complain about confusing instructions, you can promote your “Easy-to-Use” feature. ### Customer Support Instead of waiting for negative reviews, you can address common concerns before they arise. If customers often struggle with assembly, including a step-by-step guide or a how-to video can improve satisfaction and reduce complaints. ### Inventory Planning Tracking review trends can also help you predict future demand. If interest in a particular product feature is rising, you can stock up before competitors catch on. ## Tools and Techniques for Scraping Amazon Reviews Scraping Amazon reviews can be done using various tools, depending on your needs and technical expertise. Here are some of the most popular options: ### Popular Scraping Tools ✔ Scrapy: A powerful Python framework for large-scale web scraping. It’s great for extracting large amounts of data but requires some programming knowledge. ✔ Beautiful Soup: A simple Python library for parsing HTML and XML. It’s easier to use than Scrapy but is better suited for smaller projects. ✔ Playwright: A modern browser automation tool that can handle JavaScript-heavy pages, making it ideal for scraping dynamic content from Amazon. ✔ Selenium: A browser automation framework that can interact with web pages, though it’s slower compared to Playwright for large-scale scraping. Each tool has its own strengths and limitations. While Scrapy and Beautiful Soup are great for static web pages, Playwright and Selenium are better for scraping JavaScript-heavy sites. ### Why Choose Datahut Instead? Manually setting up and maintaining a scraping system is complex, time-consuming, and comes with the risk of violating Amazon’s policies. That’s where Datahut comes in. Unlike Scrapy or Beautiful Soup, Datahut provides a fully managed service—no need to write or maintain scraping scripts. We ensure that data is collected responsibly, respecting Amazon’s policies and guidelines.Whether you need occasional insights or real-time monitoring of thousands of products, Datahut’s services scale with your needs. Instead of just raw data, we provide clean, structured insights that help you improve your product, marketing, and customer service strategies. Focus on growing your Amazon business while we handle the heavy lifting of data extraction and analysis. ## Ready to Get Ahead? Let Datahut Do the Heavy Lifting! Scraping reviews manually takes forever. But [Datahut](https://www.datahut.co/?ref=blog.datahut.co) can automate the process, delivering real-time competitor insights without the hassle. Want to see how it works? Contact us for a free consultation and start outsmarting your competition today! ## FAQ Section 1\. What are the limitations of scraping Amazon reviews? Scraping Amazon reviews can be challenging due to Amazon’s anti-scraping measures, such as CAPTCHAs and IP bans. Additionally, Amazon’s TOS prohibits using automated tools for scraping, so it’s important to proceed cautiously. 2\. How often should reviews be scraped for actionable insights? The frequency of scraping depends on your business needs. For fast-moving industries like fashion or electronics, scraping reviews weekly or monthly may be necessary to stay up-to-date with customer feedback. For slower-moving industries, scraping reviews quarterly may be sufficient. 3\. Can I use scraped reviews for marketing purposes? While you can use insights from scraped reviews to inform your marketing strategies, avoiding directly copying or republishing competitor reviews is important, as this could lead to legal issues. ### Amazon Smartwatch Data Scraping for Time-Series Analysis URL: https://www.blog.datahut.co/post/how-to-scrape-amazon-s-smart-watch-data-for-time-series-analysis/ Last updated: 2026-07-23T07:48:33.000Z Amazon is the world's top-selling e-commerce site offering a wide and dynamic selection of smart watches, which is actually a comprehensive marketplace for wearable technology. There consumers can find a good selection of smart watches in terms of major brands including Apple, Samsung, and Fitbit, among others. This diversity that actually makes the smart watches segment at Amazon a very valuable hub in terms of understanding market dynamics, pricing strategies, and what works for consumers in terms of wearable technology. The web scraping system used here is two phased and sequential with the intention of gathering and processing data efficiently while being respectful to Amazon's platform. The first phase collects links of products from the Amazon smart watch category pages, while the second phase dives deeper into each product page to gather detailed information. The two-step approach ensures comprehensive data collection while maintaining the reliability and efficiency of the system. System to first phase: methodically follows on every Amazon smart watches category page and collects product URLs within a structured process. The first thing the system will be doing is developing a connection with the SQLite database to which the collected URLs will be stored, and then start methodically following on all the pages of the category section. It uses BeautifulSoup to parse the HTML content and fetch links from products and introduces smart rate limiting with some random delay to make a gentle scraping pattern. It keeps all product URLs with the date of collection in a wide database of smart watches available. The actual collector is the second phase of this process. It utilizes Playwright for dynamic content rendering and fetching product details. This stage processes each of the gathered URLs to provide all the details of the products, such as titles, features, customer ratings, price information, descriptions, technical and extra specifications. The system uses advanced error handling mechanisms along with separate tracking of all failed attempts to ensure no product information is missed. All the extracted data are carefully organized and stored within a structured SQLite database; each entry is meticulously date-stamped to enable time-series analysis. The benefits of this web scraping system go far beyond simple data collection. This system, from a market intelligence perspective, allows the tracking of price variations in time, product popularity as reviewed by metrics, and the identification of emerging trends in features and specifications. The technical implementation is robust, with good data quality, error handling, and tracking of failed URLs. Date-stamped entries enable rich temporal analysis. The advanced approach of the system toward rate limiting and user agent rotation ensures that data is collected reliably without overloading the servers of Amazon. ## Step1 :Product Link Scraping ### Setting Up the Scraping Workflow ``` import requests from bs4 import BeautifulSoup import time import random import sqlite3 ``` In this section, we start with the setup of the base process of scraping. We import all the necessary Python libraries: requests for sending HTTP requests to websites, BeautifulSoup from the bs4 package for parsing HTML content, and some additional utilities such as time, random, and sqlite3\. The time and random libraries make use of random pauses between requests to avoid overwhelming the website, and sqlite3 lets us connect to a local database to store the links of the products to be scraped. ### Setting Up the Date for Time-Series Data Scraping ``` # Step 1: Define the date variable at the top DATE_VALUE = 'October 28, 2024' # Change this value as needed ``` At the top of the script, a date variable named DATE\_VALUE is defined. This variable will allow us to manually enter the current date on which data is being scraped. For any kind of time-series analysis, the specific date of when the data is gathered needs to be tracked. With each product link that is stored, we can track the change in data, like when products might not be available or when details of products get updated. ### Connecting to the Database ``` # Step 2: Define database connection def connect_db(): """ Establishes a connection to the SQLite database. This function connects to the 'amazon.db' SQLite database file. If the database does not exist, SQLite will create a new database with the specified name. Returns: sqlite3.Connection: A connection object that represents the database. """ conn = sqlite3.connect('amazon.db') return conn ``` I define the function connect\_db() to connect to a SQLite database named amazon-timeseries-webscraping.db. SQLite is basically a lightweight, file-based database system; if the database file does not exist, SQLite automatically creates it. The function makes sure that we can structure our product data and retrieve it for analysis. This function will return a connection object to us every time we want to interact with the database and allow us to read and write data to the database. ### Creating a Table to Store Product Links ``` # Step 3: Create table if it doesn't exist (without unique constraint for the link) def create_table(conn): """ Creates the 'smart_watches_links' table in the database if it does not already exist. This function defines a table schema for storing smartwatch links. The table includes the following columns: - `id`: Primary key, auto-incremented integer identifier. - `link`: URL of the product, stored as text. This column does not enforce a unique constraint, allowing duplicate entries. - `status`: Integer indicating the status of the link (default is 0). - `date`: Text storing the date the link was added. Args: conn (sqlite3.Connection): A connection object to the SQLite database. Returns: None """ query = '''CREATE TABLE IF NOT EXISTS "smart_watches_links" ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT NOT NULL, status INTEGER DEFAULT 0, date TEXT NOT NULL)''' conn.execute(query) conn.commit() ``` The create\_table() function creates a table in the SQLite database if it doesn't exist; in this case, the table is called smart\_watches\_links, and it is meant for storing links of smartwatches product. Every row in this table is a product link, and there are several fields in it: an id - it is a unique identifier for every entry and is automatically incremented for every new row; a link - it stores the URL of the product and does not require uniqueness, allowing duplicate entries; a status - it is an integer used to track the link's status, set to 0 by default; a date - it records the date when the link was added to the database. ### Setting HTTP Headers for Web Requests ``` # Step 4: Define the base URL and headers def get_headers(): """ Returns the HTTP headers required for making requests to the website. These headers include the 'User-Agent' and 'Accept-Language' fields to mimic a legitimate browser request, helping to avoid detection as a bot during web scraping. Returns: dict: A dictionary containing HTTP headers used in the request, including: - 'User-Agent': Identifies the client browser (in this case, Chrome). - 'Accept-Language': Specifies the preferred language for the response . """ return { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) ' 'AppleWebKit/537.36 (KHTML, like Gecko) ' 'Chrome/85.0.4183.121 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', } ``` The get\_headers() function returns a set of HTTP headers that mimic a real user's browser when making requests to Amazon. Most websites block requests that appear to be coming from bots, so to avoid this we use a User-Agent string that looks like it's coming from a legitimate web browser, such as Google Chrome on Windows. The Accept-Language header is set to English, which tells the server that we want the website content in English. These headers are critical for making successful requests without getting blocked by the website. ### Extracting Product Links from a Page ``` # Step 5: Extract product links from a page def extract_product_links(soup, base_url): """ Extracts product links from a BeautifulSoup object representing a web page. This function searches for anchor tags that match specific classes indicating product links on the page. It constructs the full URLs by appending the relative link found in each anchor tag to the provided base URL. The resulting list of full product URLs is returned. Parameters: soup (BeautifulSoup): The BeautifulSoup object containing the parsed HTML of the page. base_url (str): The base URL of the website to append to the extracted relative links. Returns: list: A list of full product URLs extracted from the page. """ product_links = soup.find_all( 'a', class_="a-link-normal s-underline-text s-underline-link-text " "s-link-style a-text-normal") links = [] for link in product_links: href = link.get('href') if href: full_url = f"{base_url}{href}" links.append(full_url) return links ``` This is named extract\_product\_links(). What needs to be done here is to get a list of product links; the web page that contains all these links is given and represented as a BeautifulSoup object-called soup. It checks every such anchor tag, clearly pointed out specifically by class names in those tags which exists somewhere in the parsed HTML that contains href attribute; so it is going to return a string containing the relative URL of this product. If the href is available, then the function builds the full URL by appending this relative link to the base\_url provided. The list of complete product URLs is then collected and returned. This function is important in streamlining the URLs that would be saved in the database while performing web scraping, hence making it easy to acquire individual product pages for further extraction of data. ### Saving Product Links to the Database ``` # Step 7: Save links to database def save_links_to_db(conn, links, date): """ Saves product links to the database, allowing duplicates. This function takes a list of product links and saves each link to the specified database, with the current date and a default status of 0. It opens a cursor to execute an SQL INSERT statement for each link, storing it in the "smart_watches_links" table with the provided date value. The function allows duplicate entries, meaning multiple records with the same link can be added to the database. Once all links are saved, the database connection commits the changes. Parameters: conn (sqlite3.Connection): The connection object to the SQLite database. links (list): A list of product URLs to save in the database. date (str): The date to associate with each link. """ cursor = conn.cursor() for link in links: cursor.execute( '''INSERT INTO "smart_watches_links" (link, status, date) VALUES (?, 0, ?)''', (link, date) ) # Save the date conn.commit() ``` The function save\_links\_to\_db stores each product link in a list to a SQLite database table called "smart\_watches\_links". It requires three parameters: conn- the database connection object; links- list of product URLs; and date- a string with date for data save. It first creates a cursor for database operations and then follows in an iteration over every link in the list. For each link, it inserts a new row into the table with a link, status 0, and the provided date. Finally, after processing all links, the function commits the transaction so that all the changes are permanent in the database. ### Finding the Next Page of Results ``` # Step 8: Get the next page URL def get_next_page(soup, base_url): """ Retrieves the URL of the next page in a paginated list of search results. This function searches for the link to the next page of results on the current web page using BeautifulSoup. If the "next page" link is found, it appends the relative link to the provided base URL to construct a full URL for the next page. If no "next page" link is found, it returns None. Parameters: soup (BeautifulSoup): The BeautifulSoup object that contains the parsed HTML of the current page. base_url (str): The base URL of the website to prepend to the next page link. Returns: str or None: The full URL of the next page if found; otherwise, None. """ next_page = soup.find('a', class_="s-pagination-next") if next_page and 'href' in next_page.attrs: return f"{base_url}{next_page['href']}" return None ``` The get\_next\_page function allows navigation through pages of a paginated web search result. It uses BeautifulSoup to parse the HTML content of the current page and searches for the link with the class that indicates the "next page" button, which is commonly labeled as "s-pagination-next". When found, it extracts the relative URL from the href attribute. Then, by appending that relative URL to the base\_url, a full URL is created so that using this URL one can actually navigate the next page. If this "next page" link is not found then the function returns as None showing there are no more pages left for scraping. This comes in useful while scraping web data while having to structure the loops for multiple pages. ### Scraping All Pages ``` # Step 9: Scrape all pages and save links to the database def scrape_all_pages(conn, start_url, base_url): """ Scrapes multiple pages and saves product links to a database. This function starts at a given URL and iterates through paginated search result pages, scraping product links from each page. It retrieves and parses each page’s content, extracts product links, and saves them to the database with the provided date. The function continues to the next page if available, until no more pages are found. Parameters: conn (sqlite3.Connection): Database connection object to store product links. start_url (str): The URL of the first page to scrape. base_url (str): The base URL of the website used to construct full product URLs. Returns: None """ current_url = start_url headers = get_headers() while current_url: print(f"Scraping page: {current_url}") response = requests.get(current_url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') # Extract links and save to DB links = extract_product_links(soup, base_url) save_links_to_db(conn, links, DATE_VALUE) # Pass the date value # Find the next page URL next_page_url = get_next_page(soup, base_url) if next_page_url: current_url = next_page_url time.sleep(random.uniform(2, 5)) # Avoid blocking else: print("No more pages to scrape.") break ``` The scrape\_all\_pages function is designed to walk through and scrape multiple pages of a search results listing. At the start\_url, it fetches the HTML content by sending a request and then parse the content using BeautifulSoup. After parsing, it applies the extract\_product\_links function to gather all product links on that page, which then saves those links into a database with a given date using the save\_links\_to\_db function. To continue scraping, the function checks for the existence of a "next page" link on the current page by calling the get\_next\_page function. If there is a next page, it updates the current URL and waits a random amount of time between 2 to 5 seconds to avoid being blocked by the server. The function keeps on repeating the loop until it finds the last page where no "next page" link exists and then stops, thus ending the process of scraping. This will make for an organized way to get links from various pages. ### Main Function: Tying Everything Together ``` # Step 10: Main function to run the scraper def main(): """ Main function to initiate and run the web scraper. This function establishes a connection to the database, creates a table for storing product links, and then initiates the process to scrape multiple pages for product URLs. Once scraping is complete, it closes the database connection to ensure all data is properly saved. Parameters: None Returns: None """ base_url = "https://www.amazon.in" start_url = "https://www.amazon.in/s?i=fashion&rh=n%3A27413352031&s=popularity-rank&fs=true&ref=lp_27413352031_sar" # Connect to the database and create table conn = connect_db() create_table(conn) # Scrape all pages and save links scrape_all_pages(conn, start_url, base_url) conn.close() ``` The main function is the entry point of the scraper program. There, it initializes two very important URLs: base\_url - the main domain of the website, and start\_url-the first page with products that are to be scraped. Then, it connects to the database with the connect\_db function and creates the table structure required by calling create\_table. After initializing the database the function invokes scrape\_all\_pages, handling navigation through all the pages and saving the link to the product. On finish with scraping, the function closes the database connection as a guarantee that all saved data has been successfully retrieved. ### Running the Scraper ``` # Run the main function if __name__ == "__main__": """ Run the main function if this script is executed as the primary module. This condition checks whether the script is being run directly or imported as a module in another script. If it is the main script being run, it calls the main function to start the web scraping process. """ main() ``` This function allows to give the flow structure in its run to perform an entirely scraping process. In Python, each script has a special built-in variable called name. When the script is run directly, then name is set to " main". This check is what prevents the main function from running if this script is imported as a module in another script. For instance, it calls the main() inside it, which begins the whole process of scraping-starting from connecting to the database up to saving product links. This approach is helpful for modularity as functions can be reused elsewhere if needed without running the scraping logic automatically. ## STEP 2:Product Data Scraping From Product Links ### Library Imports for Web Scraping and Database Management ``` import sqlite3 import random import time from playwright.sync_api import sync_playwright from bs4 import BeautifulSoup ``` This module imports necessary libraries for web scraping and SQLite database for data managing. Every library added performs some specific function within the functionality of the program: sqlite3: This is a Python's built-in library that offers the SQLite database interface, meaning you can create, read, update, and delete records in a lightweight database. SQLite is very good for web scraping because data is efficiently stored without full database server overhead. random: This library is also used for generating random numbers and doing random operations. In web scraping, it is very commonly used to introduce time delays between requests, like time.sleep() to emulate the effect of human browsing, not being blocked by the server. time: There are several time-related functions that the time module has. When it comes to web scraping, it is typically utilized to stop the execution of the program for a specific amount of time. That is important not to load the server with requests and to control the timing of requests properly. playwright.sync\_api: This is a library in Playwright that is an automation library for browser-based applications. Sync\_playwright allows the controlling of browsers like Chrome, Firefox, and Safari in a synchronous manner. This makes the automation of web scraping much easier, especially when you need to navigate through certain web pages, interact with some elements, or perhaps need to extract data which JavaScript renders dynamically. bs4 (BeautifulSoup): Beautiful Soup is a library used to parse HTML and XML documents. It generates a parse tree for parsing the given HTML document and allows search queries to find and navigate through elements easily. This is highly beneficial in extracting specific data points from web pages once they are loaded into the browser, especially if one uses Playwright in scraping content rendered by JavaScript. ### Function to Load User Agent Strings for Web Scraping ``` def load_user_agents(filepath="data/user_agents.txt"): """ Load user agent strings from a specified file. This function reads user agent strings from a text file, where each user agent is expected to be on a new line. It strips any leading or trailing whitespace from each line and filters out empty lines. If no user agents are found, it raises a ValueError to indicate that the list of user agents is empty. Parameters: filepath (str): The path to the text file containing user agents. Defaults to "data/user_agents.txt". Returns: list: A list of user agent strings. """ with open(filepath, 'r') as file: agents = [ line.strip() for line in file if line.strip() ] if not agents: raise ValueError("No user agents found in file") return agents ``` The load\_user\_agents function should read a file containing user agents, used to identify web browsers and devices when making requests towards a website. Each line of the file should contain exactly one user agent. At first, the function opens a file, reads all its lines removing unnecessary spaces. It then removes all empty lines so that only the valid user agents will remain in the final list. If no user agent exists in the file, it will throw an error so that the user will be informed that the file is empty. Finally, it returns a list of user agents that can be used to make web scraping requests appear as if they come from different browsers or devices, which can help prevent the scraping from being blocked. Function to Generate Random Delay for Web Scraping ``` def get_random_delay(): """ Generate a Random Delay for Web Scraping This function returns a random floating-point number that represents a delay in seconds. The delay is generated within a specified range of 3.0 to 6.0 seconds, which can be used to space out web scraping requests. Returns: float: A random delay value between 3.0 and 6.0 seconds. """ return random.uniform(3.0, 6.0) ``` The get\_random\_delay function is designed to help avoid detection of web scraping activity as automated behavior. It returns a random floating-point number that represents a delay in seconds from 3.0 to 6.0\. This delay can be used between consecutive web scraping requests to make the requests appear more human-like. It reduces the risk of being blocked by the website being scraped through introduction of variability in the timing of requests, which makes for a more sustainable and respectful approach to scraping. ### Function to Fetch Web Page Content Using Random User Agent ``` def fetch_page_content(url, user_agents): """ Fetch Web Page Content Using a Random User Agent This function retrieves the HTML content of a specified web page by navigating to the provided URL. It uses Playwright to open a browser instance with a randomly selected user agent from the provided list. The function introduces a random delay before making the request to avoid detection by the website. Parameters: url (str): The URL of the web page to be scraped. user_agents (list): A list of user agent strings to choose from. Returns: str: The HTML content of the web page. """ user_agent = random.choice(user_agents) with sync_playwright() as p: browser = p.chromium.launch(headless=True) context = browser.new_context( user_agent=user_agent ) page = context.new_page() try: delay = get_random_delay() time.sleep(delay) page.goto(url) content = page.content() return content finally: browser.close() ``` fetch\_page\_content is the function aimed at fetching the HTML content on a web page by navigating to a specified URL. It loads a user agent from a list provided to the function to behave like different browsers, which may help to evade possible blocking by the website. Before making the request, the function introduces some random delays for requests made, so that requests seem more human-like than algorithmic. By using the Playwright, the function launches the headless browser, navigates to the URL, and captures the resulting HTML content, returning it to you as a string. This function is very useful in web scraping: while you will scrape information from website pages, you don't run much risk of getting flagged as a bot. ### Function to Extract Product Title from HTML Soup ``` def extract_title(soup): """ Extract Product Title from HTML Soup This function retrieves the product title from an HTML document represented by the BeautifulSoup object. It looks for a specific HTML element identified by the span tag with the ID 'productTitle'. If the title is found, it returns the cleaned text; otherwise, it returns None. Parameters: soup (BeautifulSoup): The BeautifulSoup object representing the HTML document. Returns: str or None: The product title as a string, or None if the title is not found. """ title_tag = soup.find('span', id='productTitle') if title_tag: title = title_tag.get_text(strip=True) return title return None ``` The extract\_title function will extract the product title from an HTML document using BeautifulSoup. It looks for a certain element, which is usually a tag with the ID productTitle, containing the name of the product. If such an element is found, the function retrieves and cleans the text removing any extra whitespace. The clean title is then returned as a string. In case the title is not available, it returns None. This is a critical function for any web scraping activity where retrieving a title for a product is required in analysis or presentation. ### Function to Extract Product Ratings from HTML Soup ``` def extract_ratings(soup): """ Extract Product Ratings from HTML Soup This function retrieves product ratings and review counts from an HTML document represented by the BeautifulSoup object. It looks for a specific div element with the ID 'averageCustomerReviews'. If the ratings and review counts are found, they are returned as a dictionary; otherwise, the function returns a dictionary with None values. Parameters: soup (BeautifulSoup): The BeautifulSoup object representing the HTML document. Returns: dict: A dictionary containing the product rating and review count, with keys 'Rating' and 'Review Count'. """ ratings = {} reviews_div = soup.find( 'div', id='averageCustomerReviews' ) if reviews_div: rating_tag = reviews_div.find( 'span', class_='a-size-base a-color-base' ) if rating_tag: rating = rating_tag.get_text(strip=True) else: rating = None review_count_tag = reviews_div.find( 'span', id='acrCustomerReviewText' ) if review_count_tag: review_count = review_count_tag.get_text(strip=True) else: review_count = None ratings['Rating'] = rating ratings['Review Count'] = review_count return ratings ``` The extract\_ratings function is used to pull ratings for a product along with the number of reviews from an HTML document. It looks for a
      element with an ID of averageCustomerReviews, which contains the relevant data. It will look inside this for tags carrying rating and review counts. Function takes the inner text out from the tag, processes the extracted string by taking away unnecessary whitespace, then returns those as values of a dictionary. Both ratings and the number of reviews are being checked first and if cannot find a respective value puts None. The final output is a dictionary containing the product's rating and review count, making it useful for data analysis in web scraping applications. ### Function to Extract Sale and Retail Prices from HTML Soup ``` def extract_prices(soup): """ Extract Prices from HTML Soup This function retrieves both sale and retail prices from an HTML document represented by the BeautifulSoup object. It looks for a specific div element with the ID 'corePriceDisplay_desktop_feature_div'. If the prices are found, they are returned as a dictionary; otherwise, the function returns a dictionary with None values. Parameters: soup (BeautifulSoup): The BeautifulSoup object representing the HTML document. Returns: dict: A dictionary containing the sale price and retail price, with keys 'Sale Price' and 'Retail Price'. """ prices = {} price_div = soup.find( 'div', id='corePriceDisplay_desktop_feature_div' ) if price_div: sale_price_tag = price_div.find( 'span', class_='a-price-whole' ) if sale_price_tag: sale_price = sale_price_tag.get_text(strip=True) else: sale_price = None retail_price_tag = price_div.find( 'span', class_='a-text-price' ) if retail_price_tag: retail_price = retail_price_tag.get_text(strip=True) else: retail_price = None if retail_price: if '₹' in retail_price: retail_price = retail_price.split('₹')[-1].strip() prices['Sale Price'] = sale_price if retail_price: prices['Retail Price'] = f'{retail_price}' else: prices['Retail Price'] = None return prices ``` This is an extract\_prices function that pulls out the sale and retail prices from an HTML document that has been represented with a BeautifulSoup object. It's finding a
      that contains its id corePriceDisplay\_desktop\_feature\_div. If the former exists, then it'll check within the div for certain span tags holding the sale and retail prices. It fetches and cleans the text from these tags, removing any unwanted whitespace. The function also checks for the currency symbol and processes the retail price accordingly. Finally, it returns a dictionary containing both the sale price and retail price, with values set to None if they cannot be found. This function is helpful for web scraping applications focused on collecting product pricing data. ### Function to Extract Product Description from HTML Soup ``` def extract_description(soup): """ Extract Product Description from HTML Soup This function retrieves the product description from an HTML document represented by the BeautifulSoup object. It first looks for an unordered list (
        ) with the class 'a-unordered-list a-vertical a-spacing-small'. If found, it collects all list items (
      • ) within that list. If no such list is found, it checks for another unordered list with a slightly different class 'a-unordered-list a-vertical a-spacing-mini'. The collected texts are returned as a list. Parameters: soup (BeautifulSoup): The BeautifulSoup object representing the HTML document. Returns: list: A list of strings containing the product description items extracted from the HTML. """ description = [] # Condition 1: Check for
          ul_lists = soup.find_all( 'ul', class_='a-unordered-list a-vertical a-spacing-small' ) if ul_lists: # If any
            elements found for Condition 1 for ul in ul_lists: list_items = ul.find_all('li') for item in list_items: text = item.get_text(strip=True) description.append(text) else: # Condition 2: If no
              elements found in Condition 1, # check for
                ul_mini = soup.find_all( 'ul', class_='a-unordered-list a-vertical a-spacing-mini' ) for ul in ul_mini: list_items = ul.find_all('li') for item in list_items: text = item.get_text(strip=True) description.append(text) return description ``` This extract\_description function is designed to extract the product description from an HTML document represented by a BeautifulSoup object. The search looks for unordered lists with the class a-unordered-list a-vertical a-spacing-small. If such lists exist, it retrieves the text of each list item and appends it to a description list. If no lists of the first class are found, it looks for another unordered list with the class a-unordered-list a-vertical a-spacing-mini. The function is very useful for web scraping applications, which require understanding the product's features or details from an e-commerce page. The result is a list of descriptive strings that give insight to the characteristics of the product. ### Function to Extract Technical Details from HTML Soup ``` def extract_technical_details(soup): """ Extract Technical Details from HTML Soup This function retrieves technical specifications of a product from an HTML document represented by a BeautifulSoup object. It first checks for a table with the ID 'productDetails_techSpec_section_1'. If found, it collects key-value pairs from the rows of that table. If this table is not present, it attempts to find an alternative table with the ID 'technicalSpecifications_section_1' and extracts the details from there. The collected technical details are returned as a dictionary. Parameters: soup (BeautifulSoup): The BeautifulSoup object representing the HTML document. Returns: dict: A dictionary containing technical details with keys as specifications and values as their corresponding information. """ technical_details = {} # First check the original table structure tech_table = soup.find( 'table', id='productDetails_techSpec_section_1' ) if tech_table: tech_rows = tech_table.find_all('tr') for row in tech_rows: key_tag = row.find( 'th', class_='a-color-secondary a-size-base prodDetSectionEntry' ) value_tag = row.find( 'td', class_='a-size-base prodDetAttrValue' ) if key_tag and value_tag: key = key_tag.get_text(strip=True) value = value_tag.get_text(strip=True) technical_details[key] = value else: # Else condition to handle the alternative HTML structure tech_table_alt = soup.find( 'table', id='technicalSpecifications_section_1' ) if tech_table_alt: tech_rows_alt = tech_table_alt.find_all('tr') for row in tech_rows_alt: key_tag_alt = row.find( 'th', class_='a-span5 a-size-base' ) value_tag_alt = row.find( 'td', class_='a-span7 a-size-base' ) if key_tag_alt and value_tag_alt: key_alt = key_tag_alt.get_text(strip=True) value_alt = value_tag_alt.get_text(strip=True) technical_details[key_alt] = value_alt return technical_details ``` This extract\_technical\_details function is supposed to scrap and return the technical details of a product from an HTML representation in the form of a BeautifulSoup object. The function will look for a particular table within the HTML that has a productDetails\_techSpec\_section\_1 ID when called since most online retailers such as Amazon use it in storing their product details. If this table is available, it scans through every row of the table searching for a specification name in the header cell (

    ) and its corresponding detail in the data cell (). Then, it pulls out and cleans text from these elements and puts them in a dictionary where keys are the names of specifications, and values are their respective details. If the table with the original ID is not found, then the function identifies an alternative table with the ID technicalSpecifications\_section\_1 and continues to repeat the extraction if this table exists. The function returns a dictionary with all details that are relevant at the end so access and sorting of product specifications becomes easy. This is very useful in web scraping applications, hence allowing a user to extract, compare, and analyze product information easily. ### Extracting Product Information with HTML Parsing ``` def extract_additional_details(soup): """ Extracts product details from a given HTML structure. This function attempts to gather product information from two possible structures commonly found on product pages. It first searches for a bullet-point list format located within a div with the ID 'detailBulletsWrapper_feature_div', extracting key-value pairs from list items ('li') that contain bold labels (in a 'span' with the class 'a-text-bold') and corresponding values. If this structure is not found, the function will look for a table structure within a table with the ID 'productDetails_detailBullets_sections1'. Here, each row ('tr') of the table provides a label ('th') and a value ('td') that are extracted and added to a dictionary. Parameters: soup (BeautifulSoup): Parsed HTML content of the webpage. Returns: dict: A dictionary containing product details, where keys are attribute labels and values are attribute details. """ details = {} # Attempt to extract details from the first structure # (ul list with detail bullets) detail_wrapper = soup.find( 'div', id='detailBulletsWrapper_feature_div' ) if detail_wrapper: print( "Extracting details from detailBulletsWrapper_feature_div..." ) ul_list = detail_wrapper.find_all('li') for li in ul_list: label = li.find( 'span', class_='a-text-bold' ) value = li.find_all('span')[-1] # Value is usually the last span if label and value: label_text = label.get_text(strip=True) value_text = value.get_text(strip=True) details[ label_text.replace(':', '') ] = value_text # Attempt to extract details from the second structure # (table format) detail_table = soup.find( 'table', id='productDetails_detailBullets_sections1' ) if detail_table: print( "Extracting details from productDetails_detailBullets_sections1..." ) rows = detail_table.find_all('tr') for row in rows: th = row.find('th') td = row.find('td') if th and td: th_text = th.get_text(strip=True) td_text = td.get_text(strip=True) details[th_text] = td_text return details ``` This function, extract\_additional\_details, is actually used to pull out specifications for products and any other information off HTML-based content. It tries to find the two common formats on pages describing products. In the first format, a bulleted list usually makes up product details, where label is paired with its value. Function detects this list by finding an ID detailBulletsWrapper\_feature\_div, then the dictionary is filled with text, which was taken from the bolded labels and the values of those labels. If the first structure is not available, then the function will look for a table structure with the ID productDetails\_detailBullets\_sections1\. Here, every row would have a label (th tag) and a value (td tag). All of them are extracted and added to the dictionary. Finally, the output will be a dictionary with attribute labels as keys and their corresponding details as values, providing a structured way of accessing product data for further use. ### Extracting Product Features with HTML Parsing ``` def extract_features(soup): """ Extracts product features from a webpage. This function attempts to gather product features from specific HTML structures found on product pages. First, it checks for an expandable "About this item" section, locating product details within a div container with the class 'a-fixed-left-grid product-facts-detail'. It extracts the label and value pairs found in left and right columns. If this section is unavailable, the function looks for a table format with class 'a-normal a-spacing-micro'. It iterates through the table rows, locating labels in 'td' elements with class 'a-span3' and values in 'td' elements with class 'a-span9'. The extracted features are returned as key-value pairs in a dictionary. Parameters: soup (BeautifulSoup): Parsed HTML content of the webpage. Returns: dict: A dictionary containing feature labels as keys and their corresponding values as entries. """ features = {} # Condition 1: Click to expand and scrape the "About this item" section expand_button = soup.find( 'span', class_='a-expander-prompt' ) if expand_button: # Simulate the click (if using Selenium or a similar tool, you'd actually click it) div_container = soup.find_all( 'div', class_='a-fixed-left-grid product-facts-detail' ) for div in div_container: key_tag = div.find( 'div', class_='a-col-left' ) value_tag = div.find( 'div', class_='a-col-right' ) if key_tag and value_tag: key = key_tag.find( 'span', class_='a-color-base' ).get_text(strip=True) value = value_tag.find( 'span', class_='a-color-base' ).get_text(strip=True) if key and value: features[key] = value # Condition 2: Scraping the features from the table structure table_container = soup.find( 'table', class_='a-normal a-spacing-micro' ) if table_container: tbody = table_container.find('tbody') feature_rows = ( tbody.find_all('tr') if tbody else [] ) for row in feature_rows: key_tag = row.find( 'td', class_='a-span3' ) value_tag = row.find( 'td', class_='a-span9' ) if key_tag and value_tag: key_span = key_tag.find( 'span', class_='a-size-base a-text-bold' ) value_span = value_tag.find( 'span', class_='a-size-base' ) key = ( key_span.get_text(strip=True) if key_span else None ) value = ( value_span.get_text(strip=True) if value_span else None ) if key and value: features[key] = value return features ``` This function extract\_features collects product feature information from structured HTML elements found on a product page. It starts trying to find details in an "About this item" expandable section. When this section is there, it will simulate a click action, handy when working with libraries such as Selenium and extract details from a grid layout in which the left column contains the feature name and the right column contains its value. If "About this item" section isn't available, then it functions like a table format search gathering the feature names and their values from the rows. It stores every feature in a dictionary with the help of a key-value pair for access of some organized feature of the product. ### Scraping Amazon Product Details with Ease ``` def scrape_amazon_product(url, user_agents): """ Scrapes product information from an Amazon product page. This function retrieves various details about a product by parsing the HTML content of its Amazon webpage. It begins by fetching the page content using the provided URL and rotating user agents to avoid detection. Key details are then extracted, including the title, features, ratings, price, description, technical specifications, and additional details. These are returned in a structured format for easier access. Parameters: url (str): The URL of the Amazon product page to scrape. user_agents (list): A list of user agent strings for rotating during requests. Returns: tuple: A tuple containing product information, with each element corresponding to a specific detail. """ content = fetch_page_content(url, user_agents) soup = BeautifulSoup(content, "html.parser") title = extract_title(soup) features = extract_features(soup) ratings = extract_ratings(soup) prices = extract_prices(soup) description = extract_description(soup) technical_details = extract_technical_details(soup) additional_details=extract_additional_details(soup) return ( title, features, ratings, prices, description, technical_details, additional_details ) ``` The scrape\_amazon\_product function is designed to gather comprehensive details from an Amazon product page. It starts by requesting the webpage's HTML content, with rotating user agents to prevent detection. The function then parses the HTML to collect specific details, such as the product title, features, ratings, price, description, and technical and additional details. Each piece of information is extracted by dedicated functions to ensure that each aspect is accurately retrieved. All collected data is returned in a tuple, providing an organized structure that allows easy access to each piece of product information for further analysis or storage. ### Setting Up a Database for Product Data Storage ``` def initialize_database(db_path): """ Initializes a SQLite database for storing product data and failed URLs. This function creates a connection to the specified SQLite database file and initializes two tables if they do not already exist. The first table, 'smart_watch_product_data', is designed to store product information such as URL, date, title, features, ratings, prices, description, technical details, and additional details, with a unique primary key based on the URL and date to prevent duplicate entries. The second table, 'failed_urls', stores URLs that fail to scrape successfully, along with the error message, timestamp, and status. After creating the tables, the connection to the database is closed to ensure data integrity. Parameters: db_path (str): The file path of the SQLite database file where tables will be created. Returns: None """ conn = sqlite3.connect(db_path) cursor = conn.cursor() # Create smart_watch_product_data table if it doesn't exist cursor.execute(""" CREATE TABLE IF NOT EXISTS smart_watch_product_data ( url TEXT, date TEXT, title TEXT, features TEXT, ratings TEXT, prices TEXT, description TEXT, technical_details TEXT, additional_details TEXT, PRIMARY KEY (url, date) ) """) # Create failed_urls table with date and status columns cursor.execute(""" CREATE TABLE IF NOT EXISTS failed_urls ( url TEXT, date TEXT, error_message TEXT, timestamp DATETIME DEFAULT CURRENT_TIMESTAMP, status INTEGER DEFAULT 0, PRIMARY KEY (url, date) ) """) conn.commit() conn.close() ``` The function initialize\_database would be used to get a structured database to put both successfully scraped product information and any failed URLs during scraping. It initializes a connection with the SQLite file specified as the database to the function then creates two tables if the tables do not exist. The smart\_watch\_product\_data table is created for storing information of the all-inclusive products and every record can be uniquely identified using a combination of the url and date fields as the primary key so that duplicate entries cannot be created. Another table, Failed\_urls is designed to keep track of those URLs which are throwing error messages while scraping, with error messages, timestamps and status field to track the failure. It ends up closing the connection once the tables are initialized to make sure data integrity and the database is ready for further operations. ### Retrieving Unscraped URLs from the Database ``` def fetch_unscraped_urls_from_main_table(db_path): """ Fetches URLs marked as unscraped from the main database table. This function connects to a specified SQLite database to retrieve URLs that have not yet been scraped. It queries the 'smart_watches_links' table for entries where the 'status' field is set to 0, indicating these URLs are pending for scraping. Each result includes the URL link and the associated date, which are then returned as a list of tuples for easy access. After retrieving the data, the database connection is closed to maintain data integrity. Parameters: db_path (str): The file path of the SQLite database to query. Returns: list: A list of tuples, each containing a URL and its associated date, for URLs marked as unscraped. """ conn = sqlite3.connect(db_path) cursor = conn.cursor() cursor.execute(""" SELECT link, date FROM smart_watches_links WHERE status = 0 """) results = cursor.fetchall() conn.close() return [(url, date) for url, date in results] ``` The fetch\_unscraped\_urls\_from\_main\_table function scans and retrieves URLs from the database that have not been scrapped yet. It connects to the SQLite database file passed as an argument and runs a query on the smart\_watches\_links table by selecting entries with a status of 0, which marks them as unscraped. The function retrieves both the URL and the date in each entry and returns it in a list of tuples. After gathering this information, the function securely closes the database connection to ensure data integrity. ### Retrieving Pending URLs from the Failed Table ``` def fetch_unscraped_urls_from_failed_table(db_path): """ Fetches URLs marked as unscraped from the failed URLs table. This function connects to the specified SQLite database to retrieve URLs that failed to be scraped previously. It queries the 'failed_urls' table, specifically selecting entries with a 'status' field set to 0, indicating they are still pending for re-scraping. Each result includes the URL and the date it was added to the failed list, and these are returned as a list of tuples. The database connection is closed after retrieval to ensure data security and maintain database integrity. Parameters: db_path (str): The file path of the SQLite database to query. Returns: list: A list of tuples, each containing a URL and its associated date, for URLs marked as unscraped in the failed URLs table. """ conn = sqlite3.connect(db_path) cursor = conn.cursor() cursor.execute(""" SELECT url, date FROM failed_urls WHERE status = 0 """) results = cursor.fetchall() conn.close() return [(url, date) for url, date in results] ``` The fetch\_unscraped\_urls\_from\_failed\_table function is used to collect URLs from the database that, on previous attempts, failed to get scraped. It connects to the SQLite database specified and then queries the failed\_urls table, looking for entries where status is set to 0-that is, pending for re-scraping. The function fetches the URL and date of each entry and returns them as a list of tuples. Finally, the database connection is closed to maintain security and integrity of data as well as for keeping the database ready for further operations after fetching the data. ### Updating URL Status in the Database ``` def update_url_status(db_path, url, date, table_name, status=1): """ Updates the scraping status of a specified URL in the database. This function connects to the specified SQLite database to update the status of a given URL based on its table location. Depending on the table provided ('smart_watches_links' or 'failed_urls'), it updates the 'status' field for the URL and date combination, marking it as scraped or pending for further action. By default, the status is set to 1 (indicating successful scraping), but this can be modified as needed. After executing the update, the database connection is closed to ensure data integrity. Parameters: db_path (str): The file path of the SQLite database to update. url (str): The URL whose status needs to be updated. date (str): The date associated with the URL entry. table_name (str): The table to update ('smart_watches_links' or 'failed_urls'). status (int, optional): The new status value (default is 1, indicating success). Returns: None """ conn = sqlite3.connect(db_path) cursor = conn.cursor() if table_name == "smart_watches_links": query = """ UPDATE smart_watches_links SET status = ? WHERE link = ? AND date = ? """ else: # failed_urls table query = """ UPDATE failed_urls SET status = ? WHERE url = ? AND date = ? """ cursor.execute(query, (status, url, date)) conn.commit() conn.close() ``` The update\_url\_status function updates the scraping status of particular URLs in the database. It connects to the SQLite database specified and updates the status field for the given URL and date in either the smart\_watches\_links or failed\_urls table. By default, it sets the status to 1, marking the URL as successfully scraped, but this can be customized. Upon closure, the function ensures that data received is securely stored after updating it by closing the database connection, thus allowing any further operations on the database without having inconsistencies. ### Logging Failed URLs in the Database ``` def save_failed_url(db_path, url, date, error_message): """ Saves a failed URL along with error details in the database. This function connects to the specified SQLite database to log URLs that failed during the scraping process. It inserts a new record into the `failed_urls` table, capturing the URL, date, and a descriptive error message associated with the failure. The status for this entry is set to `0`, indicating it is pending for re-scraping. After the insertion, the function commits the changes and closes the database connection to maintain data integrity. Parameters: db_path (str): The file path of the SQLite database to update. url (str): The URL that failed to be scraped. date (str): The date associated with the failed scraping attempt. error_message (str): A message detailing the reason for the failure. Returns: None """ conn = sqlite3.connect(db_path) cursor = conn.cursor() cursor.execute(""" INSERT INTO failed_urls (url, date, error_message, status) VALUES (?, ?, ?, 0) """, (url, date, str(error_message))) conn.commit() conn.close() ``` The save\_failed\_url function is created to capture the error URLs in a database when they fail at the time of scraping. It first establishes a connection with the SQLite database specified and then inserts a new row into the failed\_urls table with the URL itself, date of failure, and an error message in detail. This entry is indicated by the status 0, meaning it's still pending at the next attempted time to be re-scraped. The function commits changes and closes the database so that the data will safely lock up, and then the database will be ready for future operations after successfully saving information. ### Storing Product Data in the Database ``` def save_data_to_db(db_path, product_data): """ Saves product data into the database. This function connects to the specified SQLite database and inserts a new record into the `smart_watch_product_data` table. It takes a tuple containing various details about a product, such as URL, date, title, features, ratings, prices, description, technical details, and additional details. After executing the insertion, the function commits the changes to the database and closes the connection to ensure that the data is securely stored and the database is ready for future operations. Parameters: db_path (str): The file path of the SQLite database to update. product_data (tuple): A tuple containing product details in the following order: (url, date, title, features, ratings, prices, description, technical_details, additional_details). Returns: None """ conn = sqlite3.connect(db_path) cursor = conn.cursor() cursor.execute(""" INSERT INTO smart_watch_product_data (url, date, title, features, ratings, prices, description, technical_details,additional_details) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) """, product_data) conn.commit() conn.close() ``` The save\_data\_to\_db function will save the product data to the database. After establishing a connection to the SQLite database specified, it saves a new record in the smart\_watch\_product\_data table. The function accepts a tuple of necessary information about the product, including URL, date of scraping, title, features, ratings, prices, description, technical details, and others. It commits the changes after the execution of the insert operation so that data might be safely saved and would close the database connection. In this way, this step ensures that data would remain sound during any such operation of data insertion or retrieval in the database. ### Comprehensive URL Scraping Functionality ``` def scrape_urls( db_path, urls_to_scrape, table_name, user_agents ): """ Scrapes product information from a list of URLs. This function iterates through a list of URLs and attempts to scrape product data for each one. For each URL, it introduces a random delay before initiating the scrape to avoid overwhelming the server. The scraped data, which includes the title, features, ratings, prices, description, technical details, and additional details, is then saved to the database. If the scrape is successful, the URL status is updated in the specified table. In case of an error during scraping, the URL and error message are recorded, and the status is updated accordingly. Parameters: db_path (str): The file path of the SQLite database. urls_to_scrape (list of tuples): A list containing tuples of URLs and corresponding dates to scrape. table_name (str): The name of the table to update with the URL status (either 'smart_watches_links' or 'failed_urls'). user_agents (list): A list of user-agent strings to use for scraping. Returns: None """ for url, date in urls_to_scrape: try: delay = get_random_delay() print( f"Waiting {delay:.2f} seconds before " f"scraping next URL..." ) time.sleep(delay) print( f"Scraping: {url} for date: {date}" ) title, features, ratings, prices, \ description, technical_details, \ additional_details = scrape_amazon_product( url, user_agents ) # Convert data to strings for storage product_data_str = ( url, date, title, str(features), str(ratings), str(prices), str(description), str(technical_details), str(additional_details) ) # Save the scraped data save_data_to_db( db_path, product_data_str ) # Update URL status to scraped update_url_status( db_path, url, date, table_name ) print( f"Successfully scraped: {url} " f"for date: {date}" ) except Exception as e: print( f"Failed to scrape {url} for date {date}: {e}" ) if table_name == "smart_watches_links": # Only save to failed_urls if from the main table save_failed_url( db_path, url, date, str(e) ) else: # If already from failed_urls, update # status to mark as permanent failure update_url_status( db_path, url, date, "failed_urls", status=1 ) ``` Function scrape\_urls is developed to scrape product information on given URLs. Its input parameters are a database path, the list of the URLs together with respective dates, the name of the table for update and the collection of user agents. For each URL, this function waits for some random delay to avoid server blocking and then it proceeds to scrape the product details using the function scrape\_amazon\_product. Then format all these data into strings, namely title, features, ratings, prices, description, technical details etc and store in the database. Later, it marked the status of the URL as "completed" after successful scraping. If an error occurs at any stage of this process, the function will catch its message and record the URL in a failed\_urls table for further review. This structured approach enables the handling of both successful attempts to scrape and failed ones with optimal reliability to the web scraping workflow process. ### Comprehensive Web Scraping Orchestration ``` def main(): """ Main function to execute the web scraping workflow. This function orchestrates the entire web scraping process by initializing the database, fetching URLs from both the main and failed URLs tables, and performing the scraping operations. It first sets up the database by calling the `initialize_database` function, handling any potential operational errors related to database schema. Then, it retrieves URLs that have not yet been scraped from the main table and attempts to scrape product data from those URLs. If there are any URLs recorded in the `failed_urls` table, the function will also attempt to scrape those URLs as well. Finally, it notifies the user when the scraping process is complete. Returns: None """ db_path = ( "amazon-timeseries-webscraping.db" ) user_agents = load_user_agents() try: # Initialize database tables initialize_database(db_path) except sqlite3.OperationalError as e: if "duplicate column" not in str(e).lower(): raise e # Step 1: Scrape from main table print( "Starting to scrape URLs from main table..." ) urls_from_main = fetch_unscraped_urls_from_main_table( db_path ) print( f"Found {len(urls_from_main)} unscraped URLs " f"in main table" ) scrape_urls( db_path, urls_from_main, "smart_watches_links", user_agents ) # Step 2: Scrape from failed_urls table print( "\nStarting to scrape URLs from failed_urls table..." ) failed_urls = fetch_unscraped_urls_from_failed_table( db_path ) print( f"Found {len(failed_urls)} unscraped URLs in " f"failed_urls table" ) scrape_urls( db_path, failed_urls, "failed_urls", user_agents ) print("\nScraping process completed!") ``` This basic center function basically just performs the web scraping workflow, which at first starts off by taking a path to a database and loading up a list of user agents that it will make use of for actually conducting the scrape. After it's finished initializing the database via the calling initialize\_database it goes ahead and pulls down all the URLs for scrapes from the main table, along with pulling from failed\_urls. This then calls the scrape\_urls function on each batch of URLs with the attempt to fetch their product data. The work flow is designed to catch database errors so that whatever operational issues occur, this can be logged without stopping the rest of the processes. Finally, the function returns notifying the user that all its operations concerning scrapings are done and concludes an organized structure for effective scraping task management. This makes it appropriate for users who may not have an extremely extensive background in programming since the flow is logical and easy to follow. ### Main Program Execution Control ``` if __name__ == "__main__": """ Entry point for the web scraping application. This block of code checks if the current script is being run as the main program. If so, it calls the `main` function to initiate the web scraping process. This allows the script to be executed directly, while also ensuring that the web scraping operations are only performed when intended, such as when running the script from the command line. It serves as a standard practice in Python programming to organize the execution flow of the program. Returns: None """ main() ``` The if \_\_name\_\_ == "\_\_main\_\_": block, actually, constitutes the entry point to the web scraping application, because it is what checks whether or not the script is called directly, rather than imported elsewhere as a module. If the current script happens to be the main one, then it calls the main function, thereby initiating the whole web-scraping process. This design allows clear isolation of the execution logic as well as other possible imports clearly organized, hence easier code maintenance and readability. This convention ensures the scrap operation of the program only runs when it will actually run the script so will be friendly and user friendly, even for people not trained in programming. This structure is commonly found in Python and also helps build robust applications, which are also modular in their nature. ## Conclusion This guide gives an in-depth overview of developing a well-structured and efficient Amazon smartwatch web scraping system for time-series analysis. By following a two-stage process, the system is able to gather data accurately—first by web scraping product links and then web scraping detailed product information using BeautifulSoup and Playwright. The use of SQLite for storage allows for efficient tracking of product information over time, and powerful error-handling processes such as failed URL logging and retry procedures make the scraping process more reliable. Moreover, using user-agent rotation, random delays, and rate limiting makes the web scraping process remain ethical and sustainable. Connect with[ Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. FAQ SECTION 1\. Is it legal to scrape Amazon smartwatch data for time-series analysis? Web scraping legality depends on Amazon's terms of service and local laws. We recommend scraping only publicly available data and complying with ethical guidelines to avoid legal issues. 2\. What kind of smartwatch data can be scraped from Amazon? You can scrape product details like price history, reviews, ratings, stock availability, and bestseller rankings to perform time-series analysis on pricing trends and consumer demand. 3\. How often should I scrape smartwatch data for an accurate time-series analysis? The frequency depends on your analysis goals. For tracking price changes or stock availability, scraping daily or hourly may be useful. For long-term trends, weekly or monthly scraping might be sufficient. 4\. Can you provide an automated solution for scraping Amazon smartwatch data? Yes, as a web scraping service provider, we offer customized solutions for automated data extraction, ensuring efficient and structured data collection for your time-series analysis needs. ### How to Scrape eBay for Rolex Watches Over $15,000? URL: https://www.blog.datahut.co/post/how-to-scrape-ebay-for-rolex-watches-over-15-000/ Last updated: 2026-07-23T07:48:33.000Z ## Introduction Did you know web scraping can uncover hidden insights about online marketplaces, like how prices fluctuate or when products are most likely to go on sale? Web scraping is your personal assistant in tech, which automatically pulls information from websites for you. Think of this as some kind of robot that can navigate through web pages in order to find what you need, keeping it neat for you without the tiring work of copying and pasting. Instead of having to dedicate many hours to gathering data, we leave it to code to get it done in a quick and more accurate manner. In this project, we plunge into eBay, the enormous online marketplace where people buy and sell just about everything one could think of. More precisely, we will narrow our focus to Rolex watches that cost in excess of $15,000\. Why Rolex? It is among the most iconic luxury watch brands, well-renowned for its high-quality, prestigious timepieces. Watches within this price bracket are not only time-telling gadgets but collectibles, investments, and a sign of status. Scraping data about these ultra-luxury Rolex watches provides a view into one of the most intriguing slices of the luxury watch market on eBay: an opportunity to reveal trends and interesting patterns that help understand what makes high-end items desirable. ### Our Two-Step Scraping Process We are dividing our web scraping task into two major steps: - To begin with, we have specialised tools that search eBay to create a list of web addresses (URLs) for the sales of Rolex watches with a price of more than $15,000. - Getting the Details: We go to all of these URLs to collect relevant information for each watch, including its price, condition, and features. ### Technologies We're Using In our project on web scraping, we are applying various specialised technologies. Each of them plays its unique role in helping us collect and store data about Rolex watches found on eBay. Let's take a closer look at each of these tools: Scrapy is our primary web scraping framework. Think of Scrapy as a type of robot that can read web pages. We do tell Scrapy what to fetch, and it goes into the website and collects that for us. Scrapy can handle really loads of web pages really quick. It is like having a super fast reader who can go through hundreds of pages in minutes only. Scrapy is written in Python, which is one of the widely used programming languages, a favorite among many programmers, as they find it easier to understand and write. It means that we describe how we'll scrape data from web pages and how to process and store it using the Scrapy. Playwright is a very powerful browser automation tool. Many websites, such as eBay, make heavy use of JavaScript, which can cause pages to change dynamically without reloading. Basic web scrapers really struggle with handling that kind of thing. Playwright is like a web browser we can control through our code. It can click buttons, scroll down pages, and perform other actions just like a real person would, which allows it to scrape information from sections of the page that otherwise might not be viewable. Scrapy-Playwright is a special plugin which enables smooth cooperation between Scrapy and Playwright. This is the bridge between these two tools. Using Scrapy-Playwright, we can apply the full power of Scrapy in terms of defining our scraping logic, as well as use Playwright when tackling dynamic web pages. This combination is particularly useful when scraping modern websites like eBay that have lots of interactive elements. SQLite is our database. We will also need a place to put all the data that Scrapy and Playwright help collect. A filing cabinet for our computer is what SQLite represents. It keeps all of the information we have collected neatly organized so it doesn't get lost later. SQLite is easy to use and does not require a separate server, which makes it an ideal choice for projects like ours. We could use SQLite to save all the detail concerning the Rolex we find so that it shall be easy to work with this data later. These technologies work together in our project like a well-oiled machine. Scrapy defines how we extract data, interaction with the website as a real user would, these capabilities combine thanks to Scrapy-Playwright, and SQLite stores all the data we collect. Working with all of these tools, we can gather various pieces of information about prices of Rolex watches on eBay and organize them for further analysis. ### Data Cleaning Data cleaning is part and parcel of our data scraping process. When we collect data from websites, it often comes in inconsistent formats or may contain errors. To make our data more useful and accurate, we use the specific cleaning and organization tools on it: OpenRefine is a powerful tool for working with messy data. It's kind of like a smart spreadsheet that can automatically detect and correct common data issues. With OpenRefine, we can easily standardize formats, correct spelling errors, and even merge similar entries. This tool is particularly useful for cleaning up text data, like product descriptions or seller information. Pandas is a Python library that is great for data manipulation and analysis. Imagine the Swiss Army knife which can handle innumerable data formats, perform complex calculations, and even visualise our data. We use Pandas to clean numerical data — prices, for example — deal with missing values, and transform our data into a form that's easy to analyze. By applying these tools, we are assured that the data collected about Rolex watches is valid, consistent, and ready for analysis. We'll then proceed with the setup of a Scrapy project and explain how to create a Scrapy project. We will be starting an adventure in web scraping, where we will be collecting and analyzing data about luxury watches on eBay. ## Initial Setup Before starting our eBay scraping project, we'll need to set up our development environment. We'll be using Scrapy as the main web scraping framework, along with Scrapy Playwright for handling dynamic content. Open our terminal or command prompt, and if we are using a virtual environment (which is recommended for all Python projects), activate it. Then, execute the following in the terminal or command prompt: ``` pip install scrapy scrapy-playwright ``` This command installs both Scrapy and the Scrapy Playwright integration. Scrapy is our core scraping tool, while Scrapy Playwright lets us deal with JavaScript-rendered content, which is ubiquitous on modern websites like eBay. Next, we want to install the browser binaries that Playwright will use. Run this command: ``` playwright install ``` This downloads and installs the browser engines that need to simulate real browser behavior-which are Chromium, Firefox, and WebKit-needed for scraping sites with dynamic content. ## Starting the Project Now that our tools are installed, let's set up the Scrapy project structure. Within our terminal navigate to the location we want our project created. Then run: ``` scrapy startproject ebay_watches ``` The command creates a new directory called \`ebay\_watches\` with Scrapy's default project structure. It includes a \`scrapy.cfg\` file and an \`ebay\_watches\` subdirectory containing several python files. The idea behind this structure is that our code will be better organized, following conventions set by Scrappy. Now, after setting up the project, let's get into the newly created directory ``` cd ebay_watches ``` The scraper will be divided into two spiders. The first spider is responsible for gathering URLs from the Rolex watch page. The second spider then scrapes the detailed data from those URLs. To run the first spider that collects product URLs, type in this command: ``` scrapy crawl watches ``` This assumes we've already created a spider named \`watches\` in our \`spiders\` directory. It should be one that walks eBay's search results for Rolex watches over $15,000 and scrape the URLs of individual product pages. Now, when we have our collection of URLs, we run our second spider to scrape out all the detailed product data: ``` scrapy crawl watches_data ``` This spider, which we will create and name \`watches\_data\`, should read the URLs found by the first spider and visit each product page extracting such information as price, condition, and features. Don't forget to create \`spiders\` directory with both spider files, \`watches.py\` and \`watches\_data.py\`. Every spider must define in turn logic for page navigation and data extraction. By following these steps, we will have a basic Scrapy project set up and ready to scrape data from eBay. The two-spider approach allows for efficient data collection, gathering first of all the URLs and then the more detailed information. As we build our spiders, we will add more specific logic for navigating eBay's pages and extracting the very data we need about the Rolex watches. ## Scraping Product Urls In this chapter, we will discuss the details of implementing our product URL scraper for Rolex watches selling above $15,000 on eBay. We have used the Scrapy-Playwright framework; this framework integrates the mighty capabilities of Scrapy along with the potency of handling dynamic content through Playwright. Our scraping process consists of three major components: - Spider Code (watches.py) - This spider would have all its logic regarding crawling of web pages, extracting product URLs, and yielding items with a 'product\_url' field for every found URL. - Pipeline (pipelines.py) - The configuration file would include the configuration to turn on SQLitePipeline and set in that configurations the database and table names to be used to store the product URLs. - Settings (settings.py) - This file would carry the SQLitePipeline class, which deals with SQLite database connection along with storing unique product URLs. Let's examine each component in detail: ### Spider Code ``` import scrapy from scrapy_playwright.page import PageMethod ``` In programming, we often need to use tools that other people have created. That is what we are doing here. We are telling our program to use Scrapy, which is like a Swiss Army knife for web scraping. It has lots of useful tools built-in that make our job easier. We also import something called PageMethod from scrapy\_playwright. Playwright will be a tool that will help Scrapy handle websites that make very heavy use of JavaScript, like eBay's site. Think of Playwright as a puppet master for web browsers - we could control them and make them do what we want. By importing these tools, we get our toolkit ready for the scraping task. ``` class EbayWatchesSpider(scrapy.Spider): """ Spider for scraping product URLs of Rolex watches listed on eBay for over $150,000. This spider uses Scrapy Playwright to handle JavaScript rendering and interact with the web page dynamically. It starts by loading the initial page, waits for the product listings to load, and extracts the URLs of individual product listings. It also handles pagination to scrape product URLs from multiple pages. The extracted URLs are saved to an SQLite database. """ name = "watches" allowed_domains = ["ebay.com"] start_urls = ["https://www.ebay.com/b/Rolex-Watches/31387/bn_2989578?LH_BIN=1&rt=nc&_udlo=15%2C000&mag=1"] ``` Here we are creating a blueprint for our spider. In programming lingo, we call this a class. Our class, EbayWatchesSpider, based on Scrapy's Spider class, means it acquires all the basic spider abilities Scrapy provides. Now, we have to add some specific information for our spider. We will name it "watches," how Scrapy is going to refer to this spider. The allowed\_domains tells our spider it's only allowed to visit pages on ebay.com-this is how we limit our spider. Finally, start\_urls is where we tell our spider where to begin its journey. That long URL is a specific eBay search for Rolex watches priced over $15,000\. It's like giving our spider a starting point on a map. ``` def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta=dict( playwright=True, playwright_include_page=True, playwright_page_methods=[ PageMethod("wait_for_selector", "div.s-item__wrapper.clearfix"), ], ), callback=self.parse, ) ``` This is the first step of the spider. It takes our starting URL defined earlier and makes a very special kind of request. This is not just any kind of request - this is a request that uses Playwright to manipulate a web browser. We instruct Playwright to wait for it to see this part of the page (div.s-item\_\_wrapper.clearfix) before it considers the page ready to be scraped. This is important because eBay contains JavaScript, which will load most of its content, and we won't scrape until we're sure those product listings are there. It's like saying to someone, "Don't read until the page is fully loaded." After establishing this special request, we inform it to use our parse function - which we'll talk about next - to handle the response. This function is returning the request, which in Scrapy means it's passing the request off to be processed. ``` async def parse(self, response): """ Parses the response to extract product URLs and handle pagination. This method waits for the page content to load, simulates scrolling to trigger lazy-loading of additional content, and then extracts product URLs from the page. If there are more pages to scrape, it recursively follows the pagination links to continue extracting URLs. Args: response (scrapy.http.Response): The response object containing the page content. """ page = response.meta["playwright_page"] # Simulate scrolling down the page to trigger lazy-loaded content await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") # Wait for any lazy-loaded content to appear await page.wait_for_timeout(5000) # Wait 5 seconds after scrolling # Wait for the product listings to load on the page try: await page.wait_for_selector("div.s-item__wrapper.clearfix", timeout=60000) except Exception as e: self.logger.error(f"Error waiting for products to load: {e}") await page.close() # Close the page to free resources return # Optional: Log the HTML content length for debugging purposes html = await page.content() self.logger.info(f"HTML content length: {len(html)}") # Extract and yield product URLs from the page for product in response.css("div.s-item__wrapper.clearfix"): product_url = product.css("div.s-item__image-section > div > a::attr(href)").get() if product_url: yield {"product_url": product_url} # Yield the product URL as an item # Handle pagination to go to the next page, if available next_page = response.css("a.pagination__next::attr(href)").get() if next_page: self.logger.info(f"Next page URL: {next_page}") yield scrapy.Request( response.urljoin(next_page), callback=self.parse, meta=dict( playwright=True, playwright_include_page=True, playwright_page_methods=[ PageMethod("wait_for_selector", "div.s-item__wrapper.clearfix"), ], ), ) else: self.logger.info("No more pages to process.") # Close the page after processing to free up resources await page.close() ``` This is where the magic of our actual parse function begins to unravel. It is like a set of instructions for our spider to crawl the page on eBay and collect the information we want. Now, let's break it down step by step: First, our spider goes to the bottom of the page. This is important because some websites, like eBay, load more content as we scroll down. It's like checking to make sure we unrolled the whole scroll before we read it. Then, it waits 5 seconds, or 5000 milliseconds so any new content has a chance to appear. Next, we instruct our spider to wait for the listings to display. We give it up to a minute to find these listings. If it still couldn't find any after a minute, then it logs an error message-you know, writing a note in its diary-and stops working on this page, a safety measure to prevent our spider from being stuck in such situations. Our spider will now start looking for product URLs if everything has loaded correctly. It does this by going through the HTML of the page looking for specific patterns as to how eBay structures their product listings, for each and every one of these product URLs that it finds it harvests them. This is the main purpose of our spider: collecting those URLs. After it's iterated over all the items on the page, the spider looks for the "Next Page" link. If it finds such a thing, it makes up a new request to visit that page, and the whole process repeats over again on the new page. This is how our spider can iterate over all the pages of the search results, not just the first one. Now, if there is no "Next Page" link, our spider realizes that it has gotten to the very end of the result set. It logs a message saying it's done, like leaving a note saying "finished reading". Finally, our spider closes the web page. This is like closing a book when we're done reading - it helps keep things tidy and frees up computer resources. During this, our spider is using async operations (as for what 'async' and 'await' keywords are). A little like multitasking-it allows our spider to do other things while waiting for pages to load or while waiting for something to complete and be ready for the next action. Our spider uses this parse function as its heart to navigate through eBay's search results, collect all the URLs for Rolex watches, and make sure it goes through all pages of results. Using Playwright, we can interact with eBay's website as a human user would, scrolling and waiting for content. So we are able to get accurate and complete results from our scraping. ### Settings Code ``` # Bot and spider configuration BOT_NAME = "ebay_watches" SPIDER_MODULES = ["ebay_watches.spiders"] NEWSPIDER_MODULE = "ebay_watches.spiders" ``` This part sets up the basic structure of your Scrapy project. It tells Scrapy what to call the bot and where to find spider code. The bot name is "ebay\_watches" and spiders are located in the folder "ebay\_watches.spiders". It helps Scrapy know how to organize your web scraping code. ``` # URL scraping settings DOWNLOAD_DELAY = 2 # Time delay between requests CONCURRENT_REQUESTS = 1 # Number of concurrent requests ``` These settings control how the spider behaves when making requests. It sets a delay of 2 seconds between requests and limits the spider to one request at a time. This helps prevent overloading the target website and makes the scraping more polite. ``` # Configure Playwright as the download handler DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } TWISTED_REACTOR = 'twisted.internet.asyncioreactor.AsyncioSelectorReactor' ``` This section establishes Playwright as the tool for web-page downloads. Playwright is a browser automation tool that functions better with dynamic websites than Scrapy alone. It is being utilized for both HTTP and HTTPS requests. Additionally, it specifies a special reactor for dealing with asynchronous operations when working with Playwright. ``` # Set default headers for requests DEFAULT_REQUEST_HEADERS = { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "Accept-Language": "en", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36", } ``` These settings define the information sent with each request to make a spider appear more like a real web browser. It may include things like accepted content types, language, as well as a user agent string. This, therefore helps in avoiding those sites that block a spider through a technique called anti-scraping. ``` # Playwright-specific settings PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT = 60 * 1000 # 60 seconds timeout PLAYWRIGHT_BROWSER_TYPE = "chromium" PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": False, # Run browser in non-headless mode } ``` These options configure how play works. That is, it sets a timeout for loading the page and sets the type of browser to Chromium, and sets headless mode to off, meaning that when the spider is running, we can actually see the browser window, which may aid in debugging if needed. ``` # Scraping behaviour settings ROBOTSTXT_OBEY = False # Don't obey robots.txt rules RETRY_ENABLED = True RETRY_TIMES = 5 # Number of retries for failed requests RETRY_HTTP_CODES = [500, 502, 503, 504, 408] # HTTP codes to retry on ``` This section determines how the spider behaves. It disables respect for robots.txt files, which is not very polite but may be required for certain projects. It also initializes retry behaviour for failed requests. Each failed request due to certain server errors would be retried up to 5 times. ``` # AutoThrottle settings AUTOTHROTTLE_ENABLED = True AUTOTHROTTLE_START_DELAY = 5 AUTOTHROTTLE_MAX_DELAY = 60 AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0 ``` AutoThrottle automatically throttles the spider's download speed based on what the website responds with. It prevents downloading too much off of that site and getting banned. It starts at a 5-second delay that can go up to 60 seconds if needed. ``` # Miscellaneous settings REQUEST_FINGERPRINTER_IMPLEMENTATION = "2.7" FEED_EXPORT_ENCODING = "utf-8" LOG_LEVEL = 'INFO' ``` These are various other settings for Scrapy. They set the encoding for exported data, the logging level, and a specific implementation for request fingerprinting. ``` # SQLite pipeline settings ITEM_PIPELINES = { 'ebay_watches.pipelines.SQLitePipeline': 300, } SQLITE_DB = 'ebay_watches.db' SQLITE_TABLE = 'product_urls' ``` This last section creates a pipeline to store the scraped data into a SQLite database. It designates what code to use in the data handling (the SQLitePipeline) and sets the name of the database and table to be used. This means the spider will store the scraped information into a local database to be retrieved at a later time. ### Pipelines Code ``` import sqlite3 ``` The code starts by importing the sqlite3 module. This is a crucial step as it brings in all the necessary tools to work with SQLite databases in Python. ``` class SQLitePipeline: """ A Scrapy pipeline for storing product URLs in a SQLite database. This pipeline creates a SQLite database (if it doesn't exist) and a table to store unique product URLs. It handles the database connection, insertion of new URLs, and proper closure of the database connection. Attributes: db_name (str): The name of the SQLite database file. table_name (str): The name of the table to store product URLs. conn (sqlite3.Connection): The SQLite database connection. cursor (sqlite3.Cursor): The database cursor for executing SQL commands. """ ``` Now, we define the SQLitePipeline class. It is designed to work within the Scrapy framework as a pipeline for processing scraped data. In Scrapy, pipelines are aimed at processing items right after they have been scraped by spiders. In this case, our pipeline is devoted to storing URLs of products in a SQLite database. A high-level docstring in the class describes exactly what the pipeline does. This is particularly useful for developers who may need to use or maintain this code. In this case, the pipeline does the following: it will create, if they do not already exist, a SQLite database and a table where unique product URLs will be stored. Docstring also lists primary attributes of the class, which will be handy for quick reference as to what data the class will work with. ``` def __init__(self, db_name, table_name): """ Initialise the SQLitePipeline. Args: db_name (str): The name of the SQLite database file. table_name (str): The name of the table to store product URLs. """ self.db_name = db_name self.table_name = table_name ``` The init method in our SQLitePipeline class is its constructor, which gets called during the creation of a new instance of the class. There are two parameters for this method, namely db\_name and table\_name. These have enabled us to specify the name of the SQLite database file and the name of the table where we would be storing our product URLs. Accepting these as parameters makes our pipeline more flexible-it may be easily used with different database and table names without changing the code at all. It just stores these values as instance attributes, and they are thus available to be used in the other methods in the class. ``` @classmethod def from_crawler(cls, crawler): """ Create a pipeline instance from a Crawler. This class method is used by Scrapy to create an instance of the pipeline. It uses the Scrapy settings to get the database and table names. Args: crawler (scrapy.crawler.Crawler): The crawler that uses this pipeline. Returns: SQLitePipeline: An instance of the pipeline. """ return cls( db_name=crawler.settings.get('SQLITE_DB', 'ebay_watches.db'), table_name=crawler.settings.get('SQLITE_TABLE', 'product_urls') ) ``` from\_crawler designate the method as class method, denoted by decorator @classmethod. Scrapy uses this to create an instance of our pipeline using the Crawler object. In Scrapy, all related information to the scraping process are kept in the Crawler object. This method fetches database and table names from the settings of the crawler object, with fallback default values when such settings are not defined. This method allows configuring the pipeline through the settings of the Scrapy project. This pattern is common in Scrapy projects. This approach makes the pipeline more flexible and easier to configure without changing the code. ``` def open_spider(self, spider): """ Open database connection when spider is opened. This method is called by Scrapy when the spider is opened. It establishes a database connection and creates the table if it doesn't exist. Args: spider (scrapy.Spider): The spider being opened. """ # Establish database connection self.conn = sqlite3.connect(self.db_name) self.cursor = self.conn.cursor() # Create table if it doesn't exist self.cursor.execute(f''' CREATE TABLE IF NOT EXISTS {self.table_name} ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_url TEXT UNIQUE ) ''') self.conn.commit() ``` open\_spider method is called by Scrapy when a spider gets run. This is where we create our database connection. In order to create our database connection, we use sqlite3.connect function that connects us to our SQLite database file. SQLite will automatically build the file if it doesn't already exist. We also acquire a cursor object that we will make use of to execute SQL commands. We then run a SQL command to create our table if it doesn't already exist. The table has two columns: an auto-incrementing ID and a unique product URL. By using "IF NOT EXISTS" in our SQL, we ensure that we don't get errors if the table already exists from a previous run. Finally, we commit our changes to the database. This helps ensure that our database and table are ready to accept data before the spider begins its crawl. ``` def close_spider(self, spider): """ Close database connection when spider is closed. This method is called by Scrapy when the spider is closed. It ensures that the database connection is properly closed. Args: spider (scrapy.Spider): The spider being closed. """ self.conn.close() ``` The close\_spider method is the counterpart of open\_spider. Scrapy calls it when a spider finishes running. Its job is simple but important: closing the database connection. This is crucial in proper resource management. If we didn't close the connection, we might leave the database in some possibly inconsistent state or resource leakages. Obviously, closing the connection explicitly when we're done with it ensures that all our data is saved properly and that we're good stewards of system resources. ``` def process_item(self, item, spider): """ Process a scraped item. This method is called for every item pipeline component. It inserts the product URL into the database if it's not already present. Args: item (scrapy.Item): The item scraped by the spider. spider (scrapy.Spider): The spider which scraped the item. Returns: scrapy.Item: The processed item. """ # Insert the product URL into the database, ignoring if it already exists self.cursor.execute(f''' INSERT OR IGNORE INTO {self.table_name} (product_url) VALUES (?) ''', (item['product_url'],)) self.conn.commit() return item ``` The heart of our pipeline is the process\_item method. Scrapy calls this method for every item our spider scrapes. In this method, we take the product URL from the scraped item and insert it into our database. We use a SQL INSERT OR IGNORE statement, which is a nice SQLite feature. This will add the URL to the database if it is not already there but will not give an error if it is already in the database. This is perfect for ensuring we only store unique URLs. After executing the insert, we commit the change to the database to ensure it's saved. Finally, we return the item, which allows it to be processed by any subsequent pipelines in the Scrapy project. This SQLitePipeline is the right way to get unique product URLs and store them as part of a Scrapy spider workflow. The SQLitePipeline gets around database connection issues, maintains data integrity, and blends well with Scrapy's pipeline system. ## Scraping Product Data We will now go into the implementation of our products data scraping system for Rolex watches priced over $15,000 in eBay. As we have previously discussed, we are utilizing the Scrapy framework, to efficiently extract and organize data from website. We define our scraping process in three main components: - Spider(watches\_data.py): The spider would comprise the logic for crawling eBay watch listings, extracting detailed information about each watch, and yields EbayWatchesItem instances featuring url, title, sale\_price, price, discount, condition, shipping\_charge, returns, and details. - Items(items.py): The items.py contains a class called EbayWatchesItem inheriting from scrapy.Item. All those fields are declared within this class for url, title, sale\_price, price, discount, condition, shipping\_charge, returns, and details regarding such eBay watch listings. - Settings(settings.py): there should be configurations to enable the SQLitePipeline, database and table names for saving the watch data, etc; probably with settings specific for scraping on eBay (for example, about user agents or request delay). - Pipelines(pipelines.py): The pipelines.py file would hold the SQLitePipeline class. This class should be implemented such that it can handle all fields in EbayWatchesItem, not just the product URL. It will manage connections to the SQLite database and hold complete watch listing data. Let's examine each part in detail: ### Spider Code ``` import scrapy import sqlite3 from ebay_watches.items import EbayWatchesItem ``` This section is like packing our bag before a trip. We're bringing in the tools we need for our web scraping job. Scrapy is our main tool for web scraping. It's like a swiss army knife for getting information from websites. SQLite3 helps us work with our database, where we stored the links to the watch pages. It's like a filing cabinet where we keep important information. EbayWatchesItem is something we created earlier to hold all the details about each watch. Think of it as a form we'll fill out for each watch we find. ``` class SQLiteUrlSpider(scrapy.Spider): """ A spider that crawls eBay watch listings using URLs stored in a SQLite database. This spider fetches unscraped URLs from a specified table in the database, crawls each URL, extracts relevant information about the watch listing, and yields an EbayWatchesItem for each listing. Attributes: name (str): The name of the spider. db_name (str): The name of the SQLite database file. url_table (str): The name of the table containing product URLs. """ name = "watches_data" def __init__(self, *args, **kwargs): """ Initialize the SQLiteUrlSpider. Args: *args: Variable length argument list. **kwargs: Arbitrary keyword arguments. """ super().__init__(*args, **kwargs) self.db_name = 'ebay_watches.db' self.url_table = 'product_urls' ``` Our SQLiteUrlSpider is a specialized robot designed for a specific task. We give it a name, "watches\_data", and this name is how we'll refer to this spider when we want to use it. This name is important because if we have multiple spiders, we need to identify them. It is just like giving a name to each of our tools, and that's why we know which to grab when we need it. When we make our spider, we have to initialize it with some basic info. This is done in the \_\_init\_\_ method. Think of it as like a setup manual for the spider. We give it two important pieces of information: the name of our database file, (ebay\_watches.db); and the name of the table in that database where we stored our URLs (product\_urls). We're essentially giving our robot a location to find its list of tasks. The database is just as if it were a huge book of information, and the table is basically a page in that book where we have all the web addresses we would like to visit. We're making our spider smart and efficient by programming it this way. Instead of hardcoding a list of URLs or searching the entire eBay website, our spider knows exactly where to look for the information it needs to start its job. This also makes our spider flexible enough; if we wish to scrape a different set of URLs in the future, we will only need to change the database or table name and our spider will automatically update without our having to re-write its code. ``` def start_requests(self): """ Generate initial requests for the spider. This method connects to the SQLite database, fetches unscraped URLs, and yields a Request object for each URL. Yields: scrapy.Request: A request object for each unscraped URL. """ conn = sqlite3.connect(self.db_name) cursor = conn.cursor() # Fetch unscraped URLs from the database cursor.execute(f"SELECT product_url FROM {self.url_table} WHERE scraped = 0") urls = cursor.fetchall() conn.close() for url in urls: yield scrapy.Request(url=url[0], callback=self.parse, errback=self.errback) ``` The start\_requests function is where our spider actually starts its journey. Imagine this as the spider waking up and checking its to-do list. It first opens up the database we told it about earlier. That is like opening a book to the exact page where we wrote down all the web addresses we want to visit. The spider then requests the database to provide it with all those URLs that have yet to be scrapped. Just like when one goes through their checklists and crosses out those which are completed to only view the ones that are yet to be done. Once the spider has its list of URLs, it closes the database. Good practice: we don't leave our book lying open for when we are done with it, let's close it-thus making things neat and preventing any accidental changes. Then the spider makes a special request for every URL on the list. Every request is like a particular mission: "Go to this web page, and when we're done, apply the 'parse' function to understand what we found." The spider also has a backup plan: if something bad happens while visiting a page, it knows how to use the 'errback' function to handle the problem. This function is very important because it is where we transform our list of URLs into actual web scraping activities. It is efficient because it only looks at URLs it hasn't scraped before, thus saving time and preventing duplicate work. By yielding each request one at a time, we're also being gentle with the website we're scraping - instead of bombarding it with all our requests at once, we're spacing them out, which is more polite and less likely to get us blocked. ``` def parse(self, response): """ Parse the response and extract watch listing information. This method creates an EbayWatchesItem and populates it with data extracted from the response using CSS selectors. Args: response (scrapy.http.Response): The response to parse. Yields: EbayWatchesItem: An item containing the extracted watch listing information. """ item = EbayWatchesItem() # Extract data using CSS selectors item['url'] = response.url item['title'] = response.css('div.vim.x-item-title > h1 > span::text').get() item['sale_price'] = response.css('div.x-price-primary > span::text').get() item['discount'] = response.css('div.x-price-transparency > span.x-price-transparency--discount > span.ux-textspans.ux-textspans--EMPHASIS::text').get() item['price'] = response.css('div.x-price-transparency > span.x-price-transparency--discount > span.ux-textspans.ux-textspans--SECONDARY.ux-textspans--STRIKETHROUGH::text').get() item['condition'] = response.css('#mainContent > div.vim.d-vi-evo-region > div.vim.x-item-condition.mar-t-20 > div.x-item-condition-text > div > span > span:nth-child(1) > span::text').get() item['shipping_charge'] = response.css('#mainContent > div.vim.d-vi-evo-region > div.vim.d-shipping-minview.mar-t-20 > div > div > div > div:nth-child(1) > div > div > div.ux-labels-values__values.col-9 > div > div > span.ux-textspans.ux-textspans--BOLD::text').get() item['returns'] = response.css('#mainContent > div > div.vim.x-returns-minview.mar-b-20 > div > div > div > div > div > div.ux-labels-values__values.col-9 > div > div > span:nth-child(1)::text').get() # Extract additional details using a separate method item['details'] = self.extract_details(response) yield item ``` Our spider mainly does the following work in the parse function. After visiting any web page, this helps it to understand what exactly it sees. To begin with, it creates a brand new EbayWatchesItem. Think of that as a form we're going to fill out about the watch. Think of that like a checkbox list of all the information we want to collect. The URL, the title of the listing-the price, and so forth. We employ a thing called CSS selectors when completing this form. It's like we're sending some instructions on where to locate certain bits of information on the page. Imagine telling our spider, "The title is in the huge header at top," or "The price is within the box on the right-hand side of the page." The spider then follows those instructions and writes down the information found. This is a very effective technique because it enables us to target what we want specifically, with the rest of the page being irrelevant. Then, we use the extract\_details function to grab still more details about the watch. This is like turning the form over and doing the back side, filling in all the additional detail. Last but not least, we hand that completed form over to Scrapy using yield. This is like turning in our completed homework. Scrapy will now know what to do with it next, whether it's saving it to a file or passing it through for further processing, or putting it in a database. ``` def extract_details(self, response): """ Extract additional details from the watch listing page. This method parses the details section of the listing page and creates a dictionary of key-value pairs for additional watch details. Args: response (scrapy.http.Response): The response to parse. Returns: dict: A dictionary containing additional details about the watch. """ details_dict = {} details = response.css('div.ux-layout-section-evo__row > div.ux-layout-section-evo__col') for div in details: dt_text = div.css('dt.ux-labels-values__labels span.ux-textspans::text').get() dd_text = div.css('dd.ux-labels-values__values span.ux-textspans::text').get() # Only add to the dictionary if both key and value are present if dt_text and dd_text: details_dict[dt_text.strip()] = dd_text.strip() return details_dict ``` The extract\_details function is like a detective searching for clues. Its job is to find and organize all the extra information eBay has available about the watch. This is typically found in a certain section of the webpage, often in a table or list. Our function knows where to find that treasure trove of details. This process will pass over this portion of the page, looking for an information pair-a keyword like "Brand" or "Model" and its corresponding value like "Rolex" or "Submariner". It is reading a list of facts about the watch. Each time it comes across one of these pairs, it adds it to a special list called a dictionary. In our dictionary, the label becomes the key, and the value becomes, well, the value. This is a good way of presenting the information because, later on, we will easily be able to look up specific details about the watch. Our function then returns this dictionary after going through all the details on the page. This is like handing over a neatly organised list of facts about the watch. Back in our main parse function, we add this dictionary to our EbayWatchesItem under the 'details' field. This way, we capture both the main information about the listing - like price and condition - and all these extra details in one comprehensive package. ``` def errback(self, failure): """ Handle errors that occur during the crawling process. This method is called when an error occurs while processing a request. It logs the error and updates the crawler statistics. Args: failure (Failure): A twisted.python.failure.Failure object that encapsulates the error information. """ # Extract error information url = failure.request.url error_message = str(failure.value) # Update crawler statistics self.crawler.stats.inc_value('failed_urls') self.crawler.stats.inc_value(f'failed_urls/{failure.value.__class__.__name__}') # Log the error self.log(f"Error on {url}: {error_message}") # Use the pipeline to log the error if the method exists pipeline = self.crawler.engine.scraper.itemproc._middleware[0] if hasattr(pipeline, 'log_error'): pipeline.log_error(url, error_message) ``` The errback function is the safety net to our spider. After all, things don't always go according to plan when we have to deal with web scraping. The page might not load, the website could be down, or it mightn't work in the format that we expect. This function catches those problems and handles them gracefully. If there's an error, this function springs into action, marking down what URL caused the error and what the error was itself. It is like a log of issues we encounter. It's extremely useful for debugging later on; we can go back and see exactly what went wrong and where. The function also updates certain counters so that we can track how many errors we've had and what sort they are. This gives us a big-picture view of how well our scraping is going. Finally, if we've established an special way of recording errors (that's the 'pipeline' part), it uses that to make a detailed note of the error. This might involve writing to a log file or updating a database. The point here is that we're not ignoring the fact that, somewhere along the line, something bad occurred. We're acknowledging them, recording them, and setting ourselves up to deal with them. This would make our spider stronger and more reliable enough to cope with the nonpredictable nature of web scraping. ### Items Code ``` import scrapy class EbayWatchesItem(scrapy.Item): """ Defines the structure for storing data about a Rolex watch listing on eBay. This item contains fields for various details of a watch listing, including its URL, title, pricing information, condition, shipping details, and more. Each field is defined as a scrapy.Field(), which allows Scrapy to process and store the data efficiently. """ # The URL of the eBay listing url = scrapy.Field() # The title of the watch listing title = scrapy.Field() # The current sale price of the watch sale_price = scrapy.Field() # The original price of the watch (if discounted) price = scrapy.Field() # The discount amount or percentage (if applicable) discount = scrapy.Field() # The condition of the watch (e.g., "New", "Used") condition = scrapy.Field() # The shipping charge for the watch shipping_charge = scrapy.Field() # The return policy for the watch returns = scrapy.Field() # Additional details about the watch (stored as a dictionary) details = scrapy.Field() ``` In our web scraping project, we need a method to structure the information we collect about each Rolex watch. We do this using something called an EbayWatchesItem. Think of this item as a special container or a form that has a spot for each piece of information we want to gather. We create the container by defining a class inheriting from scrapy.Item, which has some special powers when working with Scrapy. Inside this class, we list out all the pieces of information we wish to collect: for instance, the URL of the listing, the title of the watch, its price, condition, etc. Each of these is called a field, and we create them using scrapy.Field(). Thus, setting up our data like this will make all the watches easier to scrape, as it provides a standard structure that we can easily populate upon finding a watch. It will be as if we fill out one standard form for all of them. In addition, this will make it far easier to deal with all this data afterwards, as we do know that every watch will have a 'title' field, a 'price' field and more. We also have a special 'details' field that can contain extra information, which might vary for each watch. This setting keeps our data organized in that we can store it without much hassle or easily analyze it later to perform whatever operations we need with the information once collected. ### Settings Code ``` import random # Bot and spider configuration BOT_NAME = "ebay_watches" SPIDER_MODULES = ["ebay_watches.spiders"] NEWSPIDER_MODULE = "ebay_watches.spiders" ``` This sets up the bot's name and where to find the spider code. It tells Scrapy what to call our bot and where to look for the spiders that do the actual scraping. ``` # Respect robots.txt rules ROBOTSTXT_OBEY = True # Configure maximum concurrent requests performed by Scrapy CONCURRENT_REQUESTS = 1 # Configure a delay for requests for the same website (default: 0) DOWNLOAD_DELAY = 1 ``` These settings make our bot behave nicely. It follows the rules set by websites, only makes one request at a time, and waits a second between requests to avoid overwhelming the server. ``` # Disable cookies (enabled by default) COOKIES_ENABLED = False # Disable Telnet Console (enabled by default) TELNETCONSOLE_ENABLED = False ``` This turns off cookies and the telnet console. It makes our bot act less like a typical browser, which can sometimes help avoid detection. ``` # List of User-Agent strings to rotate through USER_AGENTS = [ 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_12_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/60.0.3112.113 Safari/537.36', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:57.0) Gecko/20100101 Firefox/57.0', 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:52.0) Gecko/20100101 Firefox/52.0', ] # Configure default request headers DEFAULT_REQUEST_HEADERS = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language': 'en', "User-Agent": random.choice(USER_AGENTS), } ``` This sets up a list of different browser identities. Our bot will randomly pick one of these for each request, which helps it look more like a real user. ``` # Enable and configure the AutoThrottle extension (disabled by default) # AutoThrottle extension adjusts download delays dynamically RETRY_ENABLED = True RETRY_TIMES = 5 RETRY_HTTP_CODES = [500, 502, 503, 504, 408] ``` If a request fails due to server errors, this tells our bot to try again. It will retry up to 5 times for specific error codes, which helps deal with temporary server issues. ``` # SQLite database settings ITEM_PIPELINES = { 'ebay_watches.pipelines.SQLitePipeline': 300, } SQLITE_DB = 'ebay_watches.db' URL_TABLE = 'product_urls' DATA_TABLE = 'product_data' ERROR_TABLE = 'scraping_errors' ``` This sets up a SQLite database to store the scraped data. It defines tables for storing product URLs, actual product data, and any errors that occur during scraping. ### Pipeline Code ``` import sqlite3 import json ``` Here we import two very powerful tools: sqlite3 and json. sqlite3 can be envisioned as compact, portable filing cabinet that fits inside our code; one can imagine it organizes data into tables, much like spreadsheets. The json tool is like a universal translator for complicated data. The two tools help us handle information that can be nested or of several parts by simplifying it into a straightforward string that we can store and retrieve easily. These two tools working together help us handle all forms of data in our web scraping project. ``` class SQLitePipeline: """ A Scrapy pipeline for storing scraped data in a SQLite database. This pipeline handles the storage of product data, URL tracking, and error logging in separate tables within a SQLite database. It provides methods for initializing the database connection, creating necessary tables, processing scraped items, and logging errors. Attributes: db_name (str): Name of the SQLite database file. url_table (str): Name of the table storing product URLs. data_table (str): Name of the table storing scraped product data. error_table (str): Name of the table for logging scraping errors. conn (sqlite3.Connection): SQLite database connection object. cursor (sqlite3.Cursor): SQLite database cursor object. Methods: from_crawler: Class method to create a pipeline instance from a crawler. open_spider: Opens the database connection and initializes tables. close_spider: Closes the database connection. process_item: Processes and stores a scraped item in the database. log_error: Logs an error message associated with a URL. """ def __init__(self, db_name, url_table, data_table, error_table): """ Initialize the SQLitePipeline with database and table names. Args: db_name (str): Name of the SQLite database file. url_table (str): Name of the table storing product URLs. data_table (str): Name of the table storing scraped product data. error_table (str): Name of the table for logging scraping errors. """ self.db_name = db_name self.url_table = url_table self.data_table = data_table self.error_table = error_table ``` The SQLitePipeline class and its \_\_init\_\_ method are like preparing a new office for our data processing needs. When we create a new instance of this class, we are, in effect, throwing open the shop doors. The class's init method is like the blueprint for our office layout. We choose to name our main database (db\_name), which is like choosing the building for our office. Then we specify some names for different tables - url\_table, data\_table, error\_table - as if they were different departments in our office. In a table dedicated to one particular task, for instance, one table tracks the URLs we visit, another stores the actual data we obtain, and another tracks all errors encountered. By setting them up in the initializer, we are ensuring that every instance of our SQLitePipeline has everything prepared for dealing with all aspects of our data management needs. ``` @classmethod def from_crawler(cls, crawler): """ Create a pipeline instance from a crawler. This class method allows Scrapy to instantiate the pipeline with settings defined in the crawler's configuration. Args: crawler (scrapy.crawler.Crawler): The crawler instance. Returns: SQLitePipeline: An instance of the pipeline. """ return cls( db_name=crawler.settings.get('SQLITE_DB', 'ebay_watches.db'), url_table=crawler.settings.get('URL_TABLE', 'product_urls'), data_table=crawler.settings.get('DATA_TABLE', 'product_data'), error_table=crawler.settings.get('ERROR_TABLE', 'scraping_errors') ) ``` The from\_crawler method is like having a smart office manager who can set up our whole operation based on a set of instructions provided (that would be our crawler settings). This method looks into the settings of the crawler, something like a company policy document, and figures out how to set up our SQLite Pipeline. It checks for specific instructions about what to name our database and tables. If it finds these directions, it uses them. Otherwise, it has some reasonable default names. It is very useful since it will allow us to easily adjust how our pipeline is configured by only adjusting the crawler settings and not by having to dive into the code itself. It's like organizing the entire office just by updating a single document. ``` def open_spider(self, spider): """ Open database connection and initialize tables when the spider opens. This method is called when the spider is opened. It establishes a database connection, creates necessary tables if they don't exist, and adds a 'scraped' column to the URL table if it's not present. Args: spider (scrapy.Spider): The spider instance. """ self.conn = sqlite3.connect(self.db_name) self.cursor = self.conn.cursor() # Check if 'scraped' column exists in the URL table, if not, add it self.cursor.execute(f"PRAGMA table_info({self.url_table})") columns = [column[1] for column in self.cursor.fetchall()] if 'scraped' not in columns: self.cursor.execute(f''' ALTER TABLE {self.url_table} ADD COLUMN scraped INTEGER DEFAULT 0 ''') # Create the product data table if it doesn't exist self.cursor.execute(f''' CREATE TABLE IF NOT EXISTS {self.data_table} ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT UNIQUE, title TEXT, sale_price TEXT, price TEXT, discount TEXT, condition TEXT, shipping_charge TEXT, returns TEXT, details TEXT ) ''') # Create the error logging table if it doesn't exist self.cursor.execute(f''' CREATE TABLE IF NOT EXISTS {self.error_table} ( id INTEGER PRIMARY KEY AUTOINCREMENT, url TEXT, error_message TEXT, timestamp DATETIME DEFAULT CURRENT_TIMESTAMP ) ''') self.conn.commit() ``` open\_spider is kind of like the morning routine; we are opening the office for business for the first time. First, it opens a connection to our database-unlocking the main door and switching on the lights, so to speak. Then it runs a series of checks and setups. It is scanning our URL tracking table, ensuring that it has a column that marks whether a URL has been scraped or not. This column gets added if such does not exist in the table - it's somewhat like having an epiphany that we need a new filing category and then quickly doing so to make sure all is on track. It also scans for tables for the product data and errors we might encounter. If these tables don't exist yet, it will create them. It's kind of like setting up new filing cabinets if we have to because we don't have what we need. At the end of this process, our whole data storage system is now in place and operational, ready to digest whatever data our spider finds out. ``` def close_spider(self, spider): """ Close the database connection when the spider closes. Args: spider (scrapy.Spider): The spider instance. """ self.conn.close() ``` Our end-of-day routine is the close\_spider method. After all the busy work in scraping and storing data, this method closes everything down properly. Its main job is to close the connection to our database. That's very important because it ensures that all our data is properly saved as well as the closing of open connections that may cause problems later on. It's like locking all the filing cabinets, shutting down computers, and locking up the office door before we leave. It is a mundane but very important step in keeping our data safe and running smoothly with our system. ``` def process_item(self, item, spider): """ Process a scraped item and store it in the database. This method inserts or replaces the scraped item data in the product data table and updates the 'scraped' status in the URL table. Args: item (dict): The scraped item containing product data. spider (scrapy.Spider): The spider instance. Returns: dict: The processed item. """ # Convert the dictionary to a JSON string before saving it to the database details_json = json.dumps(item['details']) # Insert the item data into the product data table self.cursor.execute(f''' INSERT OR REPLACE INTO {self.data_table} (url, title, sale_price, price, discount, condition, shipping_charge, returns, details) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) ''', (item['url'], item['title'], item['sale_price'], item['price'], item['discount'], item['condition'], item['shipping_charge'], item['returns'], details_json)) # Update the 'scraped' status in the URL table self.cursor.execute(f''' UPDATE {self.url_table} SET scraped = 1 WHERE product_url = ? ''', (item['url'],)) self.conn.commit() return item ``` Then there's the process\_item method, where all the real action happens. So, every time our spider reaches out and grabs information from a website, this method processes that info and puts it away in the right place. First, it converts the data's 'details' part into a JSON-encoded string. It's kind of like taking a complex document, then summarising it so it's easy to file. Now, it gathers all the fragments of information about anything - for example, the URL of the product, its title, price, and our summarised details - and writes them to our product data table. If we have seen this product before, instead of writing a duplicate entry, it updates the information. This is similar to how we create a new file for a product or update an existing one with fresh information. In addition to storing the product data, it also updates our URL tracking table and marks this URL as 'scraped'. It is in much the same vein as checking off a task on our to-do list. ``` def log_error(self, url, error_message): """ Log an error message associated with a URL. This method inserts an error log entry into the error table. Args: url (str): The URL associated with the error. error_message (str): The error message to log. """ self.cursor.execute(f''' INSERT INTO {self.error_table} (url, error_message) VALUES (?, ?) ''', (url, error_message)) self.conn.commit() ``` We use the log\_error method to keep records of what goes wrong while we're trying to scrape. Whenever something does go wrong, like when a page refuses to load or we just can't find what we're looking for, it invokes the log\_error method. It logs both the errant URL and what exactly has gone wrong. Then it stores this information into our error table, along with the timestamp of when that error actually occurred. This essentially becomes a dedicated troubleshooting notebook where we jot down any problems we find and, importantly note exactly what happened and at what time. This can come in very handy later if we need to troubleshoot our scraper or if we want to retry those problematic URLs again. It helps learn from mistakes and continue improving the process of scraping. ## Conclusion And that's a wrap! In this blog, we learn how to scrape Rolex watches over $15,000 from eBay by using Scrapy and Playwright. We configured our spider in Scrapy, managed dynamic content with Playwright, and efficiently stored the data we scrapped. Meanwhile, we managed to handle pagination, regulate delay, and structure the data for further analysis. Web scraping might sound a bit fiddly initially, but the moment you have the right tool and approach the task in steps, it will prove to be an incredibly potent means of collecting data. In terms of suggestions for further improving this project: try to extract other watch brands, add other fields to scrape, or you could even make some analysis over it to get the trends from pricing. Hope this blog helped you understand it better! Comment below if you have any questions or suggestions. Happy coding! Connect with[ ](https://www.datahut.co/?ref=blog.datahut.co)[Datahut](https://www.datahut.co/?ref=blog.datahut.co) for top-notch web scraping services that bring you the valuable insights you need hassle-free. FAQ SECTION 1\. Is it legal to scrape eBay for Rolex watch listings? Web scraping eBay is subject to their terms of service, and scraping without permission may violate their policies. To stay compliant, use eBay’s API for structured data access or ensure that your scraping approach respects robots.txt and legal considerations. 2\. What tools are best for scraping eBay for Rolex watches over $15,000? Python libraries like BeautifulSoup and Scrapy can help scrape eBay pages, while Selenium can handle JavaScript-heavy pages. However, eBay’s API is the best option for structured data extraction, ensuring accuracy and compliance. 3\. How can I avoid getting blocked while scraping eBay? To reduce the risk of getting blocked, follow these best practices: - Use rotating proxies and user agents - Implement delays and randomized request intervals - Limit the frequency of requests to avoid detection - Prefer API access if available for reliability 4\. Can your web scraping service help extract Rolex watch data from eBay? Yes! As a web scraping service provider, we offer custom data extraction solutions to collect high-value product listings like Rolex watches. We ensure data accuracy, compliance with best practices, and automation to fetch real-time pricing, seller ratings, and listing details. Reach out to us for a tailored solution. ### Using Web Scraping to Extract Real Estate Insights URL: https://www.blog.datahut.co/post/how-to-use-web-scraping-for-real-estate-insights-from-bayut/ Last updated: 2026-07-23T07:48:33.000Z ## Introduction Would you believe me if I tell you that you can gain an iconic understanding of the comparative real estate market by utilizing web scraping? As a data analyst or a business researcher, it is now possible to acquire precise and current information from Bayut, one of the top property sites in the UAE. This blog outlines the necessary tools and techniques required in data collection and analysis that makes extracting real estate data from Bayut . ## What Is Web Scraping? Web scraping is the process of extracting information from a website in an automated manner. Instead of manually copying and pasting relevant data, users can utilize different tools, software, and scripts that will do everything for them automatically. For real estate websites such as Bayut, web scraping can be used to extract: \- Locations and specifications of properties together with their Selling prices \- Mortgate info and other related financials \- Market trends to enhance competitive analysis and decision making The reasons business adopt web scraping includes: \- Fresh listings to update their databases \- Changes in competitors’ offerings \- Market research to gain insights and information for action planning For your real estate business, here is a detailed account of how to scrape Bayut. This article explains how to get real estate data from Bayut in two simple steps . The first script is used to get property links from multiple pages on Bayut, using requests and BeautifulSoup libraries in Python to get and read HTML content. These major techniques include mimicking different user agents to make it seem like a human is visiting the website, accessing pages to harvest links from several pages, and then storing them in a SQLite database. The second script does a deep crawl of product details from the gathered links. The Playwright library allows for asynchronous browser automation. During this phase, pagination is scrolled so that dynamic contents are fetched on the pages using BeautifulSoup. By using this function, the prices, locations, specifications, and mortgage details that are available property information are obtained. The ordered information goes into a database, and faulty URLs are registered for retry scenarios. Best practice for the project includes introducing a pause between requests; the user agent should be rotated between requests, and good error handling is also as robust as possible. This is a good starter for those coming to learn real-life web scraping methods. ## Libraries and Tools Used in Bayut Web Scraping In the Bayut web scraping project, a wide range of libraries and tools available in Python is used to properly extract both static and dynamic content. For the smooth working of data retrieval and processing, each of these libraries has contributed a lot in their own individual ways. The requests library was used mainly to perform HTTP requests when gathering web content. It would make GET requests to any URL that the user may want to input and supports customized headers and cookies. For this project, requests have been used so that it would cycle through the pagination links of Bayut's website where it has used a change in the user-agent string so that it might mask it as if the view process were a human. The bs4 package offers BeautifulSoup-a power parsing tool for HTML and XML. The package allows one to easily navigate and extract the data by converting raw HTML to a structured format. The tool has been applied in the scripts: at first, property links were pulled from the website, and later, more details on the property were parsed by using tag-based selectors. The sqlite3 library is a light database management system with built-in SQL capabilities, which stores and tracks URLs and maintains a status flag to distinguish between successfully processed and failed links. This makes the system have efficient retry mechanisms and prevent redundant processing of already scraped URLs. Playwright is an advanced library for browser automation, which supports dynamic content rendering. Unlike static HTML libraries, it supports JavaScript-driven pages, allowing scrolling and clicking on buttons. In the second script, the content is dynamically loaded by Playwright to ensure that all the details of properties are visible before the data is extracted. asyncio enables tasks to be executed parallelly while concurrently handling calls. Asynchronous functions, declared by async def, make scraping even more efficient because asynchronous functions are able to run many navigation procedures concurrently, which in itself significantly reduces the overall runtime for any collection procedure. Time and random libraries create delays and randomness between requests like a human browse. In the code example, time.sleep() causes pauses during execution, and random.uniform() causes variable delays so that an anti-bot will not flag it. These libraries form a robust framework for scalable and ethical web scraping, combining automation, dynamic content handling, and data storage to optimize the Bayut project's performance and reliability. ## The Role of User Agents in Web Scraping: Mimicking Real Browsing Behavior User agents are strings that provide servers with information about the client responsible for a web request, including browser type, version, and operating system. A website will tailor its content and layout to fit the device of a user through the use of such information. Use of user agents in web scraping is also necessary because it allows mimicking the behavior of real users and escapes any detection from sites' anti-bot measure. In this web scraping project with Bayut, user agents have been used for simulating the requests coming from different browsers and devices. In order to not show repetitive patterns, which many websites flag as bot activity, the scraper will rotate user-agent strings across several requests. That improves access stability and prevents the risk of an IP ban in order to ensure continuous data collection. Adding to techniques such as random delays between requests and user-agent rotation, it makes the overall process of scraping more efficient, more reliable, and stealthier, which makes it a crucial tool for effective data harvesting. ## Efficient Progress Tracking and Recovery with SQLite Interruptions may arise from network problems, server blocks, or script errors while scraping in web scraping, and hence, data will be left incomplete. This project handles such scenarios efficiently by using SQLite as a lightweight database to track the progress and allow resumption from the point of failure. Every URL is stored with a status flag set initially to 0, which means that the data extraction is pending. Whenever the scraper successfully retrieves a URL, it updates its status to 1\. The scraper can, however, query the database to find only the URLs whose status is 0 when it stops without any prior notice so that it can proceed without duplicating work or losing the data previously acquired. SQLite is particularly beneficial for web scraping, as it is easy, efficient, and already there with Python without needing the user to have extra configuration on the server. SQLite's compact file-based storage is suitable for medium-scale data with minimal overhead. Using SQLite allows the pipeline of this scraper to be persisted across runtime, transactions safe, and easily query in general which makes it robust, scalable, and well-suited for dynamic large-scale extraction of web data. ## STEP 1 : Product URL Scraping From BAYUT ### Libraries Overview for Bayut Web Scraping ``` import requests from bs4 import BeautifulSoup import sqlite3 import random import time ``` This section incorporates libraries such as requests, BeautifulSoup, sqlite3, Playwright, asyncio, random, and time to handle HTTP requests, parse HTML, manage data storage, automate browser actions, enable asynchronous scraping, and introduce delays to make the web scraping for Bayut property data efficient and human-like. ### Defining Constants for Web Scraping Configuration ``` # Constants BASE_URL = "https://www.bayut.com" INITIAL_URL = f"{BASE_URL}/for-sale/apartments/dubai/?completion_status=ready" USER_AGENTS_FILE = "/home/user/Documents/Datahut_Internship/bayut/data/user_agents.txt" DB_FILE = "bayut_webscraping.db" HEADERS_COOKIE = 'anonymous_session_id=698fd37f-ac65-45f6-957f-fd1f0fb0e4b2; device_id=m5c20a111ecppbt7f' ``` This section defines all the constant key values needed to perform web scraping. BASE\_URL is the Bayut website main address, used as a basis to form other addresses. The INITIAL\_URL parameter defines the page from which scraping of ready-to-sell apartments in Dubai shall start. The USER\_AGENTS\_FILE parameter indicates the text file containing different user-agent strings mimicking real-user behavior. DB\_FILE is a parameter that indicates the name of the SQLite database file where the scraped data and URLs will be stored. HEADERS\_COOKIE still carries session-related cookies to communicate with the website, thus sustaining a connection during scraping. It centralizes the configuration settings, making the code easier to manage and modify. ### Loading User Agents from a File ``` # 1. Load user agents from file def load_user_agents(file_path): """ Load user agents from a specified file. Args: file_path (str): The path to the file containing a list of user agent strings, with one agent per line. Returns: list: A list of user agent strings loaded from the file. Each string is stripped of leading and trailing whitespace. """ with open(file_path, "r") as file: return [line.strip() for line in file.readlines()] ``` This function, load\_user\_agents, reads out the list of user-agent strings by reading from a specified file to simulate different web browsers while scraping. The function takes in file\_path as an argument. It's a location of the file that contains the list of user agents with one user-agent string per line. It opens the file to read all lines, removes extra spaces using strip(), and returns a list of user agent strings. These user agents help make requests appear as if sent from actual people, so, therefore, quite unlikely to have the website block their requests. ### Selecting a Random User Agent ``` # 2. Get a random user agent def get_random_user_agent(user_agents): """ Select a random user agent from a provided list. Args: user_agents (list): A list of user agent strings. Returns: str: A randomly selected user agent string from the list. """ return random.choice(user_agents) ``` The get\_random\_user\_agent function helps pick a random user-agent string from a provided list of user agents. It takes user\_agents as an argument, which is a list of user-agent strings loaded earlier. The function uses random.choice() to randomly select one user agent from the list and returns it. This randomness makes each web request look like it's coming from a different browser, helping avoid detection and blocking by websites. ### Initializing the Database and Creating a Table ``` # 3. Initialize database and create table def initialize_database(db_file): """ Initialize an SQLite database and create a table for storing product links if it does not already exist. Args: db_file (str): The file path of the SQLite database. Returns: tuple: A tuple containing: - conn (sqlite3.Connection): The SQLite connection object. - cursor (sqlite3.Cursor): The SQLite cursor object for executing database operations. """ conn = sqlite3.connect(db_file) cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS product_links ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT UNIQUE, status INTEGER DEFAULT 0 ) """) conn.commit() return conn, cursor ``` The initialize\_database function initializes an SQLite database in which the links to products are saved for web scraping. This function receives the db\_file as argument, being the name or path of the file in which the database will be saved. The function will establish a connection with the database and a cursor to perform SQL commands. The first operation will be checking whether the table product\_links already exists, in which case it creates it if it does not exist. The table product\_links has id, link and status as its three columns where id is unique, link points to the products and status tells whether the link has been processed or not set to 0 by default to efficiently store track and retry extraction of failed links. It returns the connection and cursor to use in any further database operation. ### Fetching and Parsing a Webpage ``` # 4. Fetch and parse a webpage def fetch_page(url, user_agents): """ Fetch a webpage and parse its content using BeautifulSoup. This function sends an HTTP GET request to the specified URL using randomized headers for the 'User-Agent' field to mimic browser behavior. It also includes a predefined cookie header to bypass potential session-based restrictions on the server. If the request is successful (status code 200), the HTML content of the page is parsed using BeautifulSoup and returned. Otherwise, it returns None. Args: url (str): The URL of the webpage to fetch. user_agents (list): A list of user agent strings used to randomize requests for ethical and anti-bot compliance. Returns: BeautifulSoup or None: - If the HTTP request is successful, returns a BeautifulSoup object containing the parsed HTML content of the webpage. - If the HTTP request fails (non-200 status code), returns None. """ headers = { 'User-Agent': get_random_user_agent(user_agents), 'Cookie': HEADERS_COOKIE } print(f"Fetching: {url} with User-Agent: {headers['User-Agent']}") response = requests.get(url, headers=headers) if response.status_code != 200: print(f"Failed to fetch page: {url}") return None return BeautifulSoup(response.text, "html.parser") ``` The function, fetch\_page uses BeautifulSoup to prepare the content fetched from the webpage for data extraction. It requires two arguments: the url to fetch the page and a user agent list to randomize the requests sent; the requests are, thus, more likely to look like those sent by a real browser. This randomness also helps in not getting detected by anti-bot systems. Its next call also uses the same cookie header defined, which it uses to manage sessions. It uses the requests library to send an HTTP GET request. If the request is successful with a status code of 200, it uses BeautifulSoup to parse the HTML content of the page, enabling the extraction of structured data. On failure, it prints an error message and returns None. This process makes the scraping more reliable and less likely to be blocked. ### Extracting Product Links from a Webpage ``` # 5. Extract product links from a page def extract_product_links(soup): """ Extract product links from the parsed HTML content of a webpage. This function searches the parsed HTML (BeautifulSoup object) for all product containers and extracts the individual product links. It specifically looks for `div` elements with the class `"dde89f38"`, where each product link is typically stored in an `` tag. It then constructs the full URL for each product and adds it to a list. Args: soup (BeautifulSoup): A BeautifulSoup object containing the parsed HTML content of the webpage. Returns: list: A list of full URLs pointing to individual product pages. The URLs are constructed by appending the `href` attribute of the `` tag to the base URL (`BASE_URL`). """ product_links = [] product_divs = soup.find_all("div", class_="dde89f38") for div in product_divs: a_tag = div.find("a", href=True) if a_tag: full_link = BASE_URL + a_tag['href'] product_links.append(full_link) return product_links ``` The extract\_product\_links function collects all links of products from the given webpage's parsed HTML content. To locate div elements with a specific class name "dde89f38", in which it stores all links, the script will use a BeautifulSoup object that mimics the structure of the webpage. This structure is used to collect all such tags inside each div element, which contains the actual link. As for each tag found being tag (link) with the title "Next". If such a next link is established, it gets the full base URL by pasting the "BASE\_URL" with the attribute "href". It returns a value for processing in the next page GET request. Returning None indicates "Next" linking is not made, meaning further pages are exhausted and there will be nothing for scraping. ### Saving Unique Product Links to the Database ``` # 7. Save unique links to the database def save_links_to_db(cursor, links): """ Save unique product links to an SQLite database. This function iterates through a list of product links and inserts each link into the `product_links` table of the SQLite database, ensuring that duplicate links are ignored. If an error occurs during the insertion process, it logs the error without stopping the process. Args: cursor (sqlite3.Cursor): The SQLite cursor object used to execute SQL queries on the database. links (list): A list of product URLs (strings) to be inserted into the database. Returns: None: This function does not return any value. It directly modifies the database by inserting the links. """ for link in links: try: cursor.execute(""" INSERT OR IGNORE INTO product_links (link, status) VALUES (?, 0) """, (link,)) except Exception as e: print(f"Failed to insert link {link}: {e}") ``` The function saves\_links\_to\_db saves the list of product links scraped from a webpage into an SQLite database. It accepts two parameters: the cursor, which is used to interact with the database, and links, which is a list of product URLs. Inside this function, it iterates over each link in the list and tries to insert it into the product\_links table of the database. This INSERT OR IGNORE SQL command will prevent the same link from getting inserted in duplicate. This is because, once the same link is present in the database, it won't be allowed to get inserted. If a problem occurs in inserting a link due to some error in the database, it is caught and printed, but the process doesn't stop here. In this sense, scraping is not prone to problems and infrequent errors are without any influence over the process. ### Main Function to Control the Workflow: scrape\_bayut() ``` # 8. Main function to control the workflow def scrape_bayut(): """ Main function to control the entire web scraping workflow for Bayut. This function orchestrates the web scraping process by loading user agents, initializing the database, fetching pages, extracting product links, saving them to the database, and navigating through pagination. It continuously scrapes until there are no more pages to process, handling each page's data with a delay to avoid overwhelming the server. It performs the following steps: 1. Loads a list of user agents from a specified file. 2. Initializes the SQLite database and sets up the `product_links` table. 3. Begins scraping from the initial URL. 4. Iterates over each page, fetching its content and extracting product links. 5. Removes duplicate links and saves them to the database. 6. Identifies the next page to scrape, and repeats the process until no further pages are found. 7. Introduces a random delay between requests to mimic human behavior and prevent getting blocked. 8. Closes the database connection after completing the scraping process. Args: None: This function does not take any arguments directly. Returns: None: This function does not return any value. It performs the scraping and stores the results in the database. """ # Load user agents user_agents = load_user_agents(USER_AGENTS_FILE) # Initialize database conn, cursor = initialize_database(DB_FILE) # Scraping process url = INITIAL_URL while url: soup = fetch_page(url, user_agents) if not soup: break # Extract product links page_links = extract_product_links(soup) print(f"Found {len(page_links)} product links on this page.") # Remove duplicates and save to database unique_links = list(set(page_links)) save_links_to_db(cursor, unique_links) conn.commit() print(f"Saved {len(unique_links)} unique product links to the database.") # Find the next page url = find_next_page(soup) # Add a delay between page visits if url: delay = random.uniform(2,5) print(f"Delaying for {delay:.2f} seconds before visiting the next page...") time.sleep(delay) conn.close() print(f"Scraping completed. Links saved to the database {DB_FILE} in table 'product_links'.") ``` The main function of scraping is scrape\_bayut(), which is the heart of the process that coordinates the whole flow of scraping product links from the Bayut website. It coordinates all the steps involved in the scraping process such as loading user agents initializing the database, extracting links from pages, saving them to the database, and navigating through multiple pages. This function serves as a controller which ensures all tasks are performed sequentially. The function starts with the loading of the list of user agents from a file. Then, it connects to an SQLite database and creates a table for product links. It starts scraping from a starting URL by making a request to fetch the page content through the function. After fetching the page, it proceeds with extracting the product links from the page. Links are cleaned for duplicates, and each unique link is saved in the database. It also takes care of pagination by checking if there is a "next" page to scrape. It then moves on to the next page and continues the process until there are no more pages left to scrape. To avoid overwhelming the server, the function introduces a random delay between requests, simulating human-like behavior while scraping. Once done with the scraping, the function is able to close the database connection and print out the message which says the links are saved. No value is returned by the function, but it changed the database to hold the links that were scraped. ### Entry Point to Execute the Script ``` # Execute the script if __name__ == "__main__": """ Entry point to execute the Bayut web scraping script. This block checks if the script is being executed as the main module. If it is, it calls the `scrape_bayut` function to start the web scraping process. This ensures that the scraping process is initiated only when the script is run directly, and not when it is imported as a module into another script. The script initiates the scraping of product links from the Bayut website and saves them into an SQLite database. The scraping process includes fetching pages, extracting product links, saving the links to the database, and navigating through paginated pages. Usage: - The script can be run directly from the command line, after ensuring all dependencies and configurations (e.g., user agents, database) are in place. - It will automatically start the scraping process and store the results in the database specified by `DB_FILE`. """ scrape_bayut() ``` The entry point of running the script of the web scraper is the block if \_\_name\_\_ == " \_\_main\_\_ ":. This is the block where it checks if the script is run by the end-user or if it is imported into another scrip. The function here calls the scrape\_bayut() function, which initiates the whole process of scraping. That page fetches pages from Bayut, cuts the product links and saves in SQLite DB, and handles pagination to scrape all pages available. ## STEP 2: Extracting Detailed Product Information from Individual Pages ``` import asyncio import sqlite3 import random import time from playwright.async_api import async_playwright from bs4 import BeautifulSoup ``` ### Generating a Random User-Agent for Web Scraping ``` # Function to get a random user-agent from the file def get_random_user_agent(): """ Reads a list of user-agents from a text file and returns a randomly selected user-agent string. Returns: str: A user-agent string randomly chosen from the 'data/user_agents.txt' file. """ with open('data/user_agents.txt', 'r') as file: user_agents = file.readlines() return random.choice(user_agents).strip() ``` In this section, we define the helper function get\_random\_user\_agent. This function reads user-agent strings from a text file and returns one at random. The function reads all user-agents stored in the data/user\_agents.txt file and removes any extraneous spaces or newline characters from them, before choosing one of them at random using the random.choice function. This way, each request used during the scrape will have a different user-agent, thus promoting anonymity and chances of getting blocked are lower. This dynamic way is essential to responsible and efficient web scraping. ### Establishing a Database Connection ``` # Function to connect to the SQLite database def get_db_connection(): """ Establishes a connection to the SQLite database. Returns: sqlite3.Connection: A connection object to the 'bayut_webscraping.db' database. """ conn = sqlite3.connect('bayut_webscraping.db') return conn ``` This section demonstrates the get\_db\_connection function, which connects to a SQLite- database for writing scraped data. . The function opens the bayut\_webscraping.db file and then returns the connection object used to interact with the database.This connection forms the entry point for the SQL commands used to execute insertions into data or tables. A dedicated function for database connections makes the code more modular and readable and easy to handle all the database interactions throughout the entire scraping process. ### Creating Database Tables for Web Scraping ``` def create_tables(): """ Creates necessary tables in the SQLite database for storing scraped data and failed URLs. Tables: - bayut_product_data: Stores property details such as product URL, price, location, specifications, benefits, description, property info, features, amenities, and mortgage details. Columns: - id (INTEGER): Auto-incremented primary key. - product_url (TEXT): URL of the property. - price (TEXT): Price of the property. - location (TEXT): Location of the property. - specifications (TEXT): Property specifications. - benefits (TEXT): Benefits of the property. - description (TEXT): Property description. - property_info (TEXT): Key property information. - features_and_amenities (TEXT): Features and amenities. - mortgage_details (TEXT): Mortgage calculation details. - FOREIGN KEY (product_url): References 'link' from product_links table. - failed_urls: Stores URLs that failed during scraping. Columns: - id (INTEGER): Auto-incremented primary key. - link (TEXT): URL that failed to scrape. - status (INTEGER): Scraping status (e.g., 0 for failure). - reason (TEXT): Reason for failure. Notes: - Call this function once before starting the scraping process to ensure tables are created. """ conn = get_db_connection() cursor = conn.cursor() # Create bayut_product_data table cursor.execute(""" CREATE TABLE IF NOT EXISTS bayut_product_data ( id INTEGER PRIMARY KEY AUTOINCREMENT, product_url TEXT, price TEXT, location TEXT, specifications TEXT, benefits TEXT, description TEXT, property_info TEXT, features_and_amenities TEXT, mortgage_details TEXT, FOREIGN KEY (product_url) REFERENCES product_links(link) ) """) # Create failed_urls table cursor.execute(""" CREATE TABLE IF NOT EXISTS failed_urls ( id INTEGER PRIMARY KEY AUTOINCREMENT, link TEXT, status INTEGER, reason TEXT ) """) conn.commit() conn.close() # Call this function before running the scraping process create_tables() ``` The create\_tables function initializes the required tables in the SQLite database for storing scraped data and logging failed URLs. It defines two tables: bayut\_product\_data and failed\_urls. The bayut\_product\_data table includes columns for storing property details, such as product\_url, price, location, specifications, benefits, description, property\_info, features\_and\_amenities, and mortgage\_details. Each row has a unique identifier (id), and a foreign key constraint on product\_url references the link column in the product\_links table to maintain relational integrity. The failed\_urls table captures URLs that failed to scrape, with columns for link, status (an integer representing scrape success or failure, where 0 indicates failure), and reason (describing the cause of failure). The use of CREATE TABLE IF NOT EXISTS ensures that re-running the function does not cause errors if the tables already exist. The function commits changes to save the schema and closes the connection to free resources. It should be called once before the scraping process to ensure the database structure is in place. ### Updating URL Status in the product\_links Table ``` # Function to update the status of a URL in the product_links table def update_url_status(url, status): """ Updates the scraping status of a URL in the product_links table. Args: url (str): The URL of the product link to update. status (int): The status value to set (e.g., 1 for scraped, 0 for pending). Notes: - Assumes the 'product_links' table has a 'status' column and a 'link' column. - Call this function to mark URLs as scraped or pending during the scraping process. """ conn = get_db_connection() cursor = conn.cursor() cursor.execute( "UPDATE product_links SET status = ? WHERE link = ?", (status, url) ) conn.commit() conn.close() ``` The update\_url\_status function will alter the status of a particular URL in the product\_links table as per its scraping status. It is taking two input arguments: url, which is a string for the display of the product link, and status, an integer that shows 1 if the URL has successfully been scraped and 0 for a pending scrape. The function establishes a connection to the SQLite database. It updates the status column of the corresponding link in the product\_links table. Then it commits the update and closes the connection to be conservative with its use of resources. This function assumes the product\_links table has status and link columns. This function can be used in marking URLs as processed or pending in web scraping. ### Inserting Scraped Data into bayut\_product\_data Table ``` # Function to insert scraped data into the bayut_product_data table def insert_scraped_data(data): """ Inserts the scraped property data into the bayut_product_data table. Args: data (dict): A dictionary containing the following keys: - product_url (str): URL of the property. - price (str): Price of the property. - location (str): Location of the property. - specifications (str): Property specifications. - benefits (str): Benefits of the property. - description (str): Property description. - property_info (dict): Key property information. - features_and_amenities (dict): Features and amenities. - mortgage_details (dict): Mortgage calculation details. Notes: - Converts 'property_info', 'features_and_amenities', and 'mortgage_details' dictionaries to strings before storing in the database. - Assumes 'bayut_product_data' table structure matches the data being inserted. """ conn = get_db_connection() cursor = conn.cursor() cursor.execute( """ INSERT INTO bayut_product_data ( product_url, price, location, specifications, benefits, description, property_info, features_and_amenities, mortgage_details ) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?) """, ( data['product_url'], data['price'], data['location'], data['specifications'], data['benefits'], data['description'], str(data['property_info']), str(data['features_and_amenities']), str(data['mortgage_details']) ) ) conn.commit() conn.close() ``` This function, insert\_scraped\_data, saves retrieved property information in the SQLite database's bayut\_product\_data table. It receives a dictionary as an argument; that dictionary is named data, whose keys are all the product URL, price, location, specifications, benefits, description, property information, features and amenities, and mortgage details. The inner dictionaries are flattened into string values for property\_info, features\_and\_amenities, and mortgage\_details because only those types can be used within the database table. It populates the columns by a safe parameterized query, closes the database connection after efficiently committing the transaction so that resources will not get locked. It is sure to save the data which has been scraped correctly so that it could be used or further analyzed in future. ### Inserting Failed URLs into failed\_urls Table ``` # Function to insert failed URLs into the failed_urls table with a status def insert_failed_url(url, reason, status): """ Inserts a failed URL into the failed_urls table with the specified status. Args: url (str): The URL that failed to scrape. reason (str): The reason for the failure (e.g., timeout, connection error). status (int): default 0 for URLs that need to be re-scraped. Notes: - The status column is used to track the state of URLs: - 0 indicates that the URL failed and needs to be re-scraped. - This function helps track errors and manage re-scraping of failed URLs. - The default value of 0 is inserted to indicate the URL needs to be retried. """ conn = get_db_connection() cursor = conn.cursor() cursor.execute( """ INSERT INTO failed_urls (link, status, reason) VALUES (?, ?, ?) """, (url, status, reason) ) conn.commit() conn.close() ``` The insert\_failed\_url function adds a failed URL during the scraping process into the failed\_urls table, which includes a reason for the failure and a status. The url parameter is the failed link, reason describes why the failure happened (for example, a timeout or connection error), and status is set to 0 by default, indicating that the URL needs to be retried. This function keeps track of scraping errors and manages attempts at re-scraping in an efficient manner. It commits the data to the database and closes the connection after inserting the record. ### Scrolling a Web Page to Load Dynamic Content ``` # Scraping functions for various data async def scroll_page(page, direction="down", delay=100): """ Scrolls a web page in the specified direction to load dynamic content. Args: page (playwright.async_api.Page): The Playwright page object. direction (str): Direction to scroll, either "down" or "up". Defaults to "down". delay (int): Delay between scroll steps in milliseconds. Defaults to 100 ms. Notes: - Uses JavaScript to calculate and scroll to different positions on the page. - Simulates smooth scrolling by pausing between steps. - Helps in loading content dynamically rendered during scrolling. """ scroll_height = await page.evaluate("document.body.scrollHeight") step = 100 if direction == "down": for position in range(0, scroll_height, step): await page.evaluate( f"window.scrollTo(0, {position})" ) await asyncio.sleep(delay / 1000) elif direction == "up": for position in range(scroll_height, 0, -step): await page.evaluate( f"window.scrollTo(0, {position})" ) await asyncio.sleep(delay / 1000) ``` The scroll\_page function scrolls up or down the web page in order to assist in loading dynamic content that shows up as you scroll through the page. This function uses a page object from the Playwright library to run the JavaScript needed for scrolling. The direction parameter allows you to select either "down" or "up" scrolling, which is the default, and delay sets the pause between each scroll step, defaulting to 100 milliseconds. This is a good scrolling technique when the data loads on scrolling. Thus, it can be useful in web scraping as it loads all content before scraping. ### Extracting Price from Web Page Content ``` async def extract_price(soup): """ Extracts the price from the parsed HTML content. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: The extracted price if found, otherwise "Price not found". """ price_div = soup.find( "div", class_="_61c347da" ) return ( price_div.find("span", class_="_2d107f6e").text if price_div else "Price not found" ) ``` The extract\_price function fetches the price of a product from a webpage using a BeautifulSoup object parsing the content on its HTML page. It looks for the appearance of a certain
    element bearing the \_61c347da class name inside which it finds a element with the class \_2d107f6e in order to get the text with price details. If the price is found it returns price as a string, else it returns "Price not found". This function simplifies the extraction of price information from structured web data. ### Extracting Property Location from HTML Content ``` async def extract_location(soup): """ Extracts the location from the parsed HTML content. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: The extracted location if found, otherwise "Location not found". """ location_div = soup.find( "div", class_="e4fd45f0" ) return ( location_div.text if location_div else "Location not found" ) ``` This extract\_location function finds and returns the location information from a webpage using the BeautifulSoup object, representing the parsed HTML. It searches for a
    element with the class name e4fd45f0\. If it finds the location, the function returns the text of the location; otherwise, it returns "Location not found." This function simplifies the process of getting property location details from page content. ### Extracting Property Specifications from HTML Content ``` async def extract_specifications(soup): """ Extracts the specifications from the parsed HTML content. Searches for a div with class "_14f36d85" to retrieve the specifications text. Uses `.stripped_strings` to extract and join the strings without extra whitespace. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: A comma-separated string of specifications if found, otherwise "Specifications not found". """ spec_div = soup.find( "div", class_="_14f36d85" ) return ( ", ".join(spec_div.stripped_strings) if spec_div else "Specifications not found" ) ``` The extract\_specifications function retrieves the property specifications from a webpage using a BeautifulSoup object that represents the parsed HTML. It searches for a
    element with the class name 14f36d85\. If found, it uses .strippedstrings to gather all text content without extra whitespace and returns it as a comma-separated string. If the specifications are not available, it returns "Specifications not found." This function helps in obtaining structured specifications information efficiently from the page content. ### Extracting Property Benefits from HTML Content ``` async def extract_benefits(soup): """ Extracts the benefits from the parsed HTML content. Searches for a div with the class "_34032b68 _656393c5 _701d0fe0" and then looks for an h1 element with the class "d8b96890 fontCompensation". If the benefits section is found, it returns the text content. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: The extracted benefits if found, otherwise "Benefits not found". """ overview_div = soup.find( "div", class_="_34032b68 _656393c5 _701d0fe0" ) benefits_h1 = ( overview_div.find( "h1", class_="d8b96890 fontCompensation" ) if overview_div else None ) return ( benefits_h1.text.strip() if benefits_h1 else "Benefits not found" ) ``` The extract\_benefits function extracts the benefits section from a webpage using a BeautifulSoup object that represents the parsed HTML content. It first looks for a
    element with the class \_34032b68 \_656393c5 \_701d0fe0, which contains the relevant information. Inside this div, it searches for an

    element with the class d8b96890 fontCompensation. If both elements are found, it extracts and returns the text content after stripping extra spaces. If not, the function returns "Benefits not found." ### Extracting Benefits from a Webpage ``` async def extract_benefits(soup): """ Extracts the benefits from the parsed HTML content. Searches for a div with the class "_34032b68 _656393c5 _701d0fe0" and then looks for an h1 element with the class "d8b96890 fontCompensation". If the benefits section is found, it returns the text content. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: The extracted benefits if found, otherwise "Benefits not found". """ overview_div = soup.find( "div", class_="_34032b68 _656393c5 _701d0fe0" ) benefits_h1 = ( overview_div.find( "h1", class_="d8b96890 fontCompensation" ) if overview_div else None ) return ( benefits_h1.text.strip() if benefits_h1 else "Benefits not found" ) ``` The extract\_benefits function retrieves property benefits from a webpage's HTML using BeautifulSoup. It searches for a
    with the class 34032b68 656393c5 \_701d0fe0 and then looks for an

    inside it with the class d8b96890 fontCompensation. If these elements are found, the text content of the

    is extracted and returned after removing extra spaces. If not, the function returns "Benefits not found." ### Extracting Description from a Webpage ``` async def extract_description(soup): """ Extracts the description from the parsed HTML content. Searches for a span element with the class "_3547dac9" to extract the description text. Strips any extra whitespace from the description before returning. Args: soup (BeautifulSoup): The BeautifulSoup object containing the HTML content of the page. Returns: str: The extracted description if found, otherwise "Description not found". """ description_span = soup.find( "span", class_="_3547dac9" ) return ( description_span.get_text(strip=True) if description_span else "Description not found" ) ``` The extract\_description function retrieves the property description from a webpage's HTML content using BeautifulSoup. It searches for a element with the class 3547dac9 and extracts its text content. Any extra spaces are removed using the gettext(strip=True) method. If the element is found, the description is returned; otherwise, the function returns "Description not found." ### Extracting Property Information from a Webpage ``` async def extract_property_info(soup): """ Extracts property information from the parsed HTML content. Searches for a