Submission covering Task 1 (Web Scraping), Task 2 (Exploratory Data Analysis), Task 3 (Data Visualization) and Task 4 (Sentiment Analysis). The requirement is a minimum of two to three tasks — all four are included here, so the submission is comfortably above the completion criteria.
Tasks 1–3 form one continuous pipeline on a dataset I scraped myself; Task 4 is a standalone NLP project on 20,000 real Amazon reviews.
CodeAlpha_DataAnalytics/
├── data/
│ ├── github_repos_raw.csv # 600 rows exactly as scraped
│ ├── github_repos_clean.csv # 479 unique repos, cleaned + engineered
│ ├── amazon_reviews_raw.csv # 20,000 labelled Amazon reviews
│ └── amazon_reviews_scored.csv # reviews + VADER sentiment scores
├── task1_web_scraping/
│ └── scrape_github_topics.py
├── task2_eda/
│ ├── eda_github_repos.py
│ ├── EDA_FINDINGS.md # full written analysis
│ └── outputs/ # 3 figures
├── task3_data_visualization/
│ ├── build_dashboard.py
│ └── outputs/ # 4 story charts + 1 executive dashboard
├── task4_sentiment_analysis/
│ ├── sentiment_analysis.py
│ ├── SENTIMENT_FINDINGS.md
│ └── outputs/ # 3 figures incl. word clouds
├── requirements.txt
└── README.md
pip install -r requirements.txt
python task1_web_scraping/scrape_github_topics.py # scrapes live data (~2 min)
python task2_eda/eda_github_repos.py
python task3_data_visualization/build_dashboard.py
python task4_sentiment_analysis/sentiment_analysis.pyThe CSVs are committed, so Tasks 2–4 run instantly without re-scraping.
Source: https://github.com/topics/<topic> — 15 topics × 2 pages, public pages, no API key.
Tools: requests + BeautifulSoup.
- Parses each
<article>repository card for owner, repo name, description, stars, primary language, last-updated timestamp and topic tags. - Converts GitHub's abbreviated star strings (
199k,1.2k) into real integers. - Polite scraping: a randomised 1–2 second delay between requests and a descriptive User-Agent.
- De-duplicates repositories that appear under multiple topics.
- Engineers new features for downstream analysis:
days_since_update,n_tags,desc_word_count,star_bucket.
Result: 600 raw rows → 479 unique repositories × 16 columns, a custom dataset built for this analysis.
Follows the brief exactly: questions first, then structure, then patterns, then hypothesis tests, then data-quality flags. Full write-up in task2_eda/EDA_FINDINGS.md.
Findings from the real data:
| Finding | Evidence |
|---|---|
| Popularity follows a power law | Skewness +3.65; median 37.8k stars vs mean 54.1k; the top 10% of repos hold 31.8% of all stars |
| Language is linked to popularity | Kruskal-Wallis H = 18.0, p = 0.0012 → reject H₀ across the 5 biggest languages |
| Maintenance matters | Mann-Whitney p < 0.001; repos updated within 90 days: median 41,000 stars vs 31,800 for stale ones |
| Presentation does not buy stars | Spearman ρ = +0.04 for description length and tag count, both non-significant (p > 0.39) |
| Recency correlates with stars | Spearman ρ = −0.218, p < 0.001 (fewer days since update → more stars) |
Because stars are heavily skewed, every test used is non-parametric (Spearman, Kruskal-Wallis, Mann-Whitney) and every plot uses a log scale — the assumptions of parametric tests would have been violated.
Data-quality issues flagged: 121 duplicate rows removed, 10.2% missing language values (informative, not random — mostly curated "awesome-list" repos), star counts rounded to 3 significant figures by GitHub's UI, and a clear selection bias toward popular repositories.
Five deliverables in task3_data_visualization/outputs/, each carrying its insight in the title and caption rather than leaving the reader to decode the axes:
DASHBOARD_open_source_landscape.png— six-panel executive dashboard with a KPI stripstory_01_language_leaderboard.png— volume vs impact: Python appears in the most projects, but JavaScript/TypeScript projects earn a higher median star countstory_02_topic_popularity.png— median popularity by domainstory_03_maintenance_vs_popularity.png— the maintenance effect, bar + violinstory_04_power_law.png— log-log rank plot and a cumulative-share curve showing how few projects absorb most attention
Dataset: 20,000 real Amazon product reviews (76% positive / 24% negative). Two methods compared head to head against ground-truth labels:
| Method | Type | Accuracy | Notes |
|---|---|---|---|
| VADER lexicon | Rule-based, no training | 82.7% | Works with zero labelled data; over-predicts positive on mixed reviews |
| TF-IDF + Logistic Regression | Supervised, 20k features, uni+bigrams | 89.9% (ROC-AUC 0.952) | Trained on 15,995 reviews, tested on 3,999 unseen |
The supervised model wins by 7.2 points, but the lexicon is the right tool when no labels exist — e.g. a live social-media stream.
Emotion drivers the model learned:
- Positive → love, great, awesome, easy, best, works, perfect, addictive
- Negative → uninstalled, waste, deleted, boring, useless, worst, crashes, freezes
Business read-out: ~24% of reviews are negative, so review triage can be automated at ~90% accuracy — negatives routed to support in minutes instead of being read by hand. Tracking the weekly share of negative sentiment gives a leading indicator of churn well before star ratings or refunds move.
Limitations stated honestly: sarcasm and mixed reviews ("great idea, terrible execution") defeat both methods; the class imbalance was handled with class_weight='balanced'; and reviewers are self-selected, so the sample over-represents very happy and very angry customers.
Python 3.12 · requests · BeautifulSoup4 · pandas · NumPy · SciPy · Matplotlib · Seaborn · scikit-learn · vaderSentiment · WordCloud
Author: Aditya Bansal Student ID CA/DF1/237488 — Data Analytics Intern, CodeAlpha