Skip to content

Repository files navigation

CodeAlpha — Data Analytics Internship

Submission covering Task 1 (Web Scraping), Task 2 (Exploratory Data Analysis), Task 3 (Data Visualization) and Task 4 (Sentiment Analysis). The requirement is a minimum of two to three tasks — all four are included here, so the submission is comfortably above the completion criteria.

Tasks 1–3 form one continuous pipeline on a dataset I scraped myself; Task 4 is a standalone NLP project on 20,000 real Amazon reviews.


Repository structure

CodeAlpha_DataAnalytics/
├── data/
│   ├── github_repos_raw.csv          # 600 rows exactly as scraped
│   ├── github_repos_clean.csv        # 479 unique repos, cleaned + engineered
│   ├── amazon_reviews_raw.csv        # 20,000 labelled Amazon reviews
│   └── amazon_reviews_scored.csv     # reviews + VADER sentiment scores
├── task1_web_scraping/
│   └── scrape_github_topics.py
├── task2_eda/
│   ├── eda_github_repos.py
│   ├── EDA_FINDINGS.md               # full written analysis
│   └── outputs/                      # 3 figures
├── task3_data_visualization/
│   ├── build_dashboard.py
│   └── outputs/                      # 4 story charts + 1 executive dashboard
├── task4_sentiment_analysis/
│   ├── sentiment_analysis.py
│   ├── SENTIMENT_FINDINGS.md
│   └── outputs/                      # 3 figures incl. word clouds
├── requirements.txt
└── README.md

How to run

pip install -r requirements.txt

python task1_web_scraping/scrape_github_topics.py     # scrapes live data (~2 min)
python task2_eda/eda_github_repos.py
python task3_data_visualization/build_dashboard.py
python task4_sentiment_analysis/sentiment_analysis.py

The CSVs are committed, so Tasks 2–4 run instantly without re-scraping.


Task 1 — Web Scraping

Source: https://github.com/topics/<topic> — 15 topics × 2 pages, public pages, no API key. Tools: requests + BeautifulSoup.

  • Parses each <article> repository card for owner, repo name, description, stars, primary language, last-updated timestamp and topic tags.
  • Converts GitHub's abbreviated star strings (199k, 1.2k) into real integers.
  • Polite scraping: a randomised 1–2 second delay between requests and a descriptive User-Agent.
  • De-duplicates repositories that appear under multiple topics.
  • Engineers new features for downstream analysis: days_since_update, n_tags, desc_word_count, star_bucket.

Result: 600 raw rows → 479 unique repositories × 16 columns, a custom dataset built for this analysis.

Task 2 — Exploratory Data Analysis

Follows the brief exactly: questions first, then structure, then patterns, then hypothesis tests, then data-quality flags. Full write-up in task2_eda/EDA_FINDINGS.md.

Findings from the real data:

Finding Evidence
Popularity follows a power law Skewness +3.65; median 37.8k stars vs mean 54.1k; the top 10% of repos hold 31.8% of all stars
Language is linked to popularity Kruskal-Wallis H = 18.0, p = 0.0012 → reject H₀ across the 5 biggest languages
Maintenance matters Mann-Whitney p < 0.001; repos updated within 90 days: median 41,000 stars vs 31,800 for stale ones
Presentation does not buy stars Spearman ρ = +0.04 for description length and tag count, both non-significant (p > 0.39)
Recency correlates with stars Spearman ρ = −0.218, p < 0.001 (fewer days since update → more stars)

Because stars are heavily skewed, every test used is non-parametric (Spearman, Kruskal-Wallis, Mann-Whitney) and every plot uses a log scale — the assumptions of parametric tests would have been violated.

Data-quality issues flagged: 121 duplicate rows removed, 10.2% missing language values (informative, not random — mostly curated "awesome-list" repos), star counts rounded to 3 significant figures by GitHub's UI, and a clear selection bias toward popular repositories.

Task 3 — Data Visualization

Five deliverables in task3_data_visualization/outputs/, each carrying its insight in the title and caption rather than leaving the reader to decode the axes:

  • DASHBOARD_open_source_landscape.png — six-panel executive dashboard with a KPI strip
  • story_01_language_leaderboard.png — volume vs impact: Python appears in the most projects, but JavaScript/TypeScript projects earn a higher median star count
  • story_02_topic_popularity.png — median popularity by domain
  • story_03_maintenance_vs_popularity.png — the maintenance effect, bar + violin
  • story_04_power_law.png — log-log rank plot and a cumulative-share curve showing how few projects absorb most attention

Task 4 — Sentiment Analysis

Dataset: 20,000 real Amazon product reviews (76% positive / 24% negative). Two methods compared head to head against ground-truth labels:

Method Type Accuracy Notes
VADER lexicon Rule-based, no training 82.7% Works with zero labelled data; over-predicts positive on mixed reviews
TF-IDF + Logistic Regression Supervised, 20k features, uni+bigrams 89.9% (ROC-AUC 0.952) Trained on 15,995 reviews, tested on 3,999 unseen

The supervised model wins by 7.2 points, but the lexicon is the right tool when no labels exist — e.g. a live social-media stream.

Emotion drivers the model learned:

  • Positive → love, great, awesome, easy, best, works, perfect, addictive
  • Negative → uninstalled, waste, deleted, boring, useless, worst, crashes, freezes

Business read-out: ~24% of reviews are negative, so review triage can be automated at ~90% accuracy — negatives routed to support in minutes instead of being read by hand. Tracking the weekly share of negative sentiment gives a leading indicator of churn well before star ratings or refunds move.

Limitations stated honestly: sarcasm and mixed reviews ("great idea, terrible execution") defeat both methods; the class imbalance was handled with class_weight='balanced'; and reviewers are self-selected, so the sample over-represents very happy and very angry customers.


Tech stack

Python 3.12 · requests · BeautifulSoup4 · pandas · NumPy · SciPy · Matplotlib · Seaborn · scikit-learn · vaderSentiment · WordCloud

Author: Aditya Bansal Student ID CA/DF1/237488 — Data Analytics Intern, CodeAlpha

About

Data Analytics internship at CodeAlpha — custom web-scraped GitHub dataset, statistical EDA, visual dashboard, and sentiment analysis on 20,000 Amazon reviews.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages