Excelerate × Rochester Institute of Technology

AI-Powered Data Analysis Internship

Star Performer work across data understanding, cleaning, visual analysis, predictive modeling, and recommendations for RIT student engagement.

PythonRandom ForestEDAData CleaningMachine Learning

360° Evaluation

Three perspectives · averaged
84%
Self
100%
Peer
100%
Managerial
95%
Overall average
View 360° Evaluation Report (PDF)

8,558

Student records analyzed

8,024

Records after cleaning

86.1%

Model accuracy

95%

Overall internship score

Problem & Context

Rochester Institute of Technology (RIT) and Excelerate wanted to understand what drives students to participate in opportunities after signing up. With thousands of records and many potential variables, the challenge was to move beyond guesswork and identify the real levers of engagement.

My task was to clean the dataset, explore patterns, build a predictive model, and present findings that could inform program design and outreach strategy.

Dataset & Methods

  • Dataset: 8,558 RIT student engagement records with demographic, behavioral, and opportunity-type fields.
  • Cleaning: Removed 534 records with missing learner or opportunity identifiers, removed four non-analytical or misleading columns, and engineered 14 fields. The model-ready dataset had 8,024 records across 26 columns with zero missing values.
  • Exploration: Cohort segmentation and comparative analysis across opportunity categories, signup timing, and demographic groups.
  • Modeling: Built and evaluated a Random Forest classifier to predict participation outcomes and rank feature importance.

Analysis workflow & evidence

Follow the work from the first data-quality review through cleaning, charts, modeling, and the final recommendation.

  1. 01 · Understand

    Dataset structure and data quality

    The initial review documented the dataset, variables, missing entries, and date-format issues before analysis began.

    Open Week 1 data-understanding report
  2. 02 · Clean and explore

    Cleaning, feature engineering and visual insights

    The Week 2 report records the transition from 8,558 raw rows to 8,024 records across 26 columns, with zero missing values, plus the exploratory analysis and charts.

    Open Week 2 visual-insights report
  3. 03 · Model

    Prediction and model evaluation

    The Random Forest report includes evaluation metrics, feature-importance visuals, a confusion matrix, and a predicted-participation heatmap.

    Open Week 3 prediction report
  4. 04 · Present

    Final program insights presentation

    The final presentation summarizes the analysis and recommends a real-time participation-probability dashboard for program managers; it is a proposed next step, not a claim that a live dashboard was deployed.

    Open final presentation

Key Findings

Opportunity Category is the dominant predictor

Random Forest feature importance showed that the type of opportunity — not region, age, or gender — most strongly predicted whether a learner would participate.

Events outperform internships by a wide margin

Events held a 96% active-participation rate, while internships showed a 65% rejection rate, suggesting very different learner intent and friction points.

Engagement timing predicts retention

Learners who applied 1–3 months after signup outperformed same-day applicants by 20 percentage points, indicating that rushed onboarding may lower commitment.

Recognition

This internship earned Star Performer award for the Excelerate × RIT AI-Powered Data Analysis internship, with a 95% overall score, including 84% self-evaluation, 100% peer evaluation, and 100% managerial evaluation.

Let's Talk

Need data-driven operations support?

I turn messy data into clear decisions — and clear decisions into systems that save time.