Summary
The goal of this project is to use Kaggle data to examine religious demography globally. My research on how religious populations have changed over time, as well as regional distributions and trends. The objective is to provide an interactive application that offers data-driven study of religious shifts around the world and visualizes discoveries.
Data Sources
I utilized the following publicly available datasets from Kaggle:
Provides future projections of global religious affiliations.
Includes estimated population changes by religion over time.
Contains worldwide religious distribution data by country.
Provides demographic details for major world religions.
World Religions Across Regions
Includes historical and current religious population data.
Covers multiple regions and religious groups.
This project incorporates techniques and best practices from Professors Cruz-Castro and Grant's Codio Data Wrangling Demo to ensure:
- Efficient data cleaning using
cleaning_utils.py - Effective exploratory data analysis (EDA) using
descriptive_utils.py
Project Goals
Collect and preprocess data from multiple sources.
Perform Exploratory Data Analysis (EDA) to identify trends and insights.
Develop an interactive tool (dashboard or recommendation engine) to visualize findings.
Utilize Python for data wrangling, visualization, and modeling.
Interactive Dashboard
Geospatial patterns (distribution of religions by country/region).
Time series trends (how religious populations have changed over decades).
Key statistics and visual summaries (pie charts, bar graphs, heatmaps).
Methodology
This project follows the CRISP-DM (Cross Industry Standard Process for Data Mining) framework:
1. Data Collection & Preprocessing
Acquire datasets from Kaggle and ensure licensing compliance.
Clean and preprocess the data (handling missing values, outliers, standardization).
The project utilizes cleaning_utils and descriptive_utils for efficient data processing as used in Data Wrangling Demo from Codio by Professor Castro-Cruz and Professor Grant.
Handles missing values, standardizes column names, and removes duplicates.
from cleaning_utils import clean_column_names, fill_missing_values, remove_duplicates
df = clean_column_names(df)
df = fill_missing_values(df, strategy='median')
df = remove_duplicates(df)Used for exploratory data analysis and feature visualization.
from descriptive_utils import describe_data, plot_distribution
describe_data(df)
plot_distribution(df, 'sales_amount')2. Exploratory Data Analysis (EDA)
Generate descriptive statistics and summary reports.
Create visualizations (histograms, scatter plots, correlation matrices, etc.).
3. Feature Engineering & Modeling
Identify key features influencing religious population trends.
Train predictive models (if applicable) to estimate future religious distributions.
4. Dashboard Development
Build an interactive dashboard or visualization tool using Streamlit/Plotly Dash.
Present key findings and trends through interactive components.
Repository Structure
cap5771sp25-project/
├── README.md
├── Data/
│ ├── rounded_population.csv
│ ├── ThrowbackDataThursday 201912 - Religion.csv
│ ├── WRP global data.csv
│ ├── data_access_info.txt
├── Report/
│ ├── Milestone1.pdf
│── Scripts/
│ ├── Milestone_1.py
│ ├── cleaning_utils.py
│ ├── descriptive_utils.py
│
│── Images/
│ ├── histograms/
│ ├── boxplots/
│ ├── correlation_matrices/
│
│── README.md
│── requirements.txt
Tech Stack
Programming Language: Python
Libraries: Pandas, NumPy, Matplotlib, Seaborn, Scikit-learn, Plotly/Streamlit
Data Storage: SQLite (if necessary for structured queries)
Version Control: GitHub
How to Run the Project
Terminal Command to enter Streamlit dashboard:
streamlit run Scripts/main.py
LLM & AI-Assisted Tools Usage Declaration For the development of this project, GitHub Copilot was used to assist with Exploratory Data Analysis (EDA) and code optimization. It provided automated suggestions for data cleaning, visualization improvements, and performance enhancements in Python scripts.
Additionally, Grammarly was used to refine sentence structure and improve report clarity, ensuring professional readability and coherence.
AI-Assisted Contributions: EDA Enhancements: Automated insights for data trends and visualizations. Code Optimization: Improved efficiency in looping structures, function definitions, and debugging. Report Formatting: Enhanced clarity and flow for better presentation.