Urdu Wikipedia Text Categorization System
A comprehensive Python-based system for extracting, processing, and categorizing Urdu Wikipedia articles into balanced datasets with intelligent content filtering. This project provides multiple proce
About the Project
A comprehensive Python-based system for extracting, processing, and categorizing Urdu Wikipedia articles into balanced datasets with intelligent content filtering. This project provides multiple processing modes including file generation, Excel output, and Jupyter notebook support for interactive analysis.
Features
Core Features
- Balanced Dataset Creation: Smart algorithm that creates roughly equal numbers of Short, Medium, and Large chunks
- Intelligent Content Filtering: Rejects repetitive content and enforces minimum quality standards
- Multi-mode Processing: Supports disk output, Excel generation, print mode, and Jupyter notebook execution
- Real-time Progress Tracking: Visual feedback with distribution monitoring during processing
- Incremental Saves: Excel and CSV files updated every 10 processed files
Text Processing
- Smart Text Splitting: Intelligent text chunking respecting Urdu punctuation (Û) boundaries
- Size Categorization: Automatically categorizes text into Short (1-50 words), Medium (51-100 words), and Large (101-300 words)
- Content Validation: Minimum 6 words for short content, repetitive pattern detection
- Duplicate Detection: Prevents duplicate URLs and sentences from being processed
Data Management
- Content Categorization: Classifies articles into 5 categories: Islam, History, Science, Personality, and General
- Template-based Output: Generates structured XML output files with metadata
- Excel Export: Creates organized Excel files for data analysis
- Natural Sorting: Proper numerical sorting of input files (article1 before article100)
Quality Assurance
- Repetition Detection: Analyzes word frequency to reject repetitive content (threshold: 60%)
- Minimum Word Count: Enforces at least 6 words for short content
- API Integration: Wikipedia API with proper headers and error handling
- Robust Error Handling: Comprehensive error handling for network requests and data processing
Project Structure
âââ app.py # Main processing script (command-line) with balanced distribution
âââ app.ipynb # Jupyter notebook version with interactive widgets
âââ Sample Data set/ # Input Wikipedia articles (article1.txt, article2.txt, etc.)
âââ txtOutput/ # Generated text files by size and type
â âââ Short/ # Short articles (1-50 words)
â â âââ NC/ # News/Current Affairs
â â âââ LR/ # Literature/Research
â â âââ HR/ # Historical Records
â â âââ NP/ # Notable Personalities
â âââ Medium/ # Medium articles (51-100 words)
â â âââ NC/
â â âââ LR/
â â âââ HR/
â â âââ NP/
â âââ Large/ # Large articles (101-300 words)
â âââ NC/
â âââ LR/
â âââ HR/
â âââ NP/
âââ ExcelOutput/ # Excel files for analysis
â âââ short.xlsx # Short articles data
â âââ medium.xlsx # Medium articles data
â âââ large.xlsx # Large articles data
â âââ combined.xlsx # All articles combined
âââ README.md
âââ requirements.txt
Installation
- Clone or download this repository
- Create a virtual environment (recommended):
python -m venv venv - Activate the virtual environment:
- Windows:
venv\Scripts\activate - macOS/Linux:
source venv/bin/activate
- Windows:
- Install required packages:
pip install -r requirements.txt
Required Dependencies
pandas- Data manipulation and Excel exportrequests- Wikipedia API callsopenpyxl- Excel file generationipywidgets- Jupyter notebook progress bars (for notebook version only)
Usage
Method 1: Command Line Processing (Recommended)
The app.py script provides a unified approach with multiple output modes and balanced distribution:
Command Line Usage
# Generate files to disk python app.py --mode disk # Generate Excel files only python app.py --mode excel # Print results to console python app.py --mode print # Generate both files and Excel (recommended) python app.py --mode both # Allow duplicate URLs (process all documents even with same URL) python app.py --mode disk --allow-duplicate-urls # Run content validation tests python app.py test
Processing Modes
- disk: Creates structured text files in
txtOutput/folder organized by size and subcategory - excel: Generates Excel files in
ExcelOutput/folder for analysis - print: Displays results in console for debugging
- both: Combines disk and excel modes (recommended for complete output)
Features During Processing
- Real-time Progress: Shows current file being processed and completion percentage
- Distribution Tracking: Displays balanced distribution (S:X M:Y L:Z) during processing
- Incremental Saves: Excel files saved every 10 files to prevent data loss
- Graceful Interruption: Press Ctrl+C to stop processing and save current progress
Method 2: Jupyter Notebook (Interactive)
For interactive analysis and experimentation with visual progress bars:
-
Install Jupyter if not already installed:
pip install jupyter ipywidgets -
Launch Jupyter notebook:
jupyter notebook -
Open
app.ipynband run the cells interactively
Notebook Features
- Interactive Progress Bars: Visual progress tracking with ipywidgets
- Real-time Distribution Display: See balanced distribution as processing happens
- Cell-by-cell Execution: Run individual components for testing
- Data Analysis Tools: Built-in summary statistics and sample viewing
- Export Options: Flexible export to Excel with custom filtering
Configuration
Size Categories with Balanced Distribution
- Short: 1-50 words (minimum 6 words, non-repetitive)
- Medium: 51-100 words
- Large: 101-300 words
The system uses an intelligent balancing algorithm that:
- Tracks the count of each size category in real-time
- Prioritizes underrepresented categories to maintain balance
- Adds 30% randomness to prevent overly deterministic patterns
- Aims for roughly equal distribution across all three sizes
Content Validation Rules
Short Content Requirements
- Minimum Words: At least 6 words
- Repetition Threshold: Less than 60% repetitive content
- Unique Word Ratio: Must have sufficient vocabulary diversity
Repetition Detection
The system analyzes word frequency to detect repetitive patterns:
repetition_ratio = 1 - (unique_words / total_words) # Content rejected if repetition_ratio > 0.6 (60%)
Examples of rejected content:
- "ÙØµÙÙ Ú©ÙÙØ¯Ú¯Ø§Ù." (too short, only 2 words)
- "ØªØ§Ø±ÛØ®. ØªØ§Ø±ÛØ®." (too repetitive)
- Single word or phrase repeated multiple times
Content Categories (Keyword-based Classification)
- Islam: Religious content, Quran, Hadith, Islamic practices
- Keywords: Ø§Ø³ÙØ§Ù Ø Ù Ø³ÙÙ Ø ÙØ±Ø¢ÙØ ØØ¯ÛØ«Ø ÙÙ Ø§Ø²Ø Ø±ÙØ²ÛØ Ø²Ú©ÙØ©Ø ØØ¬Ø Ø¯Ø¹Ø§Ø Ø³ÛØ±ØªØ ØµØØ§Ø¨ÛØ ÙÙÛØ ØªÙØ³ÛØ±Ø Ø´Ø±ÛØ¹ØªØ اÙÙÛØ اÛÙ Ø§ÙØ تصÙÙØ عÙÛØ¯ÛØ Ø¬ÛØ§Ø¯Ø Ø³ÙØª
- History: Historical events, civilizations, wars, empires
- Keywords: ØªØ§Ø±ÛØ®Ø ÙØ¯ÛÙ Ø ÙØ±ÙÙ ÙØ³Ø·ÛØ Ø«ÙØ§ÙØªØ Ø¬ÙÚ¯Ø Ø³ÙØ·ÙØªØ Ø¨Ø§Ø¯Ø´Ø§ÛØ Ø§Ø³ØªØ¹Ù Ø§Ø±Ø ÙÛØ§ Ø¯ÙØ±Ø آثار ÙØ¯ÛÙ ÛØ ØªØØ±ÛÚ©Ø Ø§ÙÙÙØ§Ø¨Ø ÙÙØ¬ÛØ ØªÙ Ø¯ÙØ ØªÛØ°ÛØ¨Ø Ø¯ÙØ± ØºÙØ§Ù ÛØ عاÙÙ Û Ø¬ÙÚ¯Ø Ø³ÙØ·Ùت عث٠اÙÛÛØ ÛÙØ±Ù¾Û ÙØ´Ø§Û ثاÙÛÛØ ÙØ¯ÛÙ ÛÙÙØ§Ù
- Science: Physics, chemistry, biology, technology, research
- Keywords: Ø³Ø§Ø¦ÙØ³Ø Ø·Ø¨ÛØ¹ÛØ§ØªØ Ú©ÛÙ ÛØ§Ø ØÛØ§ØªÛØ§ØªØ Ø±ÛØ§Ø¶ÛØ§ØªØ ÙÙÚ©ÛØ§ØªØ جÛÙÛØ§ØªØ Ù¹ÛÚ©ÙØ§ÙÙØ¬ÛØ Ø§ÛÙÙÙØ¬ÛØ Ù ØµÙÙØ¹Û Ø°ÛØ§ÙØªØ Ù ÛکاÙÚ©Ø³Ø Ø±ÙØ¨ÙÙ¹Ú©Ø³Ø Ù¾ÛØªÚ¾ÙÙÙØ¬ÛØ Ø¬ØºØ±Ø§ÙÛÛØ آب Ù ÛÙØ§Ø Ø§Ø±ØªÙØ§Ø¡Ø Ú©ÙØ§ÙÙ¹Ù Ø Ú©Ø§Ø¦ÙØ§ØªØ ØªÙØ§ÙØ§Ø¦ÛØ تØÙÛÙ
- Personality: Biographies, character traits, leadership
- Keywords: Ø³ÙØ§ÙØ Ø¹Ù Ø±ÛØ Ø´Ø®ØµÛØªØ Ø²ÙØ¯Ú¯ÛØ Ø§ÙØ³Ø§ÙÛ Ø®ØµÙØµÛØ§ØªØ Ú©Ø±Ø¯Ø§Ø±Ø ÙØ§Ø¦Ø¯Ø ÙÛÚØ±Ø ترÙÛØ Ø®ÙØ§Ø¨Ø ÛÙØ±Ø Ø¹Ø§Ø¯Ø§ØªØ Ù Ø¹Ø±ÙÙ Ø§ÙØ±Ø§Ø¯Ø رÛÙÙ Ø§Ø ØÙصÙÛ Ø§ÙØ²Ø§Ø¦ÛØ Ø°ÛØ§ÙØªØ Ø°Ø§ØªÛ ØªØ±ÙÛØ اثر Ù Ø±Ø³ÙØ®Ø Ù Ø§ÛØ±Ø Ø°Ù Û Ø¯Ø§Ø±ÛØ ÙØ¸Ø±ÛÛ
- General: Geography, sports, culture, literature, economics
- Keywords: ٠عÙÙÙ Ø§ØªØ Ø±ÛØ§Ø¶ÛØ Ø´ÛØ±Ø Ø¬Ø²ÛØ±ÛØ ØµÙØ¨ÛØ Ú©Ú¾ÛÙØ Ú©Ú¾ÛÙÙÚº Ú©Û Ø§ØµÙÙØ ÙÙ¹Ø¨Ø§ÙØ Ú©Ø±Ú©Ù¹Ø Ø§ÛØ´ÛØ§Ø¦Û Ù Ù Ø§ÙÚ©Ø Ø¯Ø±ÛØ§Ø Ù¾ÛØ§ÚØ Ø¢Ø¨Ø§Ø¯ÛØ زباÙÛÚºØ Ù ÙØ³ÛÙÛØ ÙÙÙÙØ Ú©ØªØ¨Ø Ø§Ø¯Ø¨Ø Ù Ø¹ÛØ´ØªØ تعÙÛÙ Ø Ø«ÙØ§ÙØªØ Ø±ÛØ§Ø¶ÛØ§ØªÛ Ù Ø³Ø§ÙØ§Øª
Subcategories (Random Assignment)
- NC: News/Current Affairs
- LR: Literature/Research
- HR: Historical Records
- NP: Notable Personalities
Output Formats
Text Files (XML Structure)
Generated files follow this template structure:
<features domain="general" document_size="medium" type="HR" url="https://ur.wikipedia.org/wiki?curid=12345"> <source> [Urdu text content here] </source> </features>
Excel Files
Excel output includes the following columns:
- source: Original filename
- title: Wikipedia article title
- id: Wikipedia article ID (curid)
- url: Wikipedia article URL
- document_size: Size category (short/medium/large)
- domain: Content category (Islam/History/Science/Personality/General)
- sentence: Processed text content
Key Features
Balanced Distribution Algorithm
The system implements an intelligent balancing algorithm to ensure equal representation:
- Real-time Tracking: Monitors count of Short, Medium, and Large chunks
- Dynamic Prioritization: Calculates ratios and prioritizes underrepresented sizes
- Smart Ordering: Sorts sizes by current ratio (ascending) to favor underrepresented
- Controlled Randomness: 30% chance to shuffle for variety while maintaining balance
- Iterative Processing: Continuously adjusts priorities as new chunks are created
Example output:
Distribution: S:639 M:638 L:637 (nearly perfect balance!)
Smart Text Processing
Iterative Chunk Extraction
The system processes text iteratively:
- Gets balanced size order based on current distribution
- Tries to extract chunk for each size in priority order
- Validates content (word count, repetition check)
- Updates global counters and adjusts priorities
- Repeats until all text is processed or remaining text is invalid
Sentence Boundary Detection
- Urdu Punctuation Awareness: Respects Urdu sentence endings (Û) for natural text splitting
- Smart Cutting: Finds last sentence boundary within word limit range
- Fallback Mechanism: Cuts at max words if no good boundary found
- Minimum Word Enforcement: Ensures chunks meet minimum word requirements
Content Quality Validation
Repetition Detection Algorithm
def is_repetitive_content(text, threshold=0.6): words = extract_words(text) unique_words = count_unique(words) repetition_ratio = 1 - (unique_words / total_words) return repetition_ratio > threshold # 60% threshold
Short Content Validation
def is_valid_short_content(text): if word_count < 6: return False # Too short if is_repetitive_content(text): return False # Too repetitive return True
Duplicate Prevention
- URL Deduplication: Tracks seen URLs to prevent duplicate processing
- Sentence Deduplication: Tracks seen sentences to prevent duplicate chunks
- Optional Override:
--allow-duplicate-urlsflag for special cases
API Integration
The system uses the Wikipedia API with proper headers:
- User-Agent header for identification (
UnifiedApp/1.0) - Timeout handling (10 seconds)
- Error handling for HTTP errors and JSON parsing
- Lightweight API calls for better performance
- Extracts article title and content for categorization
Progress Tracking
Command Line
- File-by-file progress with completion percentage
- Real-time distribution display (S:X M:Y L:Z)
- Progress bar updates every 10 files
- Incremental save notifications
Jupyter Notebook
- Interactive progress bars with ipywidgets
- Visual percentage completion
- Real-time status updates
- Distribution summary in widget
Error Handling
The system includes comprehensive error handling for:
- Network Issues: Graceful handling of connection failures
- Wikipedia API Errors: Rate limiting (429), forbidden (403), not found (404)
- JSON Parsing: Handles malformed or empty API responses
- File I/O: Encoding fallbacks (UTF-8 with error replacement)
- Keyboard Interrupts: Ctrl+C saves progress before exiting
- Invalid Content: Warns about unprocessable text (too short/repetitive)
Graceful Degradation
try: process_all(args.mode, allow_duplicate_urls=args.allow_duplicate_urls) except KeyboardInterrupt: print("\n\nâ ï¸ Ctrl+C detected. Saving progress...") finally: if args.mode in ["excel", "both"]: save_excels() # Always save what we have
Performance Considerations
Memory Efficiency
- Streaming Processing: Processes files one at a time
- Incremental Saves: Writes to disk every 10 files to free memory
- Lazy Loading: Only loads current file content into memory
- Efficient Data Structures: Uses sets for O(1) duplicate detection
Processing Speed
- Natural Sorting: Pre-sorts files once for correct order
- Regex Optimization: Compiled patterns for faster matching
- Batch Operations: Groups related operations (e.g., Excel saves)
- Progress Tracking: Minimal overhead with efficient string formatting
Scalability
- Large Datasets: Can handle thousands of files
- Interrupt Recovery: Save progress allows resuming from interruption
- Configurable Batch Size: Adjust save frequency (currently every 10 files)
- Resource Management: Proper file handle cleanup
Testing
The system includes built-in test functionality:
python app.py test
This runs comprehensive tests for:
- Repetitive Content Detection: Tests various repetitive patterns
- Valid Content Validation: Ensures quality content passes filters
- Balanced Processing: Verifies distribution algorithm works correctly
- Sample Output: Shows example chunks with word counts
Test Output Example
𧪠Testing Content Validation Logic
==================================================
â Testing Repetitive/Invalid Content (should be rejected):
1. 'ÙØµÙÙ Ú©ÙÙØ¯Ú¯Ø§Ù.' (2 words)
Repetitive: True, Valid for short: False
â
Testing Valid Content (should be accepted):
1. 'اصÙÛ Ø§Ù
رÛÚ©Û. ÛÙØ±Ù¾Û ÙÙØ¢Ø¨Ø§Ø¯ÛÙÚº Ú©Û Ø§Ø¨ØªØ¯Ø§...' (35 words)
Repetitive: False, Valid for short: True
ð Balanced Distribution:
Short: 3 chunks (33.3%)
Medium: 3 chunks (33.3%)
Large: 3 chunks (33.3%)
Troubleshooting
Common Issues
-
Unbalanced Distribution
- The algorithm should create roughly equal numbers of each size
- If imbalanced, check if source text has sufficient variety
- Verify content validation isn't rejecting too many chunks
-
Too Many Warnings About Unprocessed Words
- Usually indicates repetitive or very short content
- Check source data quality
- Adjust repetition threshold if needed (currently 60%)
-
JSONDecodeError: Empty API responses
- Check internet connection
- Verify Wikipedia article IDs are valid
- API may be temporarily unavailable
-
403 Forbidden Errors: Wikipedia blocking requests (FIXED)
- System includes proper User-Agent headers
- Should not occur with current implementation
-
File Not Found Errors: Missing input files
- Ensure
Sample Data setfolder exists - Check for article*.txt files in the folder
- Verify file paths and permissions
- Ensure
-
Unicode/Encoding Errors: Text encoding issues
- System automatically handles encoding with fallbacks
- Uses UTF-8 with error replacement as fallback
Debug Mode
For troubleshooting, you can modify the code to add debug prints:
# In process_all function print(f"Processing file: {fname}") print(f"Found {len(matches)} documents in file") print(f"Current distribution: {global_size_counts}") print(f"Chunk created: {size} with {count_words(chunk)} words")
Monitoring Distribution
Watch the real-time distribution during processing:
â
Completed article1.txt: 5 chunks | Distribution: S:2 M:2 L:1
â
Completed article2.txt: 8 chunks | Distribution: S:5 M:5 L:3
If distribution becomes heavily skewed, it may indicate:
- Source text predominantly contains one size range
- Content validation rejecting specific sizes
- Need to adjust word count ranges
Example Workflow
Complete Processing Pipeline
-
Prepare Input Data
# Ensure Sample Data set folder contains article files # Files should be in format: <doc url="...">content</doc> -
Run Processing
# Process with both disk and Excel output python app.py --mode both -
Monitor Progress
ð Processing article1.txt (1/100) â Completed article1.txt: 5 chunks | Distribution: S:2 M:2 L:1 Files |ââââââââââââââââââââââââââââââââââââââââ| 10/100 (10.0%) processed ð¾ Saving progress... | Distribution: S:15 M:14 L:13 -
Review Output
# Check text files ls txtOutput/Short/NC/ ls txtOutput/Medium/LR/ ls txtOutput/Large/HR/ # Check Excel files ls ExcelOutput/ # short.xlsx, medium.xlsx, large.xlsx, combined.xlsx -
Analyze Results
- Open Excel files for data analysis
- Check distribution balance in summary
- Review sample content for quality
Jupyter Notebook Workflow
-
Launch Notebook
jupyter notebook app.ipynb -
Run Cells Sequentially
- Configuration cell
- Helper functions
- Processing function
- Execute processing with progress bar
-
View Results
- Summary statistics automatically displayed
- Sample content from each size category
- Export filtered datasets
Algorithm Details
Balanced Distribution Algorithm
def get_balanced_size_order(): total = sum(global_size_counts.values()) if total == 0: # First time: random order return shuffle(["Short", "Medium", "Large"]) # Calculate ratios ratios = {size: count/total for size, count in global_size_counts.items()} # Sort by ratio (ascending) - prioritize underrepresented sorted_sizes = sorted(ratios.keys(), key=lambda x: ratios[x]) # 30% chance to shuffle for variety if random.random() < 0.3: random.shuffle(sorted_sizes) return sorted_sizes
Iterative Processing Flow
Input Text
â
Get Balanced Size Order (based on current distribution)
â
Try Each Size in Priority Order
â
Extract Chunk (respect sentence boundaries)
â
Validate Content (word count, repetition)
â
If Valid: Save Chunk, Update Counters
â
Repeat with Remaining Text
â
Output: Balanced Dataset
Data Quality Metrics
Expected Output Quality
- Balanced Distribution: ±5% variance between sizes
- Minimum Word Count: All short content â¥6 words
- Repetition Ratio: All content <60% repetitive
- Unique Sentences: No duplicate sentences across dataset
- Valid URLs: All chunks linked to valid Wikipedia articles
Sample Statistics (from test run)
ð Total records created: 1914
ð Breakdown by size:
Short: 639 (33.4%)
Medium: 638 (33.3%)
Large: 637 (33.3%)
ð·ï¸ Breakdown by domain:
History: 920
General: 366
Islam: 256
Personality: 187
Science: 110
Uncategorized: 75
Contributing
- Fork the repository
- Create a feature branch
- Make your changes
- Test with
python app.py test - Ensure balanced distribution is maintained
- Submit a pull request
License
This project is for educational and research purposes. Please respect Wikipedia's terms of service and API usage guidelines.
Acknowledgments
- Wikipedia API for providing access to Urdu content
- Pandas library for efficient data manipulation
- ipywidgets for interactive notebook progress tracking
Contact
For questions or issues, please create an issue in the repository or contact the development team.
Project Timeline
Technologies
External Links
Related Projects
Projects built with similar technologies.
Online Html Editor And Viewer
The Online HTML Editor and Viewer is a simple web application built with Flask that allows users to write and preview HTML code in real-time.
Qrgen
A premium, feature-rich QR Code Generator engineered with Python (Flask) and a pristine Glassmorphism frontend.
Examina Ai
Using AI, It transforms raw study materials into structured, verified examination sets with support for institutional export formats like Moodle XML.