linear-coding-agent

Author	SHA1	Message	Date
David Blanc Brioir	9e657cbf29	Add unified memory tools: search_memories, trace_concept_evolution, check_consistency, update_thought_evolution_stage - Create memory/mcp/unified_tools.py with 4 new handlers: - search_memories: unified search across Thoughts and Conversations - trace_concept_evolution: track concept development over time - check_consistency: verify statement alignment with past content - update_thought_evolution_stage: update thought maturity stage - Export new tools from memory/mcp/__init__.py - Register new tools in mcp_server.py with full docstrings These tools complete the Ikario memory toolset to match memoryTools.js expectations. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-30 17:24:30 +01:00
David Blanc Brioir	376d77ddfa	fix: Replace text2vec-transformers with BGE-M3 in MCP retrieval tools The text2vec-transformers Docker service was removed in Jan 2026, but retrieval_tools.py still used near_text() which requires it. Now uses GPU embedder (BGE-M3) with near_vector() like flask_app.py. Changes: - Add GPU embedder singleton (get_gpu_embedder) - search_chunks_handler: near_text → near_vector + BGE-M3 - search_summaries_handler: near_text → near_vector + BGE-M3 Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-30 00:34:50 +01:00
David Blanc Brioir	119ad7ebc3	fix: Show RAG context sidebar when chunks are received The sidebar content was hidden by default (display: none) but displayContext() never made it visible when chunks arrived. Added sidebarContent.style.display = 'block' to show the context panel with all RAG chunks. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-15 21:48:39 +01:00
David Blanc Brioir	2a8098f17a	chore: Clean up obsolete files and add Puppeteer chat test - Remove obsolete documentation, examples, and utility scripts - Remove temporary screenshots and test files from root - Add test_chat_backend.js for Puppeteer testing of chat RAG - Update .gitignore Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-15 21:40:56 +01:00
David Blanc Brioir	1bf570e201	refactor: Rename Chunk_v2/Summary_v2 collections to Chunk/Summary - Add migrate_rename_collections.py script for data migration - Update flask_app.py to use new collection names - Update weaviate_ingest.py to use new collection names - Update schema.py documentation - Update README.md and ANALYSE_MCP_TOOLS.md Migration completed: 5372 chunks + 114 summaries preserved with vectors. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-14 23:59:03 +01:00
David Blanc Brioir	e78c3ae292	docs: Update README to reflect 6 collections (3 RAG + 3 Memory) Architecture clarification: - Updated schema section: 4 → 6 collections - Clarified separation: RAG (3) vs Memory (3) - Removed Document collection references - Updated collection names: Chunk → Chunk_v2, Summary → Summary_v2 Schema changes reflected: - RAG: Work, Chunk_v2, Summary_v2 (schema.py) - Memory: Conversation, Message, Thought (memory/schemas/memory_schemas.py) Vectorization details: - All 5 vectorized collections use GPU embedder (BAAI/bge-m3, RTX 4070) - Manual vectorization with Python PyTorch CUDA - 1024 dimensions, cosine similarity Updated diagrams: - Architecture mermaid diagram shows 6 collections - Pipeline diagram updated to 6 collections - Added memory/ module structure Updated examples: - Replaced Chunk with Chunk_v2 in all code examples - Added Memory collections documentation - Clarified separation of concerns Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-09 14:36:08 +01:00
David Blanc Brioir	53f6a92365	feat: Remove Document collection from schema BREAKING CHANGE: Document collection removed from Weaviate schema Architecture simplification: - Removed Document collection (unused by Flask app) - All metadata now in Work collection or file-based (chunks.json) - Simplified from 4 collections to 3 (Work, Chunk_v2, Summary_v2) Schema changes (schema.py): - Removed create_document_collection() function - Updated verify_schema() to expect 3 collections - Updated display_schema() and print_summary() - Updated documentation to reflect Chunk_v2/Summary_v2 Ingestion changes (weaviate_ingest.py): - Removed ingest_document_metadata() function - Removed ingest_document_collection parameter - Updated IngestResult to use work_uuid instead of document_uuid - Removed Document deletion from delete_document_chunks() - Updated DeleteResult TypedDict Type changes (types.py): - WeaviateIngestResult: document_uuid → work_uuid Documentation updates (.claude/CLAUDE.md): - Updated schema diagram (4 → 3 collections) - Removed Document references - Updated to reflect manual GPU vectorization Database changes: - Deleted Document collection (13 objects) - Deleted Chunk collection (0 objects, old schema) Benefits: - Simpler architecture (3 collections vs 4) - No redundant data storage - All metadata available via Work or file-based storage - Reduced Weaviate memory footprint Migration: - See DOCUMENT_COLLECTION_ANALYSIS.md for detailed analysis - See migrate_chunk_v2_to_none_vectorizer.py for vectorizer migration Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-09 14:13:51 +01:00
David Blanc Brioir	c90864e9f7	docs: Remove obsolete documentation files from library_rag Cleaned up 15 obsolete MD files that were temporary session reports and outdated documentation, now replaced by the organized docs/ structure and comprehensive README.md. Removed files: - ANALYSE_ARCHITECTURE_WEAVIATE.md (superseded by docs/migration-gpu/) - ANALYSE_RAG_FINAL.md (session report) - ANALYSE_RESULTATS_RESUME.md (session report) - COMPLETE_SESSION_RECAP.md (session report) - EXPLICATION_SUMMARY_CHUNK.md (old technical docs) - FIX_HIERARCHICAL.md (session report) - INTEGRATION_SUMMARY.md (session report) - PLAN_LLM_SUMMARIZER.md (old planning docs) - QUICKSTART_SUMMARY_SEARCH.md (superseded by README.md) - README_SEARCH.md (superseded by README.md) - REFACTOR_SUMMARY.md (session report) - SESSION_SUMMARY.md (session report) - TTS_INSTALLATION_GUIDE.md (not used) - WEAVIATE_GUIDE_COMPLET.md (superseded by README.md) - WEAVIATE_SCHEMA.md (superseded by schema.py comments) Retained documentation: ✓ README.md (main documentation) ✓ docs/ (organized migration and project docs) ✓ docs_techniques/ (technical specifications) ✓ .claude/CLAUDE.md (Claude Code instructions) ✓ examples/ (usage examples) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-09 12:52:32 +01:00
David Blanc Brioir	eb2bf45281	chore: Update project configuration and improve chat prompts Configuration Updates: - .claude/settings.local.json: Add security permissions for WebFetch, WebSearch, nvidia-smi - package.json: Add puppeteer dependency for browser automation tests - package-lock.json: Update lockfile with puppeteer@24.34.0 and dependencies - Remove root .env.example (superseded by generations/library_rag/.env.example) Flask App Improvements: - Enhanced chat prompt to REQUIRE "Sources utilisées" section in responses - Added explicit warnings against inventing citations not in provided passages - Improved source citation format with mandatory author, work, and passage number - Strengthened instructions to prevent hallucinated references Benefits: - Chat responses now consistently include proper source citations - Better academic rigor in philosophical analyses - Prevents LLM from inventing non-existent references - Automated testing infrastructure with Puppeteer Related to GPU embedder migration testing and validation. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-09 12:45:20 +01:00
David Blanc Brioir	a3d5e8935f	refactor: Remove Docker text2vec-transformers service (GPU embedder only) BREAKING CHANGE: Docker text2vec-transformers service removed Changes: - Removed text2vec-transformers service from docker-compose.yml - Removed ENABLE_MODULES and DEFAULT_VECTORIZER_MODULE from Weaviate config - Updated architecture comments to reflect Python GPU embedder only - Simplified docker-compose to single Weaviate service Architecture: Before: Weaviate + text2vec-transformers (2 services) After: Weaviate only (1 service) Vectorization: - Ingestion: Python GPU embedder (manual vectorization) - Queries: Python GPU embedder (manual vectorization) - No auto-vectorization modules needed Benefits: - RAM: -10 GB freed (no text2vec-transformers container) - CPU: -3 cores freed - Architecture: Simplified (one service instead of two) - Maintenance: Easier (no Docker service dependencies) Validation: ✅ Weaviate starts correctly without text2vec-transformers ✅ Existing data accessible (5355 chunks preserved) ✅ API endpoints respond correctly ✅ No errors in startup logs Migration: GPU embedder already tested and validated See: TESTS_COMPLETS_GPU_EMBEDDER.md Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-09 12:07:09 +01:00
David Blanc Brioir	17dfe213ed	feat: Migrate Weaviate ingestion to Python GPU embedder (30-70x faster) BREAKING: No breaking changes - zero data loss migration Core Changes: - Added manual GPU vectorization in weaviate_ingest.py (~100 lines) - New vectorize_chunks_batch() function using BAAI/bge-m3 on RTX 4070 - Modified ingest_document() and ingest_summaries() for GPU vectors - Updated docker-compose.yml with healthchecks Performance: - Ingestion: 500-1000ms/chunk → 15ms/chunk (30-70x faster) - VRAM usage: 2.6 GB peak (well under 8 GB available) - No degradation on search/chat (already using GPU embedder) Data Safety: - All 5355 existing chunks preserved (100% compatible vectors) - Same model (BAAI/bge-m3), same dimensions (1024) - Docker text2vec-transformers optional (can be removed later) Tests (All Passed): ✅ Ingestion: 9 chunks in 1.2s ✅ Search: 16 results, GPU embedder confirmed ✅ Chat: 11 chunks across 5 sections, hierarchical search OK Architecture: Before: Hybrid (Docker CPU for ingestion, Python GPU for queries) After: Unified (Python GPU for everything) Files Modified: - generations/library_rag/utils/weaviate_ingest.py (GPU vectorization) - generations/library_rag/.claude/CLAUDE.md (documentation) - generations/library_rag/docker-compose.yml (healthchecks) Documentation: - MIGRATION_GPU_EMBEDDER_SUCCESS.md (detailed report) - TEST_FINAL_GPU_EMBEDDER.md (ingestion + search tests) - TEST_CHAT_GPU_EMBEDDER.md (chat test) - TESTS_COMPLETS_GPU_EMBEDDER.md (complete summary) - BUG_REPORT_WEAVIATE_CONNECTION.md (initial bug analysis) - DIAGNOSTIC_ARCHITECTURE_EMBEDDINGS.md (technical analysis) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-09 11:44:10 +01:00
David Blanc Brioir	0c8ea8fa48	fix: Correct Work titles and improve LLM metadata extraction Fixes issue where LLM was copying placeholder instructions from the prompt template into actual metadata fields. Changes: 1. Created fix_work_titles.py script to correct existing bad titles - Detects patterns like "(si c'est bien...)", "Titre corrigé...", "Auteur à identifier" - Extracts correct metadata from chunks JSON files - Updates Work entries and associated chunks (44 chunks updated) - Fixed 3 Works with placeholder contamination 2. Improved llm_metadata.py prompt to prevent future issues - Added explicit INTERDIT/OBLIGATOIRE rules with ❌/✅ markers - Replaced placeholder examples with real concrete examples - Added two example responses (high confidence + low confidence) - Final empty JSON template guides structure without placeholders - Reinforced: use "confidence" field for uncertainty, not annotations Results: - "A Cartesian critique... (si c'est bien le titre)" → "A Cartesian critique of the artificial intelligence" - "Titre corrigé si nécessaire (ex: ...)" → "Computationalism and The Case When the Brain Is Not a Computer" - "Titre de l'article principal (à identifier)" → "Computationalism in the Philosophy of Mind" All future document uploads will now extract clean metadata without LLM commentary or placeholder instructions. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 23:59:25 +01:00
David Blanc Brioir	0c3b6c5fea	feat: Auto-create Work entries during document ingestion Adds automatic Work object creation to ensure all uploaded documents appear on the /documents page. Previously, chunks were ingested but Work entries were missing, causing documents to be invisible in the UI. Changes: - Add create_or_get_work() function to weaviate_ingest.py - Checks for existing Work by sourceId (prevents duplicates) - Creates new Work with metadata (title, author, year, pages) - Returns UUID for potential future reference - Integrate Work creation into ingest_document() flow - Add helper scripts for retroactive fixes and verification: - create_missing_works.py: Create Works for already-ingested documents - reingest_batch_documents.py: Re-ingest documents after bug fixes - check_batch_results.py: Verify batch upload results in Weaviate This completes the batch upload feature - documents now properly appear on /documents page immediately after ingestion. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 23:34:06 +01:00
David Blanc Brioir	b8d94576de	fix: Correct Weaviate ingestion for Chunk_v2 schema compatibility Fixes batch upload ingestion that was failing silently due to schema mismatches: Schema Fixes: - Update collection names from "Chunk" to "Chunk_v2" - Update collection names from "Summary" to "Summary_v2" Object Structure Fixes: - Replace nested objects (work: {title, author}) with flat fields - Use workTitle and workAuthor instead of nested work object - Add year field to chunks - Remove document nested object (not used in current schema) - Disable nested objects validation for flat schema Impact: - Batch upload now successfully ingests chunks to Weaviate - Single-file upload also benefits from fixes - All new documents will be properly indexed and searchable Testing: - Verified with 2-file batch upload (7 + 11 chunks = 18 total) - Total chunks increased from 5,304 to 5,322 - All chunks properly searchable with workTitle/workAuthor filters Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 23:25:36 +01:00
David Blanc Brioir	b70b796ef8	feat: Add multi-file batch upload with sequential processing Implements comprehensive batch upload system with real-time progress tracking: Backend Infrastructure: - Add batch_jobs global dict for batch orchestration - Add BatchFileInfo and BatchJob TypedDicts to utils/types.py - Create run_batch_sequential() worker function with thread.join() synchronization - Modify /upload POST route to detect single vs multi-file uploads - Add 3 batch API routes: /upload/batch/progress, /status, /result - Add timestamp_to_date Jinja2 template filter Frontend: - Update upload.html with 'multiple' attribute and file counter - Create upload_batch_progress.html: Real-time dashboard with SSE per file - Create upload_batch_result.html: Final summary with statistics Architecture: - Backward compatible: single-file upload unchanged - Sequential processing: one file after another (respects API limits) - N parallel SSE connections: one per file for real-time progress - Polling mechanism to discover job IDs as files start processing - 1-hour timeout per file with error handling and continuation Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 22:41:52 +01:00
David Blanc Brioir	7a7a2b8e19	feat: Improve chat page filters layout - Works filter section: Increase max-height from 250px to 70vh (full screen) - Context RAG section: Closed by default (display: none) - Mobile responsive: Adjust works filter to 50vh on mobile - Enhances visibility of available works at page load Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 22:31:07 +01:00
David Blanc Brioir	2f34125ef6	feat: Add Memory system with Weaviate integration and MCP tools MEMORY SYSTEM ARCHITECTURE: - Weaviate-based memory storage (Thought, Message, Conversation collections) - GPU embeddings with BAAI/bge-m3 (1024-dim, RTX 4070) - 9 MCP tools for Claude Desktop integration CORE MODULES (memory/): - core/embedding_service.py: GPU embedder singleton with PyTorch - schemas/memory_schemas.py: Weaviate schema definitions - mcp/thought_tools.py: add_thought, search_thoughts, get_thought - mcp/message_tools.py: add_message, get_messages, search_messages - mcp/conversation_tools.py: get_conversation, search_conversations, list_conversations FLASK TEMPLATES: - conversation_view.html: Display single conversation with messages - conversations.html: List all conversations with search - memories.html: Browse and search thoughts FEATURES: - Semantic search across thoughts, messages, conversations - Privacy levels (private, shared, public) - Thought types (reflection, question, intuition, observation) - Conversation categories with filtering - Message ordering and role-based display DATA (as of 2026-01-08): - 102 Thoughts - 377 Messages - 12 Conversations DOCUMENTATION: - memory/README_MCP_TOOLS.md: Complete API reference and usage examples All MCP tools tested and validated (see test_memory_mcp_tools.py in archive). Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 18:08:13 +01:00
David Blanc Brioir	187ba4854e	chore: Major cleanup - archive migration scripts and remove temp files CLEANUP ACTIONS: - Archived 11 migration/optimization scripts to archive/migration_scripts/ - Archived 11 phase documentation files to archive/documentation/ - Moved backups/, docs/, scripts/ to archive/ - Deleted 30+ temporary debug/test/fix scripts - Cleaned Python cache (__pycache__/, .pyc) - Cleaned log files (.log) NEW FILES: - CHANGELOG.md: Consolidated project history and migration documentation - Updated .gitignore: Added .log, .pyc, archive/ exclusions FINAL ROOT STRUCTURE (19 items): - Core framework: agent.py, autonomous_agent_demo.py, client.py, security.py, progress.py, prompts.py - Config: requirements.txt, package.json, .gitignore - Docs: README.md, CHANGELOG.md, project_progress.md - Directories: archive/, generations/, memory/, prompts/, utils/ ARCHIVED SCRIPTS (in archive/migration_scripts/): 01-11: Migration & optimization scripts (migrate, schema, rechunk, vectorize, etc.) ARCHIVED DOCS (in archive/documentation/): PHASE_0-8: Detailed phase summaries MIGRATION_README.md, PLAN_MIGRATION_WEAVIATE_GPU.md Repository is now clean and production-ready with all important files preserved in archive/. Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 18:05:43 +01:00
David Blanc Brioir	7045907173	feat: Optimize chunk sizes with 1000-word limit and overlap Implemented chunking optimization to resolve oversized chunks and improve semantic search quality: CHUNKING IMPROVEMENTS: - Added strict 1000-word max limit (vs previous 1500-2000) - Implemented 100-word overlap between consecutive chunks - Created llm_chunker_improved.py with overlap functionality - Added 3 fallback points in llm_chunker.py for robustness RE-CHUNKING RESULTS: - Identified and re-chunked 31 oversized chunks (>2000 tokens) - Split into 92 optimally-sized chunks (max 1995 tokens) - Preserved all metadata (workTitle, workAuthor, sectionPath, etc.) - 0 chunks now exceed 2000 tokens (vs 31 before) VECTORIZATION: - Created manual vectorization script for chunks without vectors - Successfully vectorized all 92 new chunks (100% coverage) - All 5,304 chunks now have BGE-M3 embeddings DOCKER CONFIGURATION: - Exposed text2vec-transformers port 8090 for manual vectorization - Added cluster configuration to fix "No private IP address found" - Increased worker timeout to 600s for large chunks TESTING: - Created comprehensive search quality test suite - Tests distribution, overlap detection, and semantic search - Modified to use near_vector() (Chunk_v2 has no vectorizer) Scripts: - 08_fix_summaries_properties.py - Add missing Work metadata to summaries - 09_rechunk_oversized.py - Re-chunk giant chunks with overlap - 10_test_search_quality.py - Validate search improvements - 11_vectorize_missing_chunks.py - Manual vectorization via API Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-08 17:37:49 +01:00
David Blanc Brioir	ca221887eb	docs: Update README for schema changes and Docker config - Add 'summary' vectorized field to Chunk collection description - Update vectorization strategy (text/summary/keywords) - Add HNSW + RQ vector index configuration section - Correct Docker config: BGE-M3 ONNX is CPU-only (not CUDA) - Add llm_summarizer.py and summary generation scripts to project structure - Update annexe with accurate GPU/VRAM information - Remove incorrect GPU configuration example 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-07 23:10:36 +01:00
David Blanc Brioir	636ad6206c	feat: Add vectorized summary field and migration tools - Add 'summary' field to Chunk collection (vectorized with text2vec) - Migrate from Dynamic index to HNSW + RQ for both Chunk and Summary - Add LLM summarizer module (utils/llm_summarizer.py) - Add migration scripts (migrate_add_summary.py, restore_.py) - Add summary generation utilities and progress tracking - Add testing and cleaning tools (outils_test_and_cleaning/) - Add comprehensive documentation (ANALYSE_.md, guides) - Remove obsolete files (linear_config.py, old test files) - Update .gitignore to exclude backups and temp files 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-07 22:56:03 +01:00
David Blanc Brioir	feb215dae0	revert: Remove max-height from works-list (causes double scrollbar) - Removed max-height: 300px from .works-list - Keeps only the Unicode encoding fix (→ to ->) - Avoids having two scrollbars in the works filter section 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-04 16:50:29 +01:00
David Blanc Brioir	6596a4e32f	fix: Resolve works filter display and encoding issues Problem 1: Only 3 works visible despite 8/10 badge - Added max-height: 300px and overflow-y: auto to .works-list - Now all 10 works are scrollable in the filter section Problem 2: UnicodeEncodeError with → character in console - Replaced Unicode arrow (→) with ASCII arrow (->) in print statements - Fixes 'charmap' codec error on Windows console 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-04 16:47:28 +01:00
David Blanc Brioir	a73ed2d98e	chore: Add autonomous agent infrastructure and cleanup old files - Disable CLAUDE.md confirmation rules for autonomous agent operation - Add utility scripts: check_linear_status.py, check_meta_issue.py, move_issues_to_todo.py - Add works filter specification: prompts/app_spec_works_filter.txt - Update .linear_project.json with works filter issues - Remove old/stale scripts and documentation files - Update search.html template This commit completes the infrastructure for the autonomous agent that successfully implemented all 13 works filter issues (LRP-136 to LRP-148). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-04 16:42:42 +01:00
David Blanc Brioir	fe085c7ebe	LRP-148: Add user guide documentation for works filter - WORKS_FILTER.md: Complete user documentation in French - Feature overview and location - Selection/deselection instructions - Quick action buttons (Tout/Aucun) - Badge counter explanation - Collapse functionality - Default behavior and localStorage persistence - Impact on semantic search - Recommended use cases (comparative study, focus, exclusion) - Responsive mobile support - API Reference section: - GET /api/get-works endpoint documentation - POST /chat/send selected_works parameter - Error codes and validation - Troubleshooting guide: - No works displayed - Filter not working - How to reset selection - Chunks count explanation 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 16:32:25 +01:00
David Blanc Brioir	c533f67e2f	LRP-146: Add unit tests for works filter backend routes - Test /api/get-works route: - Unique works extraction with correct chunk counts - Sorting by author then title - Connection failure and query exception handling - Edge cases: empty database, missing title/author - Test /chat/send selected_works parameter: - Accepts empty list (search all works) - Accepts valid work title list - Rejects non-list types (string, dict) - Rejects mixed types in list - Verifies parameter passed to background thread - Test rag_search works filter: - No filter when selected_works is empty/None - Contains_any filter applied when works selected 18 tests, all passing, no real Weaviate calls (fully mocked) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 16:27:41 +01:00
David Blanc Brioir	82da123ef7	feat: Implement works filter UI for chat page (LRP-139, 140, 141, 143) - Add works filter section HTML above Context RAG sidebar - Add CSS styles for works filter with checkboxes, badges, and collapse - Implement JavaScript for loading works from /api/get-works - Add localStorage persistence for selected works - Integrate selected_works parameter with /chat/send API call - Add Tout/Aucun buttons for quick selection - Add collapsible section with chevron toggle - Responsive design for mobile screens 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 14:46:48 +01:00
David Blanc Brioir	7e8367863d	LRP-138: Implement Weaviate filter for selected_works in chat search - Add selected_works parameter to rag_search() function - Build Weaviate filter using Filter.by_property("workTitle").contains_any() - Add selected_works parameter to diverse_author_search() function - Pass selected_works from run_chat_generation to diverse_author_search - Preserve work filter in fallback search path - Add logging for applied work filters The filter allows restricting RAG search to specific works selected by the user. When selected_works is empty or None, all works are searched (no filter). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 14:32:47 +01:00
David Blanc Brioir	930615d239	feat: Add selected_works parameter to /chat/send route - Add optional selected_works parameter to /chat/send endpoint - Validate that selected_works is a list of strings - Pass parameter to run_chat_generation function - Backward compatible (works without the parameter) - Add logging for selected_works filter Linear issue: LRP-137 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 13:29:37 +01:00
David Blanc Brioir	d106e91d56	feat: Add /api/get-works route for works filtering - Add new API endpoint GET /api/get-works - Returns JSON array of all unique works with metadata - Each work includes: title, author, chunks_count - Results sorted by author then title - Proper error handling for Weaviate connection issues - Fixed gRPC serialization issue with nested objects Linear issue: LRP-136 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-04 13:23:24 +01:00
David Blanc Brioir	8c0e1cef0d	refactor: Integrate summary search into dropdown and fix hierarchical mode Previously created a separate page for summary search, which was redundant since hierarchical mode already demonstrates the summary→chunk pattern. Refactored to integrate summary-only mode as a dropdown option in the main search interface, reducing code duplication by ~370 lines. Also fixed critical bug in hierarchical search where return_properties excluded the nested "document" object, causing source_id to be empty and all sections to be filtered out. Solution: removed return_properties to let Weaviate return all properties including nested objects. All 4 search modes now functional: - Auto-detection (default) - Simple chunks (10% visibility) - Hierarchical summary→chunks (variable) - Summary-only (90% visibility) Tests: 14/14 passed for dropdown integration, hierarchical mode confirmed working with 13 passages across 4 section groups. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-03 17:59:58 +01:00
David Blanc Brioir	b76e56e62e	refactor: Suppression tous fonds beiges header section - Retiré fond beige dégradé du section-header - Retiré fond beige des lignes 2 et 3 - Retiré padding et border-radius des lignes 2-3 - Présentation ultra-épurée : texte simple + icônes - Garde uniquement bordure accent en bas du header	2026-01-02 00:07:17 +01:00
David Blanc Brioir	c0cef02990	refactor: Présentation strictement IDENTIQUE lignes 2 et 3 - Ligne 2 et 3 ont exactement le même style CSS - Même couleur (var(--color-accent)) - Même background beige (rgba(125, 110, 88, 0.08)) - Même padding (0.25rem 0.5rem) - Même border-radius (4px) - Seule différence : icône et contenu - Présentation ultra-cohérente visuellement	2026-01-02 00:03:13 +01:00
David Blanc Brioir	77473f9060	refactor: Uniformisation complète police et style lignes 2-3 - Retiré font-weight: 600 du titre section - Lignes 2 et 3 ont maintenant exactement le même style - Police par défaut, pas de variations de graisse - Présentation ultra-simplifiée et cohérente - Seule différence : couleurs (accent vs text-strong)	2026-01-02 00:01:59 +01:00
David Blanc Brioir	a8dbe40d50	refactor: Harmonisation police lignes 2 et 3 du header section - Ligne 2 (hiérarchie) : police normale, pas de font-size - Ligne 3 (titre) : police normale, pas de font-size ni font-family spéciale - Changé h4 en span pour cohérence typographique - Gardé font-weight: 600 sur le titre pour légère emphase - Résultat : lignes 2 et 3 visuellement cohérentes	2026-01-02 00:00:24 +01:00
David Blanc Brioir	3d20a54d06	refactor: Réorganisation header section en 3 lignes claires Ligne 1 : Auteur \| Œuvre \| Similarité \| Nb passages Ligne 2 : 🗂️ Hiérarchie (chapterTitle) Ligne 3 : 📂 Titre section Plus compact et hiérarchie mieux visible avant le titre	2026-01-01 23:59:00 +01:00
David Blanc Brioir	6a2ec10d7b	feat: Ajout auteur, œuvre et hiérarchie dans header section - Badge auteur (récupéré du premier chunk de la section) - Badge œuvre (récupéré du premier chunk de la section) - Hiérarchie complète avec icône 🗂️ (chapterTitle du premier chunk) Ex: "Peirce: CP 7.316" - Fond beige léger pour la hiérarchie - Affichage au-dessus du titre de section Structure header de section: 1. Auteur + Œuvre (badges) 2. Titre section avec icône 📂 3. Hiérarchie complète (chapterTitle) 4. Similarité + nombre passages 5. Résumé LLM 6. Concepts	2026-01-01 23:54:44 +01:00
David Blanc Brioir	9c63ef84da	feat: Amélioration hiérarchie visuelle sections/chunks - Header section avec fond beige dégradé distinct des chunks - Icône 📂 + label "Section :" explicite avant le titre - Titre section en plus gros (1.2em, font-weight 600) - Badge nombre de passages en couleur accent - Zone chunks avec fond blanc pur pour contraster - Bordure section plus épaisse (2px) et arrondie (10px) - Summary text avec fond blanc semi-transparent pour lisibilité - Label "Concepts :" avant la liste des concepts Résultat: Hiérarchie visuelle très claire entre section et passages	2026-01-01 23:31:31 +01:00
David Blanc Brioir	1cec07b284	feat: Group chunks under sections in hierarchical search - Stage 2 now searches chunks for EACH section using section summary as query - Chunks distributed across sections (limit / sections_limit) - Template displays sections with nested chunks underneath - Each section shows: title, summary, concepts, chunk count, and passages - Removes separate global passages list - now fully grouped by section Structure: Section 1 → Chunks 1-3, Section 2 → Chunks 4-6, etc.	2026-01-01 18:25:11 +01:00
David Blanc Brioir	65adc02d6e	fix: Hide duplicate summary text when identical to title Problem: Sections showed title twice (once as title, once as summary_text) Cause: summary_text contains same content as title in current data Solution: Only show summary_text if different from title and section_path Condition: summary_text != title AND summary_text != section_path	2026-01-01 16:16:50 +01:00
David Blanc Brioir	109d16b223	fix: Correct Jinja2 template syntax error (missing endif removal) Error: 'Encountered unknown tag else' - endif was closing the if block too early Fix: Removed extra {% endif %} before {% else %} - Line 232: Removed incorrect closing tag - The {% else %} at line 234 is part of the hierarchical/simple mode conditional - Proper structure: if hierarchical ... else simple ... endif Tests: - Template syntax validates ✓ - Search page loads ✓ - Hierarchical mode works ✓	2026-01-01 15:54:44 +01:00
David Blanc Brioir	d824269606	fix: Adapt hierarchical display for mismatched sectionPath formats Root cause: - Summary.sectionPath: '635. As for the subject...' (paragraph numbers) - Chunk.sectionPath: 'Peirce: CP 4.47 > 47. §3 THE NATURE...' (canonical refs) - No way to match them with prefix/equal filters Solution (workaround until summaries are regenerated): - Show sections as context (relevant high-level topics found) - Show chunks globally (top 20 most relevant passages) - Don't try to group chunks under sections UI changes: - '📚 Sections pertinentes trouvées' (context cards with summary) - '📄 Passages les plus pertinents' (top chunks, not grouped) - Cleaner, more honest representation of what we found Next steps to fully fix: - Regenerate Summary collection with correct sectionPath format - Or create a mapping between Summary titles and Chunk sectionPaths	2026-01-01 15:51:11 +01:00
David Blanc Brioir	47cf21867f	fix: Use prefix matching for sectionPath to find chunks in sections Problem: - Summary.sectionPath: "Peirce: CP 2.504" - Chunk.sectionPath: "Peirce: CP 2.504 > 504. Text..." - Filter.equal() found 0 matches (no exact match exists) Solution: - Single semantic query to get all relevant chunks - Distribute chunks to sections using Python startswith() - This correctly matches chunks to their parent sections Performance improvement: - 1 query instead of N queries (one per section) - Python-side filtering is fast for small result sets Result: Chunks should now appear in their corresponding sections	2026-01-01 15:45:37 +01:00
David Blanc Brioir	474edf75e5	fix: Display work/author metadata and improve section titles Backend fix: - Remove return_properties from hierarchical chunk query - Weaviate returns nested objects (work, document) when return_properties is not specified - This allows chunks to have work.author and work.title available Frontend improvements: - Truncate long section titles to 80 chars with ellipsis - Hide section_path if identical to title (avoid duplication) - Work and author badges should now display correctly in chunk metadata	2026-01-01 15:42:03 +01:00
David Blanc Brioir	80464f9f69	feat: Add author/work/hierarchy display and align colors with design charter Hierarchical search improvements: - Display author and work for each chunk using badge-author and badge-work - Show section hierarchy (sectionPath) in chunk metadata - Add 📍 icon for section path in headers Color alignment with charter: - Replace Bootstrap colors (#007bff, #28a745, #6c757d) with charter variables - section-group: border and shadow use accent colors (125,110,88) - section-header: border uses var(--color-accent) - chunk-item: border-left uses var(--color-accent-alt) - Mode badges: hierarchical=accent-alt, simple=accent - Concept badges: subtle beige background with accent border - Alert boxes: beige background instead of yellow Visual improvements: - Add hover transform effect on chunks (translateX) - Smoother color transitions using CSS variables	2026-01-01 15:39:07 +01:00
David Blanc Brioir	f49279fee3	fix: Remove nested objects from return_properties to fix gRPC serialization error - Remove 'document' from Summary query return_properties - Remove 'work' from Document query return_properties - Nested objects (OBJECT datatype) cause gRPC proto serialization error - Weaviate should return nested objects automatically without explicit request - Fixes: 'proto: invalid type: map[string]interface {}' error	2026-01-01 15:30:54 +01:00
David Blanc Brioir	9c6ba3f4a1	fix: Prevent context manager conflict by never calling simple_search from hierarchical_search - Add @contextmanager decorator for proper exception handling - Remove all simple_search() calls from within hierarchical_search() - Return mode='error' to signal fallback needed - Handle fallback in search_passages() (outside context manager) - This eliminates 'generator didn't stop after throw()' error	2026-01-01 15:27:14 +01:00
David Blanc Brioir	22ac9a030e	fix: Never call simple_search from exception handler during context cleanup	2026-01-01 15:19:33 +01:00
David Blanc Brioir	4492814891	fix: Exit context manager before calling simple_search in exception handler	2026-01-01 15:16:44 +01:00
David Blanc Brioir	8153ea35a4	fix: Prevent context manager conflict in hierarchical_search ## Problem "generator didn't stop after throw()" error when hierarchical_search falls back to simple_search. Both functions use 'with get_weaviate_client()', creating nested context managers on the same generator. ## Solution - Use ValueError("FALLBACK_TO_SIMPLE") signal instead of calling simple_search() inside the context manager - Catch ValueError in except block and call simple_search() outside context - Applied to all 3 fallback points: 1. No Weaviate client 2. No summaries found (Stage 1) 3. No sections after filtering ## Result Fallback now works correctly without context manager conflicts. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>	2026-01-01 15:10:06 +01:00

1 2

68 Commits