1. Добавь еще сохранение последней выбранной модели 2. Ollama-web это работа по API сайта ollama.com (через ключ) 3. openrouter модели отвечаю кракозябрами 4. nvidia не подтягивает модели с сайта 5. HuggingFace 404 Client Error: Not Found for url: https://api-inference.huggingface.co/models/meta-llama/Llama-3.2-1B 6. добавь автоматическую подгрузки моделей а так же выстрый поиск модели в списке # AI Model Tester - Specification ## Project Overview - **Project name**: AI Model Tester - **Type**: Web application (single-page) - **Core functionality**: Test AI models across multiple providers (Ollama, Ollama-web, OpenRouter, NVIDIA, HuggingFace) with unified interface - **Target users**: Developers testing AI model integrations ## UI/UX Specification ### Layout Structure - **Single page application** with vertical flow - **Sections**: 1. Header - Logo + title 2. Provider selector - Horizontal tabs/cards 3. Configuration panel - API key input, model selector 4. Query section - Textarea for prompt, thinking toggle 5. Response area - Streaming response display ### Visual Design - **Theme**: Dark mode with glassmorphism elements - **Typography**: - Font: 'JetBrains Mono' for code, 'Outfit' for UI - Headings: Bold, large - **Effects**: - Glass cards with backdrop-filter blur - Subtle glow on active elements - Smooth transitions (0.3s ease) - Animated gradient borders on focus ### Components 1. **Provider Tabs** - 5 providers as clickable cards - Active state: glowing border + background tint - Icons for each provider 2. **API Key Input** - Password input with show/hide toggle - Save indicator - Validation feedback 3. **Model Selector** - Dropdown loaded dynamically per provider - Refresh button for Ollama - Loading state with skeleton 4. **Thinking Toggle** - Custom styled checkbox - Label "Enable reasoning" / "Режим размышления" - Glow effect when enabled 5. **Query Input** - Large textarea with syntax highlighting feel - Character count - Submit button with loading state 6. **Response Display** - Markdown rendering - Copy button - Streaming animation (typing effect) - Thinking process reveal if enabled ## Functionality Specification ### Core Features 1. **Provider Management** - Select from: Ollama (local), Ollama-web, OpenRouter, NVIDIA, HuggingFace - Store API keys per provider in config.json - Display saved key status (masked) 2. **Model Discovery** - Ollama: GET /api/tags from local instance - Ollama-web (ollama.com): GET https://ollama.com/api/tags (Bearer key) - OpenRouter: https://openrouter.ai/api/v1/models - NVIDIA: https://integrate.api.nvidia.com/v1/models - HuggingFace: https://api-inference.huggingface.co/models (list via inference) 3. **API Key Storage** - Save to config.json via backend - Keys stored encrypted (simple base64 for demo) - Load on page init 4. **Query Execution** - Send request to appropriate API - Handle streaming responses - Display thinking/reasoning if enabled - Error handling with user-friendly messages 5. **Thinking Mode** - Enable: send to provider with reasoning enabled - Display reasoning in collapsible section - Default: disabled ### Data Flow - Frontend → Flask backend → Provider APIs - Config stored in JSON file on server ### Edge Cases - No API key entered - prompt to enter - Invalid API key - show error - Model list empty - show "No models found" - Network error - show retry option - Ollama not running - detect and show message ## Acceptance Criteria - [ ] All 5 providers selectable and functional - [ ] API keys save to config.json - [ ] Models load dynamically per provider - [ ] Queries send and receive responses - [ ] Thinking mode toggles correctly - [ ] Responsive on desktop (mobile optional) - [ ] Dark theme with glassmorphism working - [ ] Error states handled gracefully ## Files - `server.py` - Flask backend - `index.html` - Frontend - `config.json` - Data storage (auto-created)