Commit History
Add Qwen 2.5 1.5B (GGUF Q2_K) model (~676 MB) to catalog 4b88bd9
Remove legacy static frontend prototype and set root API status endpoint fc5e529
fix(llm): use Qwen 2.5 3B Q2_K/IQ2_M available in bartowski repo and sanitize fallback repo names 0de6c12
feat(llm): add Qwen 2.5 3B 1.6-bit IQ1_S and expose dynamic model catalog in GET /api/models 0c6ae44
refactor(llm): modularize LLM with BaseLLMEngine ABC, engines provider, use_rag bypass, and remove legacy parser 8433fda
fix(rag): clean markdown syntax, reframe prompt, and add repetition penalty to prevent SLM regurgitation 5d5d9fd
perf: set default to Qwen 0.5B Q2_K 2-bit quantization, increase n_ctx to 3072 and max tokens to 1024 47a6eb0
fix(llm): map Ollama model names to HF models and isolate Pinecone session namespaces 2b7fc7a
Add session manager component with targeted session termination and lifecycle APIs 357974c
perf: add real-time SSE token streaming, stop tokens, and embedding singleton 49fbea4
perf: add OpenMP thread pinning, passive wait policy, and core utilization tuning 534bb06
feat: enable 2-bit Q2_K/IQ2 GGUF quantization and CPU flash_attn batching 65ebd96
perf: add startup warmup, vectorstore caching, 512-token detailed prompt 9d4d031
feat: add dynamic set-engine endpoint and llama.cpp vs transformers toggle f5f1b16
build: install pre-compiled CPU wheel for llama-cpp-python a2d85b4
added cpp inference 940dce9
added docs api correction edc09af
added diff models a083144
Changed Embedding model b7d61ce
rename to pdf in controller and changed embedding 1cbfd3a
modularized chuncker and pipeline to service folder 9fe5723
Modularized code and detached FAISS d383dfc
Added optional include 23b848e
Always query Pinecone regardless of session_id; filter only when session scoped fdf50e0
Switch LLM to Qwen2.5-0.5B local; add retrieval debug logging 5711e8c
Running Qwen0.5B in HF space 5c1cf07
Add Pinecone PDF ingestion pipeline 21ae65e
Deploy update 343f3a5
add browser-side PDF text extraction and parsing support 866953f
update frontend UI theme and threshold-based prompt retrieval fallback 41ef344
feat: support client-side context to bypass FAISS retrieval d49e446
feat: design clean compact single-line chat input box ae446b8
feat: deploy premium dark mode default and glassmorphic text bar updates 3fae50e
feat: add /api/ping endpoint for frontend cold-start detection and queueing handshake d0043be
fix max_tokens=0 Bad Request error in ChatHuggingFace 50fb034
Improve heartbeat, error handling, and general answering logic f420708
Optimize for mobile and improve WCAG accessibility 34c51c6
HARI PREETHAM commited on
Downgrade numpy to <2 to fix FAISS compatibility 339fc37
HARI PREETHAM commited on
Update RAG system with new chatbot UI and unified server 3fe38c2
HARI PREETHAM commited on
Switch to Mistral-7B for serverless inference reliability fb0bf22
HARI PREETHAM commited on
Fix HuggingFaceEndpoint validation error - remove repo_id when endpoint_url is used 94a7e9b
HARI PREETHAM commited on
Re-enable FAISS allow_dangerous_deserialization for new LangChain version 287934d
HARI PREETHAM commited on
Upgrade LangChain to v0.2 to support new HuggingFace Router URL ffd5a03
HARI PREETHAM commited on
Fix HuggingFaceHub router URL using environment variable 1549d93
HARI PREETHAM commited on
Update HuggingFace API Endpoint to router.huggingface.co cfba20f
HARI PREETHAM commited on
Fix HuggingFaceEndpoint validation error 74eac38
HARI PREETHAM commited on
Fix FAISS deserialization error d5280c7
HARI PREETHAM commited on
Deploying RAG backend 6ac0077
HARI PREETHAM commited on