How is interactive Q&A / chat implemented?
Chat over the repo; deep-research mode; streaming; conversation memory; MCP exposure.
Verdict
deepwiki-open serves chat over a WebSocket at /ws/chat, with an HTTP streaming fallback. Each turn runs the same research_chat() RAG path as page generation: FAISS top-20 chunks plus the history the client sends, wrapped in <turn> blocks. The server stores no history. Deep Research is a five-step prompt sequence (plan, updates, conclusion) that the frontend drives turn by turn. A separate Codemap mode builds a step-by-step code guide in two LLM calls and streams it as NDJSON. There is no MCP server.
OpenDeepWiki serves chat at POST /api/v1/chat/stream as Server-Sent Events, with typed content, thinking, tool_call, tool_result and done events. The agent reads the generated pages first (ReadDoc) and then the source (ReadFile, Grep, ListFiles). It can also call configured MCP servers and skills. Sessions and messages are saved in the database. There is no separate deep-research loop; the system prompt asks for step-by-step reasoning instead. It also runs an MCP server. The global /api/mcp endpoint lists repositories, routes a question to them by token scoring, and searches or reads generated pages. The per-repo endpoint /api/mcp/{owner}/{repo} offers SearchDoc, GetRepoStructure and ReadFile, but its SearchDoc uses a literal SQL LIKE.
For a chat that you can use in an IDE agent, or that must read actual source files, OpenDeepWiki is the stronger choice. For quick similarity-based answers and the guided Deep Research and Codemap modes in a single-user setup, use deepwiki-open.
Per-project answers
AsyncFuncAI/deepwiki-open
answeredInteractive Q&A over a repository is implemented through a WebSocket-based chat system. The frontend's Ask component (src/components/Ask.tsx) provides the UI, with three modes: Fast, Deep Research, and Codemap. When a user submits a question, the frontend opens a WebSocket to /ws/chat (src/utils/websocketClient.ts, line 64), sends a ChatCompletionRequest as JSON, and receives streaming text chunks. On WebSocket failure, it falls back to HTTP POST at /api/chat/stream (src/app/api/chat/stream/route.ts).
The backend endpoint is handle_websocket_chat() in api/routers/chat.py (line 20). It delegates to research_chat() (api/services/research.py), which: (1) prepares or loads the FAISS RAG index for the repo, (2) retrieves relevant document chunks, (3) loads conversation history from the Memory class into XML-style <turn> blocks, (4) picks the appropriate iteration prompt, and (5) streams the LLM response via ChatStreamer.respond_stream(). If the prompt exceeds 7500 tokens, RAG context is dropped and the model answers from its training data alone (line 146-148).
Deep Research mode (api/prompts.py lines 60-151) runs up to 5 automated iterations: iteration 1 produces a research plan, iterations 2-4 investigate deeper aspects, and iteration 5 produces a final conclusion. Each iteration builds on the previous conversation history. The frontend's Ask.tsx auto-continues through iterations (line 620-637), extracting research stages from the response (## Research Plan, ## Research Update N, ## Final Conclusion) and displaying a stage navigator. A max of 5 iterations is hard-coded (research.py line 219), with force-complete logic on the frontend (Ask.tsx line 501-508).
Codemap mode (api/services/codemap.py) is a separate multi-phase pipeline: RAG retrieval → two-step JSON generation (skeleton + enrich with Mermaid diagrams) → citation grounding against real source files → NDJSON streamed over /ws/codemap. Results are displayed as structured step-by-step guides with a side-by-side CodeViewer for cited source files.
Conversation history is maintained client-side in the conversationTurns array (Ask.tsx line 92), with model/provider selection exposed through a modal. Unlike the wiki generation pipeline, Q&A does not cache results on the server.
AIDotNet/OpenDeepWiki
answeredChat over the repo. The chat assistant is served through ChatAssistantEndpoints (src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs:14). The primary endpoint is POST /api/v1/chat/stream (line 150), which accepts a ChatRequest containing message history, model ID, and a DocContextDto (owner, repo, branch, language, current document path, catalog menu). The response is Server-Sent Events (SSE) with text/event-stream content type and X-Accel-Buffering: no to disable nginx buffering.
Streaming implementation. ChatAssistantService.StreamChatAsync() (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372) creates an agent session via Microsoft.Agents.AI. It uses agent.RunStreamingAsync() to yield token-by-token SSE events of types: content (text chunks), thinking (reasoning tokens from models that support it, e.g. DeepSeek R1), tool_call (with incremental or complete arguments), tool_result, done (final token counts), and error. The tool-call handling supports both OpenAI streaming format (StreamingChatCompletionUpdate with ToolCallUpdates, line 590–664) and Anthropic streaming format (RawMessageStreamEvent with content_block_start/delta/stop events, line 666–803), including thinking_delta and input_json_delta for extended thinking models.
Agent tools. The chat agent receives ReadDoc (line 503) for reading generated wiki pages with line ranges, ReadFile/ListFiles/Grep from GitTool (line 479) for source code access when the repository clone is available, plus configurable MCP tools and Skill tools. The system prompt (line 963–1138) instructs a structured workflow: Understand → Gather (using tools) → Analyze → Respond, with source-code verification required when docs are unclear.
Conversation memory. Session state is tracked in the database via ChatSessionImpl and SessionManager (src/OpenDeepWiki/Chat/Sessions/). Messages are persisted through a DatabaseMessageQueue with dead-letter handling. Usage accounting records token consumption via AiUsageAccounting per session.
Deep-research mode. The system prompt includes internal_thinking instructions (line 998–1022) that encourage the agent to reason step-by-step before responding, using tags for structured deliberation.
MCP exposure. The chat is also available through MCP endpoints (/api/mcp and /api/mcp/{owner}/{repo}, per README), where the same agent tools are exposed via the MCP protocol. Repository-scoped MCP endpoints bind the tool context to a specific repo. The McpToolConverter (src/OpenDeepWiki/Services/Chat/McpToolConverter.cs) converts MCP provider configs into AI tools at request time.