# How is interactive Q&A / chat implemented?

> Open-source DeepWiki — a good answer covers: Chat over the repo; deep-research mode; streaming; conversation memory; MCP exposure.

Canonical page: https://llms-technical-reviews.com/open-source-deepwiki/q/qa/

## Verdict

[deepwiki-open](/p/deepwiki-open/) serves chat over a WebSocket at `/ws/chat`, with an HTTP streaming fallback. Each turn runs the same `research_chat()` RAG path as page generation: FAISS top-20 chunks plus the history the client sends, wrapped in `<turn>` blocks. The server stores no history. Deep Research is a five-step prompt sequence (plan, updates, conclusion) that the frontend drives turn by turn. A separate Codemap mode builds a step-by-step code guide in two LLM calls and streams it as NDJSON. There is no MCP server.

[OpenDeepWiki](/p/opendeepwiki/) serves chat at `POST /api/v1/chat/stream` as Server-Sent Events, with typed `content`, `thinking`, `tool_call`, `tool_result` and `done` events. The agent reads the generated pages first (`ReadDoc`) and then the source (`ReadFile`, `Grep`, `ListFiles`). It can also call configured MCP servers and skills. Sessions and messages are saved in the database. There is no separate deep-research loop; the system prompt asks for step-by-step reasoning instead. It also runs an MCP server. The global `/api/mcp` endpoint lists repositories, routes a question to them by token scoring, and searches or reads generated pages. The per-repo endpoint `/api/mcp/{owner}/{repo}` offers `SearchDoc`, `GetRepoStructure` and `ReadFile`, but its `SearchDoc` uses a literal SQL `LIKE`.

For a chat that you can use in an IDE agent, or that must read actual source files, OpenDeepWiki is the stronger choice. For quick similarity-based answers and the guided Deep Research and Codemap modes in a single-user setup, use deepwiki-open.

## Per-project answers

### AsyncFuncAI/deepwiki-open (answered)

Interactive Q&A over a repository is implemented through a WebSocket-based chat system. The frontend's `Ask` component (`src/components/Ask.tsx`) provides the UI, with three modes: Fast, Deep Research, and Codemap. When a user submits a question, the frontend opens a WebSocket to `/ws/chat` (`src/utils/websocketClient.ts`, line 64), sends a `ChatCompletionRequest` as JSON, and receives streaming text chunks. On WebSocket failure, it falls back to HTTP POST at `/api/chat/stream` (`src/app/api/chat/stream/route.ts`).

The backend endpoint is `handle_websocket_chat()` in `api/routers/chat.py` (line 20). It delegates to `research_chat()` (`api/services/research.py`), which: (1) prepares or loads the FAISS RAG index for the repo, (2) retrieves relevant document chunks, (3) loads conversation history from the `Memory` class into XML-style `<turn>` blocks, (4) picks the appropriate iteration prompt, and (5) streams the LLM response via `ChatStreamer.respond_stream()`. If the prompt exceeds 7500 tokens, RAG context is dropped and the model answers from its training data alone (line 146-148).

Deep Research mode (`api/prompts.py` lines 60-151) runs up to 5 automated iterations: iteration 1 produces a research plan, iterations 2-4 investigate deeper aspects, and iteration 5 produces a final conclusion. Each iteration builds on the previous conversation history. The frontend's `Ask.tsx` auto-continues through iterations (line 620-637), extracting research stages from the response (`## Research Plan`, `## Research Update N`, `## Final Conclusion`) and displaying a stage navigator. A max of 5 iterations is hard-coded (research.py line 219), with force-complete logic on the frontend (Ask.tsx line 501-508).

Codemap mode (`api/services/codemap.py`) is a separate multi-phase pipeline: RAG retrieval → two-step JSON generation (skeleton + enrich with Mermaid diagrams) → citation grounding against real source files → NDJSON streamed over `/ws/codemap`. Results are displayed as structured step-by-step guides with a side-by-side CodeViewer for cited source files.

Conversation history is maintained client-side in the `conversationTurns` array (Ask.tsx line 92), with model/provider selection exposed through a modal. Unlike the wiki generation pipeline, Q&A does not cache results on the server.


Citations: [api/services/research.py:60-311](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L60-L311) · [src/components/Ask.tsx:110-144](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/src/components/Ask.tsx#L110-L144) · [src/utils/websocketClient.ts:64-96](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/src/utils/websocketClient.ts#L64-L96) · [api/routers/chat.py:20-66](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/chat.py#L20-L66) · [api/prompts.py:60-151](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/prompts.py#L60-L151) · [api/services/codemap.py:211-298](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/codemap.py#L211-L298)

### AIDotNet/OpenDeepWiki (answered)

**Chat over the repo.** The chat assistant is served through `ChatAssistantEndpoints` (src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs:14). The primary endpoint is `POST /api/v1/chat/stream` (line 150), which accepts a `ChatRequest` containing message history, model ID, and a `DocContextDto` (owner, repo, branch, language, current document path, catalog menu). The response is Server-Sent Events (SSE) with `text/event-stream` content type and `X-Accel-Buffering: no` to disable nginx buffering.

**Streaming implementation.** `ChatAssistantService.StreamChatAsync()` (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372) creates an agent session via `Microsoft.Agents.AI`. It uses `agent.RunStreamingAsync()` to yield token-by-token SSE events of types: `content` (text chunks), `thinking` (reasoning tokens from models that support it, e.g. DeepSeek R1), `tool_call` (with incremental or complete arguments), `tool_result`, `done` (final token counts), and `error`. The tool-call handling supports both **OpenAI streaming format** (`StreamingChatCompletionUpdate` with `ToolCallUpdates`, line 590–664) and **Anthropic streaming format** (`RawMessageStreamEvent` with `content_block_start/delta/stop` events, line 666–803), including `thinking_delta` and `input_json_delta` for extended thinking models.

**Agent tools.** The chat agent receives `ReadDoc` (line 503) for reading generated wiki pages with line ranges, `ReadFile`/`ListFiles`/`Grep` from `GitTool` (line 479) for source code access when the repository clone is available, plus configurable MCP tools and Skill tools. The system prompt (line 963–1138) instructs a structured workflow: Understand → Gather (using tools) → Analyze → Respond, with source-code verification required when docs are unclear.

**Conversation memory.** Session state is tracked in the database via `ChatSessionImpl` and `SessionManager` (src/OpenDeepWiki/Chat/Sessions/). Messages are persisted through a `DatabaseMessageQueue` with dead-letter handling. Usage accounting records token consumption via `AiUsageAccounting` per session.

**Deep-research mode.** The system prompt includes `internal_thinking` instructions (line 998–1022) that encourage the agent to reason step-by-step before responding, using  tags for structured deliberation.

**MCP exposure.** The chat is also available through MCP endpoints (`/api/mcp` and `/api/mcp/{owner}/{repo}`, per README), where the same agent tools are exposed via the MCP protocol. Repository-scoped MCP endpoints bind the tool context to a specific repo. The `McpToolConverter` (src/OpenDeepWiki/Services/Chat/McpToolConverter.cs) converts MCP provider configs into AI tools at request time.


Citations: [src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs:1-245](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs#L1-L245) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:186-300](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L186-L300) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:463-827](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L463-L827) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:963-1138](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L963-L1138)
