Anthropic ClaudeLLM・AI開発⭐ リポ 0品質スコア 50/100

llamaindex-development

Name: llamaindex-development
Author: mindrally

LlamaIndexを用いたRAGアプリケーション、ベクターストア、ドキュメント処理、クエリエンジンの構築から本番AIアプリケーションの開発まで、専門的なガイダンスを提供します。

description の原文を見る

Expert guidance for LlamaIndex development including RAG applications, vector stores, document processing, query engines, and building production AI applications.

SKILL.md 本文

LlamaIndex Development

LlamaIndexを使用したRAG(Retrieval-Augmented Generation)アプリケーション、データインデックス、LLM駆動アプリケーションのPython開発の専門家です。

主要な原則

簡潔で技術的な応答と正確なPythonの例を提供する
関数型・宣言型プログラミングを使用し、可能な限りクラスを避ける
コード品質、保守性、パフォーマンスを優先する
目的を反映した説明的な変数名を使用する
PEP 8スタイルガイドに従う

コード構成

ディレクトリ構造

project/
├── data/                 # Source documents and data
├── indexes/              # Persisted index storage
├── loaders/              # Custom document loaders
├── retrievers/           # Custom retriever implementations
├── query_engines/        # Query engine configurations
├── prompts/              # Custom prompt templates
├── transformations/      # Document transformations
├── callbacks/            # Custom callback handlers
├── utils/                # Utility functions
├── tests/                # Test files
└── config/               # Configuration files

命名規約

ファイル、関数、変数にはsnake_caseを使用
クラスにはPascalCaseを使用
プライベート関数にはアンダースコアをプレフィックスする
説明的な名前を使用する(例: create_vector_index, build_query_engine)

ドキュメント読み込み

Document Loadersの使用

from llama_index.core import SimpleDirectoryReader
from llama_index.readers.file import PDFReader, DocxReader

# Load from directory
documents = SimpleDirectoryReader(
    input_dir="./data",
    recursive=True,
    required_exts=[".pdf", ".txt", ".md"]
).load_data()

# Load specific file types
pdf_reader = PDFReader()
documents = pdf_reader.load_data(file="document.pdf")

カスタムローダー

from llama_index.core.readers.base import BaseReader
from llama_index.core import Document

class CustomLoader(BaseReader):
    def load_data(self, file_path: str) -> list[Document]:
        # Custom loading logic
        with open(file_path, 'r') as f:
            content = f.read()

        return [Document(
            text=content,
            metadata={"source": file_path}
        )]

テキスト分割と処理

Node Parsing

from llama_index.core.node_parser import (
    SentenceSplitter,
    SemanticSplitterNodeParser,
    MarkdownNodeParser
)

# Simple sentence splitting
splitter = SentenceSplitter(
    chunk_size=1024,
    chunk_overlap=200
)
nodes = splitter.get_nodes_from_documents(documents)

# Semantic splitting (preserves meaning)
from llama_index.embeddings.openai import OpenAIEmbedding

semantic_splitter = SemanticSplitterNodeParser(
    embed_model=OpenAIEmbedding(),
    breakpoint_percentile_threshold=95
)

# Markdown-aware splitting
markdown_splitter = MarkdownNodeParser()

チャンキングのベストプラクティス

埋め込みモデルのコンテキストウィンドウに基づいてチャンクサイズを選択する
チャンク間のコンテキスト保持のためにオーバーラップを使用する
可能な限りドキュメント構造を保持する
フィルタリングと検索のためにメタデータを含める
より良い一貫性のためにセマンティック分割を使用する

Vector Storesとインデックス

インデックスの作成

from llama_index.core import VectorStoreIndex, StorageContext
from llama_index.vector_stores.chroma import ChromaVectorStore
import chromadb

# In-memory index
index = VectorStoreIndex.from_documents(documents)

# With persistent vector store
chroma_client = chromadb.PersistentClient(path="./chroma_db")
chroma_collection = chroma_client.get_or_create_collection("my_collection")

vector_store = ChromaVectorStore(chroma_collection=chroma_collection)
storage_context = StorageContext.from_defaults(vector_store=vector_store)

index = VectorStoreIndex.from_documents(
    documents,
    storage_context=storage_context
)

サポートされているVector Stores

Chroma(ローカル開発)
Pinecone(本番環境、マネージド)
Weaviate(本番環境、自ホスト型またはマネージド)
Qdrant(本番環境、自ホスト型またはマネージド)
PostgreSQL with pgvector
MongoDB Atlas Vector Search

インデックスの永続化

from llama_index.core import StorageContext, load_index_from_storage

# Persist index
index.storage_context.persist(persist_dir="./storage")

# Load index
storage_context = StorageContext.from_defaults(persist_dir="./storage")
index = load_index_from_storage(storage_context)

Query Engines

基本的なQuery Engine

from llama_index.core import VectorStoreIndex

index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine(
    similarity_top_k=5,
    response_mode="compact"
)

response = query_engine.query("What is the main topic?")
print(response.response)

Response Modes

refine: 各ノードを通じて段階的に回答を改善する
compact: LLMに送信する前にチャンクを結合する
tree_summarize: ツリーを構築して要約する
simple_summarize: 切り詰めて要約する
accumulate: 各ノードからの応答を蓄積する

高度なQuery Engine

from llama_index.core.query_engine import RetrieverQueryEngine
from llama_index.core.postprocessor import SimilarityPostprocessor

query_engine = RetrieverQueryEngine.from_args(
    retriever=index.as_retriever(similarity_top_k=10),
    node_postprocessors=[
        SimilarityPostprocessor(similarity_cutoff=0.7)
    ],
    response_mode="compact"
)

Retrievers

カスタムRetrievers

from llama_index.core.retrievers import VectorIndexRetriever

# Basic retriever
retriever = VectorIndexRetriever(
    index=index,
    similarity_top_k=10
)

# Retrieve nodes
nodes = retriever.retrieve("search query")

ハイブリッド検索

from llama_index.core.retrievers import QueryFusionRetriever

# Combine multiple retrieval strategies
retriever = QueryFusionRetriever(
    [
        index.as_retriever(similarity_top_k=5),
        bm25_retriever,  # Keyword-based
    ],
    num_queries=4,
    use_async=True
)

埋め込み

埋め込みモデル

from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.core import Settings

# OpenAI embeddings
Settings.embed_model = OpenAIEmbedding(
    model="text-embedding-3-small",
    dimensions=512  # Optional dimension reduction
)

# Local embeddings
Settings.embed_model = HuggingFaceEmbedding(
    model_name="BAAI/bge-small-en-v1.5"
)

LLM構成

LLMのセットアップ

from llama_index.llms.openai import OpenAI
from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings

# OpenAI
Settings.llm = OpenAI(
    model="gpt-4o",
    temperature=0.1
)

# Anthropic
Settings.llm = Anthropic(
    model="claude-sonnet-4-20250514",
    temperature=0.1
)

Agents

Agentsの構築

from llama_index.core.agent import ReActAgent
from llama_index.core.tools import QueryEngineTool, ToolMetadata

# Create tools from query engines
tools = [
    QueryEngineTool(
        query_engine=documents_query_engine,
        metadata=ToolMetadata(
            name="documents",
            description="Search through documents"
        )
    ),
    QueryEngineTool(
        query_engine=code_query_engine,
        metadata=ToolMetadata(
            name="codebase",
            description="Search through code"
        )
    )
]

# Create agent
agent = ReActAgent.from_tools(
    tools,
    llm=llm,
    verbose=True
)

response = agent.chat("Find information about X")

パフォーマンス最適化

キャッシング

from llama_index.core import Settings
from llama_index.core.llms import LLMCache

# Enable LLM response caching
Settings.llm = OpenAI(model="gpt-4o")
Settings.llm_cache = LLMCache()

非同期操作

# Use async for better performance
response = await query_engine.aquery("question")

# Batch processing
responses = await asyncio.gather(*[
    query_engine.aquery(q) for q in questions
])

埋め込み最適化

可能な場合は埋め込みをバッチ処理する
精度が許す限り、より小さい埋め込みディメンションを使用する
繰り返されるドキュメントの埋め込みをキャッシュする
コスト対応アプリケーションにはローカルモデルを使用する

エラーハンドリング

from llama_index.core.callbacks import CallbackManager, LlamaDebugHandler

# Debug handler for troubleshooting
debug_handler = LlamaDebugHandler()
callback_manager = CallbackManager([debug_handler])

Settings.callback_manager = callback_manager

テスト

ドキュメントローダーと変換のユニットテストを実施する
既知のクエリで検索品質をテストする
インデックスの永続化とロードを検証する
Query Engineのレスポンスをテストする
検索メトリクス(精度、再現率)を監視する

依存関係

llama-index
llama-index-embeddings-openai
llama-index-llms-openai
llama-index-vector-stores-chroma
chromadb
python-dotenv
pydantic

ライセンス: Apache-2.0(寛容ライセンスのため全文を引用しています) · 原本リポジトリ

詳細情報

作者: mindrally
リポジトリ: mindrally/skills
ライセンス: Apache-2.0
最終更新: 不明

GitHubで原本を見る →フィードバックを送る

Source: https://github.com/mindrally/skills / ライセンス: Apache-2.0