Elasticsearch Index Design
Design explicit index mappings from access patterns, review existing mappings for type and storage mistakes, and apply corrections through a new index plus reindex when field types must change.
<!-- begin-partial: preamble -->Environment Configuration
This skill executes Elasticsearch operations through the elastic CLI. If the
elastic CLI is not installed, tell the user what it is needed for. Do
not guess credentials, call the HTTP API directly, or attempt other workarounds.
This skill references operations in HTTP-shorthand form (e.g., GET /, GET /_cat/indices, GET /{index}/_mapping,
GET /{index}/_settings/index.mode, POST /_query). The Operations table at the end of this document
maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API
directly.
Process
-
Gather access patterns per field. Before choosing types, list how each field is used. For every field capture:
- Search — full-text match, phrase, relevance scoring?
- Filter — exact term, terms set, prefix?
- Aggregate — terms, cardinality, histogram, stats?
- Sort — ascending/d descending in result sets?
- Retrieve only — returned in
_sourcebut never queried?
The decision: classify each field into one primary access pattern (search, exact, numeric metric, date, boolean, structured object, or retrieve-only). Missing access-pattern data is a blocker — ask the user rather than guessing. Call
GET /to confirm connectivity; when reviewing an existing index, callGET /{index}/_mappingto ground the discussion in the current mapping. -
Choose field types from access patterns. Map each field to the minimal type set that satisfies its pattern. Read Field Type Decisions and Multi-Field Patterns before proposing mappings.
Key judgments:
Pattern Mapping Full-text search only text(no keyword sub-field)Filter / agg / sort only keyword(nottext)Full-text search and sort or aggregation textwithfields.keywordmulti-fieldDecimal price or metric double,float, orscaled_float— nottextor integerTimestamp dateTrue/false flag booleanFree-form key/value map with many distinct keys flattened— not dynamicobjectMulti-field rule: When a field must be searchable and sortable/aggregatable (e.g. product
name), map it astextwith akeywordsub-field — search onname, sort and aggregate onname.keyword. Mapping as onlytextor onlykeywordis wrong for that combined pattern.Explicit mapping rule: For new indices, always define mappings explicitly with
PUT /{index}. Do not rely on dynamic mapping for production indices — the first document can lock in wrong types (strings astext, ambiguous numbers askeyword).Index settings: Set deliberate
number_of_shardsandnumber_of_replicasin the samePUT /{index}request when the deployment allows it (Self-Managed / Elastic Cloud Hosted). On Serverless, omit shard and replica counts (Elastic manages them); still supply explicit mappings. State chosen values or document that defaults apply.Example —
productsindex optimized for search plus sort/agg on name:{ "settings": { "number_of_shards": 1, "number_of_replicas": 1 }, "mappings": { "properties": { "name": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } }, "price": { "type": "double" }, "created": { "type": "date" }, "in_stock": { "type": "boolean" } } } }Create with
PUT /productspassing thesettingsandmappingsblocks. Verify withGET /products/_mapping. -
Guard against mapping explosion and storage bloat. On high-volume indices, type mistakes multiply cost. Read Mapping Explosion and Storage Bloat and apply these review checks:
- Analyzed-but-not-searched fields — Fields used only for filter and aggregation (
url, HTTPstatus_code,tags, IDs) must bekeyword, nottext.textwastes space; aggregations ontextrequire fielddata or a.keywordsub-field that should not exist if the field is not searched. message.keywordwithoutignore_above— A keyword sub-field on a large full-text body indexes the entire raw string as one term. Flag this anti-pattern; remove the sub-field when only full-text search is needed, or addignore_abovewhen a bounded exact-match sub-field is truly required.- Dynamic free-form objects —
objectwith"dynamic": trueon user-supplied key/value data with thousands of distinct keys causes mapping explosion. Recommendflattened(or strict dynamic / allowlist strategy). doc_values: false— On fields retrieved in hits but never sorted, aggregated, or filtered (e.g. display-onlysession_id), set"doc_values": falseonkeywordto save disk at scale.scaled_float— For metrics with bounded precision (e.g.response_time_ms), preferscaled_floatwith an appropriatescaling_factorover plainfloat/doublewhen storage dominates.
Prefer
"dynamic": "strict"on the root mapping unless unknown fields are an explicit requirement. - Analyzed-but-not-searched fields — Fields used only for filter and aggregation (
-
Apply design: create new index and reindex when types change. Elasticsearch cannot change an existing field's type in place. When review finds wrong types (text→keyword, object→flattened, float→scaled_float, doc_values changes on existing fields), state clearly that fixes require a new index and reindex — not a mapping update on the live index.
Workflow for correcting an existing high-volume index such as
events:- Design the corrected mapping on a new index name (e.g.
events-v2) incorporating all fixes from steps 2–3. - Create the destination with
PUT /events-v2and the full correctedmappings(andsettingswhere applicable). - Copy documents with
POST /_reindex— for large indices usewait_for_completion=falseand track the task. Source:{ "index": "events" }, destination:{ "index": "events-v2" }. - Verify with
GET /events-v2/_count(compare to source count) andGET /events-v2/_mapping(confirm types). - Cut over reads and writes (index alias swap or application config) after validation.
Example corrected excerpt for the
eventsreview pattern:{ "mappings": { "properties": { "@timestamp": { "type": "date" }, "event_id": { "type": "keyword" }, "session_id": { "type": "keyword", "doc_values": false }, "url": { "type": "keyword" }, "status_code": { "type": "keyword" }, "response_time_ms": { "type": "scaled_float", "scaling_factor": 100 }, "tags": { "type": "keyword" }, "message": { "type": "text" }, "labels": { "type": "flattened" } } } }Do not attempt in-place mapping fixes for these type changes — they are rejected or leave data inconsistent. For greenfield indices, a single
PUT /{index}before first ingest avoids reindex entirely. - Design the corrected mapping on a new index name (e.g.
Review checklist
When the user supplies a mapping JSON and usage notes, walk this checklist in order:
- Match each field's type to its stated access pattern (see step 2).
- Flag
texton filter/agg-only fields; flag missing multi-fields where search and sort/agg share one logical field. - Flag
message.keyword(or similar) withoutignore_aboveon large analyzed text. - Flag dynamic
objecton high-cardinality free-form maps; recommendflattened. - Propose retrieve-only and numeric storage optimizations (
doc_values: false,scaled_float). - State that type changes require a new index and
POST /_reindex, then show the corrected mapping and reindex plan.
Examples
"Users search product names and also sort and aggregate on them" — one logical field, two access patterns, so use a
text field with a keyword multi-field:
{
"mappings": {
"properties": {
"product_name": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } }
}
}
}
"A status field is only ever filtered and aggregated, never full-text searched" — use keyword, not text:
{ "mappings": { "properties": { "status": { "type": "keyword" } } } }
"Free-form labels object with unbounded keys" — avoid mapping explosion with flattened:
{ "mappings": { "properties": { "labels": { "type": "flattened" } } } }
Guidelines
- Minimal mapping — Map only what access patterns require; every sub-field and analyzed form adds indexed data.
- Never guess access patterns — Wrong type choice is expensive to fix at scale.
- Verify after create — Always confirm with
GET /{index}/_mapping; useGET /{index}/_countafter reindex. - Cross-skill boundary — Copying documents between indices is
POST /_reindex(see the reindex skill for slicing, throttling, and task tracking). Loading files into a new index is bulk ingest, not index design.
Reference material
- Field Type Decisions — access-pattern-to-type table and common mistakes
- Multi-Field Patterns — text+keyword,
ignore_above, anti-patterns - Mapping Explosion and Storage Bloat —
flattened,doc_values, dynamic objects
Operations
| HTTP API (shorthand) | elastic CLI command |
|---|---|
GET / | elastic es info |
GET /{index}/_mapping | elastic es indices get-mapping --index '<index>' |
PUT /{index} | elastic es indices create --index '<index>' --mappings '<json>' --settings '<json>' |
POST /_reindex | elastic es reindex --source '<json>' --dest '<json>' |
POST /_reindex?wait_for_completion=false | elastic es reindex --wait-for-completion false --source '<json>' --dest '<json>' |
GET /{index}/_count | elastic es count --index '<index>' |