Utilities

The nse._utils module contains internal utility functions used by the NSE library. These functions support common operations such as filesystem path handling, file downloads, archive extraction, and other internal processing.

The module is not part of the public NSE API and is not intended to be used directly by library users. It is documented here as a reference for developers and contributors working on the NSE library.

nse._utils.consume_archive(file: Path, folder: Path, extract_files: List[str] | None = None) → Path

Extract a .zip or .gz archive and delete the original file.

For .zip files:
  • If extract_files is provided, all listed members are extracted, and the path to the last member in the list is returned.

  • Otherwise, the first member in the archive’s name list is extracted and its path returned.

For .gz files: the decompressed contents are written to a sibling file whose name is the archive’s stem (foo.csv.gz → foo.csv). Only single-stream .gz files are supported; .tar.gz archives are decompressed to a .tar file rather than extracted.

Parameters:
  • file (Path) – Path to the archive to extract.

  • folder (Path) – Directory into which contents are extracted.

  • extract_files (Optional[List[str]]) – Optional list of member names to extract from a zip archive. If None, the first member is extracted. Must be non-empty if provided. Defaults to None.

Returns:

Path to the extracted (or decompressed) file.

Return type:

Path

Raises:
  • ValueError – If file has a suffix other than .zip or .gz, or if extract_files is provided as an empty list.

  • KeyError – If a name in extract_files is not present in the zip archive.

  • zipfile.BadZipFile – If the file is not a valid zip archive.

  • OSError – If file I/O fails during extraction or decompression.

Note

The original archive at file is deleted after a successful extraction. If deletion fails — for example due to a permission error or, on Windows, because the file is still held open by another process — the failure is logged and the extracted file is still returned. A FileNotFoundError (the archive was already removed) is silently ignored.

Note

If extraction fails partway through, files already written to folder are not rolled back, and the original archive is left in place.

nse._utils.prepare_path(path: str | Path, is_folder: bool = False)

Resolve and validate a filesystem path.

Expands ~, converts to an absolute path, and — if is_folder is True — ensures the path is a directory, creating it (including parents) if it does not exist.

Parameters:
  • path (Union[str, Path]) – The path to resolve. Strings are converted to Path.

  • is_folder (bool) – Default False. If True, treat path as a directory and enforce/create it.

Returns:

The resolved absolute path.

Return type:

Path

Raises:

NotADirectoryError – If is_folder is True and path exists but does not point to a directory. This includes symlinks whose target is a regular file.

Note

When is_folder is False, the path is resolved but neither validated nor created.

Note

If is_folder is True and path does not exist, it is created with mkdir(parents=True, exist_ok=True). Concurrent callers racing to create the same directory will not raise.

Note

path is resolved with Path.resolve() before any checks, so symlinks are fully followed. A broken symlink is therefore indistinguishable from a non-existent path, and its target directory will be created when is_folder is True.

nse._utils.split_date_range(from_date: date, to_date: date, max_chunk_size: int = 365) → List[Tuple[date, date]]

Split a date range into non-overlapping, inclusive chunks.

Each chunk spans at most max_chunk_size days (inclusive of both endpoints). The next chunk begins one day after the previous chunk’s end.

Parameters:
  • from_date (datetime.date) – The starting date of the range (inclusive).

  • to_date (datetime.date) – The ending date of the range (inclusive).

  • max_chunk_size (int) – Default 365. Maximum number of days in each chunk, counted inclusively (max_chunk_size=1 yields one-day chunks). Must be positive.

Returns:

A list of (start_date, end_date) tuples, ordered chronologically. Each tuple is inclusive of both endpoints and consecutive tuples do not overlap.

Return type:

List[Tuple[datetime.date, datetime.date]]

Raises:

ValueError – If max_chunk_size is less than or equal to 0.

Note

If from_date > to_date, an empty list is returned. No exception is raised.