ALDashboard.translation
- ALDashboard.translation
- hashlib
- json
- math
- os
- re
- tempfile
- ThreadPoolExecutor
- lru_cache
- Lock
- Any
- Callable
- List
- Optional
- Tuple
- Union
- Literal
- cast
- ET
- zipfile
- config
- in_celery
- setup
- astparser
- DAError
- functions
- word
- DA
- interview_cache
- parse
- pdftk
- util
- core
- setup_translation
- send_file
- redirect
- flash
- pandas
- xlsxwriter
- DAFile
- language_name
- get_config
- log
- DAEmpty
- NamedTuple
- Dict
- chat_completion
- get_default_model
- StableValue
- stable_values_to_preserve
- DEFAULT_MAX_FRAGMENTS_PER_BATCH
- alignment_problems
- estimate_cost
- estimate_tokens
- plan_batches
- tiktoken
- runtime
- Lexer
- MAX_MAKO_RETRIES
- MAX_BATCH_SPLIT_DEPTH
- DEFAULT_MAX_PARALLEL_REQUESTS
- DEFAULT_TRANSLATION_MODEL
- TRANSLATION_MODEL_FALLBACK_CHAIN
- is_valid_mako_block
- small_model_for_fallback
- DEFAULT_LANGUAGE
- __all__
- gpt_is_available
- may_have_mako
- may_have_html
- translate_fragments_gpt
- Translation
- translation_file
hashlib
json
math
os
re
tempfile
ThreadPoolExecutor
lru_cache
Lock
Any
Callable
List
Optional
Tuple
Union
Literal
cast
ET
zipfile
config
in_celery
setup
astparser
DAError
functions
word
DA
interview_cache
parse
pdftk
util
core
setup_translation
send_file
redirect
flash
pandas
xlsxwriter
DAFile
language_name
get_config
log
DAEmpty
NamedTuple
Dict
chat_completion
get_default_model
StableValue
stable_values_to_preserve
DEFAULT_MAX_FRAGMENTS_PER_BATCH
alignment_problems
estimate_cost
estimate_tokens
plan_batches
tiktoken
runtime
Lexer
MAX_MAKO_RETRIES
MAX_BATCH_SPLIT_DEPTH
DEFAULT_MAX_PARALLEL_REQUESTS
DEFAULT_TRANSLATION_MODEL
TRANSLATION_MODEL_FALLBACK_CHAIN
is_valid_mako_block
def is_valid_mako_block(text: str) -> Tuple[bool, Optional[str]]
Return True if the provided text parses as Mako without raising an error. Empty strings are treated as valid.
This lexes the template rather than rendering it. Rendering runs whatever
Python the draft translation happens to contain, and it runs once per
drafted fragment, so lexing is both the safe choice and much the faster one.
translation_validation checks translation files the same way.
small_model_for_fallback
@lru_cache(maxsize=1)
def small_model_for_fallback() -> Optional[str]
The small model this server's provider offers, or None if it cannot say.
The named fallback chain is a list of OpenAI models, which is no use to a server pointed at some other endpoint. ALToolbox resolves a "small" model from the docassemble configuration, then the configured model sets, then the endpoint's own model list, so it is the one fallback that does not assume who the provider is.
Cached: resolving it can call the models endpoint, and this is consulted once per fragment that needs a retry. A server changing providers mid-process is not a case worth re-querying for.
DEFAULT_LANGUAGE
__all__
gpt_is_available
def gpt_is_available() -> bool
Return True if the GPT API is available.
may_have_mako
def may_have_mako(text: str) -> bool
Return True if the text appears to contain any Mako code, such as ${...} or % at the beginning of a line.
may_have_html
def may_have_html(text: str) -> bool
Return True if the text appears to contain any HTML code, such as <p> or <div>.
translate_fragments_gpt
def translate_fragments_gpt(
fragments: Union[str, List[Tuple[int, str]]],
source_language: str,
tr_lang: str,
interview_context: Optional[str] = None,
special_words: Optional[Dict[int, str]] = None,
model: Optional[str] = DEFAULT_TRANSLATION_MODEL,
openai_base_url: Optional[str] = None,
max_output_tokens: Optional[int] = None,
max_input_tokens: Optional[int] = None,
openai_api: Optional[str] = None,
reasoning_effort: Optional[Literal["minimal", "low", "medium",
"high"]] = "low",
max_fragments_per_batch: int = DEFAULT_MAX_FRAGMENTS_PER_BATCH,
max_parallel_requests: int = DEFAULT_MAX_PARALLEL_REQUESTS,
request_cost_callback: Optional[Callable[[float], None]] = None
) -> Dict[Union[int, str], str]
Use an AI model to translate a list of fragments (strings) from one language to another and provide a dictionary with the original text and the translated text.
You can optionally provide an alternative model, but it must support JSON mode.
Arguments
fragments- A list of strings to be translated.source_language- The language of the original text.tr_lang- The language to translate the text into.special_words- A dictionary of special words that should be translated in a specific way.model- The GPT model to use. Defaults to DEFAULT_TRANSLATION_MODEL.openai_base_url- The base URL for the OpenAI API. If not provided, the default OpenAI URL will be used.max_output_tokens- The maximum number of tokens to generate in the output.max_input_tokens- The maximum number of tokens in the input. If not provided, it will be set to 4000.openai_api- The OpenAI API key. If not provided, it will use the key from the configuration.reasoning_effort- Controls the reasoning effort for thinking models like GPT-5. Defaults to "low".max_fragments_per_batch- How many fragments to translate per request. One request per fragment is slow; a whole interview in one request comes back misaligned.max_parallel_requests- How many requests to keep in flight at once.request_cost_callback- Optional callback receiving the estimated cost of each completed API request, including retries.
Returns
A dictionary where the keys are the indices of the fragments and the values are the translated text.
Translation Objects
class Translation(NamedTuple)
file: DAFile
an XLSX or XLIFF file
untranslated_words: `(
int # Word count for all untranslated segments that are not Mako or HTML )`
untranslated_segments: int
Number of rows in the output that have untranslated text - one for each question, subquestion, field, etc.
total_rows: int
preserved_values: int
undrafted_segments: int
estimated_cost_usd: float
Rough API cost of the AI drafts, if any
translation_file
def translation_file(yaml_filename: str,
tr_lang: str,
use_gpt=False,
use_google_translate=False,
openai_api: Optional[str] = None,
max_tokens=4000,
interview_context: Optional[str] = None,
special_words: Optional[Dict[int, str]] = None,
model: Optional[str] = None,
openai_base_url: Optional[str] = None,
max_input_tokens: Optional[int] = None,
max_output_tokens: Optional[int] = None,
reasoning_effort: Optional[Literal["minimal", "low",
"medium",
"high"]] = None,
validate_mako: Optional[bool] = True) -> Translation
Return a tuple of the translation file in XLSX format, plus a count of the number of words and segments that need to be translated.
The word and segment count only apply when filetype="XLSX".
This code was adjusted from the Flask endpoint-only version in server.py. XLIFF support was removed for now but can be added later.
Arguments
yaml_filename- Fully qualified interview YAML path.tr_lang- Target translation language (ISO code).use_gpt- Whether to include GPT draft translations.use_google_translate- Placeholder for legacy Google support.openai_api- API key override.max_tokens- Legacy max token setting (kept for backward compatibility).interview_context- Optional context prompt to send with GPT calls.special_words- Optional glossary to enforce terminology.model- Preferred OpenAI model.openai_base_url- Override the OpenAI base URL.max_input_tokens- Optional override for input token limits.max_output_tokens- Optional override for completion token limits.reasoning_effort- Reasoning effort setting, used for GPT-5 models.validate_mako- When True, retry GPT translations that break Mako syntax (default).