Introduce `collect_pattern_addrs` to gather all bound addresses from a
destructuring pattern. This enables the optimizer to remove unused
destructuring definitions, similar to how simple unused definitions are
handled. Added integration tests to verify this behavior and other
optimizer robustness scenarios.
The Binder and related types have been refactored to use the new
`NodeKind` enum instead of the previous `BoundKind`. This commit updates
all references to use the new structure, ensuring consistency across the
compiler's AST representation.
Key changes include:
- Replacing `BoundKind` with `NodeKind` in the Binder's `bind` and
`visit` methods.
- Updating pattern matching and field access to reflect the new enum
variants (e.g., `NodeKind::Def` instead of `BoundKind::Define`).
- Adjusting identifier bindings to use the new `IdentifierBinding` enum.
- Reflecting changes in `DefBinding`, `AssignBinding`, and
`LambdaBinding` structures.
- Ensuring all newly created nodes use `NodeKind` and the appropriate
metadata.
The `AssignBinding` struct now uses `Option<Address<L>>` to
differentiate between simple assignments and destructuring assignments.
For destructuring, the addresses are managed by the `Identifier` nodes
within the target pattern.
A new `GetField` node kind has been introduced to represent optimized
field access, which is generated during the lowering phase for the
`RuntimePhase`. This change aids in more efficient code generation for
accessing fields.
This commit introduces two new files:
- `.claudeignore`: Specifies files and directories that should be
ignored by Claude, such as snapshot files and the `target` directory.
- `docs/bytecode-architecture.md`: Documents the architecture for
bytecode generation, outlining the hybrid execution model, the
transformation pipeline (Specializer and Lowering), and the role of
Tail Call Optimization (TCO).
This commit introduces a new `NodeKind` enum that serves as a unified
representation for AST nodes across different compilation phases. It
replaces the separate `SyntaxKind` and `BoundKind` enums, allowing
phase-specific information to be attached through generic type
parameters.
This change simplifies the AST structure and makes it easier to manage
node information consistently throughout the compilation process. The
`NodeKind` enum now holds all possible AST node variants, with
phase-specific data encapsulated within `CompilerPhase` trait
implementations.
This commit reduces unnecessary cloning for fields like `addr`, `kind`,
and `field` within `BoundKind` and `VM::GetField`. This improves
performance by avoiding redundant memory allocations.
This commit refactors the AST substitution logic to correctly handle
nested expansions and improve the inlining of lambda expressions.
The changes include:
- Extracting the core value from potentially expanded AST nodes before
checking their kind.
- Ensuring that lambda expressions with empty upvalues are correctly
substituted.
- Adding support for inlining global get expressions as AST
substitutions.
- Improving the `UsageInfo` collection to correctly track assigned
values.
Introduce generic `CompilerPhase` trait to unify different stages of the
bound AST. Rename `LocalSlot` to `VirtualId` and introduce `StackOffset`
for clearer distinction between compile-time and run-time addressing.
Update type aliases and implementations to reflect these changes.
The VM's field access logic has been refactored to clone `Rc` pointers
for addresses and field identifiers. This change aims to improve
efficiency by avoiding unnecessary dereferencing and copying of values
when accessing fields within records and objects.
Additionally, the `Value::Object` handling in `BoundKind::GetField` has
been refined to more explicitly manage `RecordSeries` and `StreamNode`
types, ensuring that field access is handled polymorphically and
efficiently without copying underlying data structures.
This commit makes several changes related to how closures are handled:
- **Dumper:** Removes redundant introspection of closure ASTs, as this
is no longer relevant.
- **Optimizer:** Updates `try_inline` to accept `ExecNode` for
parameters and adds a check to ensure the `function_node` is a
`Lambda` before attempting inlining.
- **Environment:** Simplifies closure creation by passing an `ExecNode`
for parameters and ensuring the correct `stack_size` is used.
- **VM:** Adjusts `Closure` to store an `ExecNode` for parameters and
updates `VM::unpack` to accept an `ExecNode`.
Introduces a `StackAllocator` to manage local slot mapping during
lowering. This allows for efficient calculation of stack sizes for
lambdas and the root node. Tail position marking is now handled more
accurately.
Replace manual resizing of global arrays with simple pushes. This
simplifies the logic and removes the need for index calculations within
the registration functions.
Change the `slots` field in `TypeContext` from a `Vec` to a `HashMap`.
This allows for sparse local variable storage and avoids the need to
pre-allocate for all slots, which is more efficient when not all locals
are used.
Add configuration for the repomix MCP server to the Gemini settings.json
file. This will allow Gemini to use repomix for code analysis and
understanding.
This commit replaces the `UntypedNode` enum with the more accurately
named `SyntaxNode`. This change is primarily for clarity and better
reflects the role of these nodes as representing the structure of the
source code prior to semantic analysis.
The corresponding enum `UntypedKind` has also been renamed to
`SyntaxKind` to maintain consistency.
No functional changes are introduced by this refactoring; it is purely a
renaming and organizational update.
Simplified the AST architecture by removing the overly complex Node<K,
T> structure and replacing it with specialized types for
each phase:
- Introduced UntypedNode in nodes.rs for the Parser and Macro system,
eliminating redundant ty: () initializations.
- Defined a new generic Node<T> in bound_nodes.rs for all compiler
stages, where T represents the phase-specific metadata.
- Updated the Binder to act as the bridge between UntypedNode and
Node<()>.
- Simplified type aliases: TypedNode, AnalyzedNode, and ExecNode now
leverage the streamlined Node<T> structure.
- Updated Parser, Macros, Type-Checker, Optimizer, and VM to reflect
the architectural changes.
- Fixed import chains and resolved needless borrows via clippy.
This change improves type safety, reduces boilerplate code in the
parser, and makes the compiler stages more idiomatic and
readable.
This commit refactors the purity analysis from using a
`HashMap<GlobalIdx, Purity>` to a `Vec<Purity>` indexed by
`GlobalIdx.0`.
This change simplifies the data structure and makes purity lookups more
efficient. The `Binder`'s `globals` field has also been removed as it
was not being used.
The `Binder` struct now accepts a `Rc<RefCell<HashMap<Symbol,
(GlobalIdx, Identity)>>>` for global names. This allows the binder to
directly access and manage global symbols during the binding process,
simplifying the logic and improving efficiency.
Additionally, `CompilerScope` now implements `Default`, and a type alias
`BindingResult` has been introduced for clarity. The `bind_root`
function signature has been updated to accept the `globals` argument.
The `Binder` and `FunctionCompiler` initialization has been refactored
to accept initial scopes and slot counts. This allows for more flexible
management of the compiler's state, particularly for the root scope.
The `bind_root` function now returns the final scopes and slot count of
the root function compiler, enabling the `Environment` to update its
state with these new bindings.
The global variable handling has been integrated into the scope
management, removing the direct reliance on a shared `HashMap` for
globals. This promotes a more consistent approach to symbol resolution.
The `LocalInfo` struct now includes a `purity` field to track the purity
of local variables and function arguments. This is initialized to
`Purity::Impure` for local variables and arguments within functions and
when registering native functions. Additionally, the `Environment`
struct is updated to include `root_scopes` and `root_slot_count` to
support parallel mode compilation by populating the root scope with
initial function and native function information.
Removes global variables and related logic from the Binder, simplifying
its responsibilities to managing local scopes and function compilers.
This change also streamlines the `define_variable` and `resolve_local`
methods within `FunctionCompiler` and removes unused fields from
`FunctionCompiler` and `Binder`.
Introduces `fixed_scope_idx` to `Binder` and `Environment` to track
immutable scopes.
This allows enforcing that definitions (`def`) cannot occur in scopes
that are considered fixed or immutable, such as the global scope or the
initial bootstrap scope.
The `FunctionCompiler` is updated to use `fixed_scope_idx` and `is_root`
to determine scope immutability.
The TCO (Tail Call Optimization) module has been renamed to `lowering`.
This change better reflects the module's broader responsibility, which
includes not only TCO but also general AST transformations and
preparation for VM execution.
The `optimize` function has been renamed to `lower` to align with the
module's new name.
The `calculate_stack_size` method was duplicated in `bound_nodes.rs` and
`tco.rs`. This commit extracts it into a single standalone function in
`tco.rs` to avoid duplication. The stack size is now calculated and
stored in the `ExecNode`'s `ty` field during the TCO optimization phase.
Replaces manual iteration for pushing `Value::Void` with `Vec::resize`
for more efficient stack management. This change improves performance by
reducing redundant operations and simplifies the code.
Store the pre-calculated stack size of lambda bodies in the
RuntimeMetadata. This allows the VM to directly access this information
without recalculating it.
This commit introduces a new mechanism for calculating the required
stack size for closures at bind time, rather than storing it in the AST.
This is achieved by adding a `calculate_stack_size` method to
`BoundNode`.
Key changes include:
- `BoundNode::calculate_stack_size`: Recursively traverses the bound AST
to find the maximum `LocalSlot` index used.
- Recursion blocker for lambdas: Ensures that nested lambdas are not
visited during stack size calculation for the outer closure,
preventing incorrect stack size estimations.
- VM update: The `VM::run` method now accepts and uses the
pre-calculated stack size for setting up call frames.
- `Closure` struct update: Stores the `stack_size` directly within the
`Closure` struct.
- Binder and TypeChecker updates: Minor adjustments to accommodate the
new strategy and error reporting.
- Optimizer enhancements: Constant and pure-lambda propagation for
`Define` statements within blocks to ensure inlinability.
This refactoring addresses issues related to accurate stack size
management and improves the overall robustness of the binder and VM.
The previous implementation of the binder had a slightly complex way of
handling the root scope. This commit simplifies that logic by directly
checking if the current scope is the root scope and the scope stack has
only one element. This makes the code more readable and maintainable.
This commit refactors the Binder implementation to better align with the
updated scope structure.
Key changes include:
- Introducing `ExprContext` to distinguish between statement and
expression contexts.
- Modifying `define_variable` to handle both global and local scope
definitions more robustly.
- Updating `resolve_variable` to correctly handle upvalue captures,
especially for global addresses.
- Adjusting `bind` and related functions to use the new `ExprContext`.
- Updating the refactoring log to reflect completed phases.
Implement strict block scoping and distinguish between statements and
expressions.
This change introduces hierarchical scopes for functions and blocks,
enforces that `def` is a statement with no return value, and prevents
statements from appearing in expression positions or as the last element
of a block. A new example `design-flaw.myc` is added to demonstrate
the uninitialized variable issue solved by these changes.
The `Analyzer` is now responsible for determining the purity of `Define`
bindings. Previously, this responsibility was left to the caller.
Added a new example `design-flaw.myc` to test this change.
Updated `soa_series.myc` due to the purity change.
The 'again' macro is a new addition to the macro system, allowing for
recursive macro expansion. This commit introduces the necessary AST
nodes, macro expansion logic, and an integration test to ensure its
functionality. The `soa_series.myc` example has also been updated to
demonstrate its usage.
The `Parameter` kind in `UntypedKind` was only used for identifiers that
were being declared or referenced. This commit renames it to
`Identifier` and updates all the necessary code to reflect this change.
This simplifies the AST and makes it more consistent.
Additionally, a new macro `repeat` has been added to
`src/ast/system.myc`.
The `system.myc` file, containing macros and core utilities, has been
removed as its functionality is now handled by an embedded string. This
simplifies the build process and eliminates a file that was no longer
strictly necessary.
Additionally, the `with_cwd()` method on the `Environment` has been
removed. The environment now automatically adds the current working
directory and its `rtl` subdirectory to the search paths during
initialization. This streamlines the setup for file-based usage and
ensures standard search paths are always considered.
- Add `rtl/system.myc` as the system library, containing core utilities
like `while` and `cache`.
- Remove the duplicate `prelude.myc` from `src/ast/rtl/`.
- Modify `Environment::collect_dependencies` to implicitly add `#use
system` when necessary.
- Update `Environment::with_cwd` to also add the `rtl` subdirectory to
search paths.
- Add `target/` to `.geminiignore` to exclude build artifacts.
- Adjust example `sma.myc` to use the new `cache` macro and reduce
sample data size.
- Update `rtl/sma.myc` to correctly initialize the `history` series with
its `length`.
- Add `preload_dependencies` calls to `dump_ast` and `run_script` for
proper module loading.
The `PipelineNode` has been refactored into `StreamNode`. This commit
updates the example files to reflect this change and also updates the
`series` constructor to accept a lookback parameter.
The `PipelineNode` was an artifact of a previous implementation and is
no longer needed. The `StreamNode` is the appropriate type for
representing the output of a pipe operation.
The `series` function now requires a `lookback` argument to specify the
maximum number of items to store. This makes the behavior of series more
explicit and prevents accidental unbounded memory growth.
This commit extracts the logic for inlining function calls into a new
private
method, `try_inline`. This improves code organization and reduces
duplication in the `Optimizer`'s `visit_node` method.
The extracted method handles the preparation of substitutions and the
recursive visiting of the inlined node. It also ensures that the
substitution map state is correctly synchronized back to the calling
context.
The type checker now correctly handles unwrapping `Stream(T)` types for
lambda arguments.
A new `build_map_stream` function has been added to `rtl/streams.rs` to
facilitate field access on streams.
The `create-random-ohlc` function now returns a `Stream` instead of a
`Series`.
Examples have been updated to reflect these changes and demonstrate
stream usage.